<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[AI & Software Engineering at Production Scale — A Newsletter]]></title><description><![CDATA[Subscribe for practical architecture insights and proven engineering heuristics.]]></description><link>https://nidly.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!Kox9!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3870fb32-be58-47d8-819c-07dadafc1198_1024x1024.png</url><title>AI &amp; Software Engineering at Production Scale — A Newsletter</title><link>https://nidly.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 19:08:53 GMT</lastBuildDate><atom:link href="/__u/nidly.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Alireza Rahmani Khalili]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[alirezarahmani@live.com]]></webMaster><itunes:owner><itunes:email><![CDATA[alirezarahmani@live.com]]></itunes:email><itunes:name><![CDATA[Alireza Rahmani Khalili]]></itunes:name></itunes:owner><itunes:author><![CDATA[Alireza Rahmani Khalili]]></itunes:author><googleplay:owner><![CDATA[alirezarahmani@live.com]]></googleplay:owner><googleplay:email><![CDATA[alirezarahmani@live.com]]></googleplay:email><googleplay:author><![CDATA[Alireza Rahmani Khalili]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[AI Is Making Your Software Architecture Drift]]></title><description><![CDATA[What I learned after repeatedly approving AI-generated code that worked locally and weakened the system globally.]]></description><link>https://nidly.substack.com/p/ai-is-making-your-software-architecture</link><guid isPermaLink="false">https://nidly.substack.com/p/ai-is-making-your-software-architecture</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 31 Aug 2026 07:10:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zhQ2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!zhQ2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!zhQ2!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!zhQ2!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!zhQ2!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!zhQ2!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!zhQ2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2382402,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/209974860?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!zhQ2!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!zhQ2!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!zhQ2!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!zhQ2!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19e06cb-edad-4635-a3f4-4aa681eb0af5_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I was <a href="/__u/nidly.substack.com/p/i-replaced-my-entire-code-review?r=a3p8i">reviewing</a> AI-generated code for a relatively ordinary feature. The implementation was clean. The tests passed; The naming matched conventions we&#8217;d settled on months earlier. Nothing in the diff looked irresponsible.</p><p>I still did not like it. The model had solved the feature by crossing a boundary we had deliberately kept in place for years. It had imported a repository from another part of the system and used it directly, bypassing the application boundary that was supposed to sit in between.</p><p>It wasn&#8217;t doing anything dramatic. From the model&#8217;s point of view, the repository was there, accessible, and already used elsewhere. It looked like just another dependency.</p><blockquote><p>The code worked. That was the problem.</p></blockquote><p>I approved it anyway. That part matters, because the rest of this story makes less sense if I pretend I caught the problem immediately and stopped it. I didn&#8217;t. I left a small comment about extracting an interface later, treated it as follow-up work, and moved on.</p><p>I had approved several changes like this before I realized they were not separate implementation shortcuts. They were the same architectural decision being made repeatedly, without anyone explicitly deciding to make it.</p><div><hr></div><h2>The code that looked fine</h2><p>AI-generated code is often impressive at the level you check first. It follows the naming conventions you already have. It reuses abstractions instead of inventing new ones. It writes tests, sometimes more thorough ones than a tired engineer would write at 6pm on a Thursday. It handles the edge cases you&#8217;d expect, and occasionally catches one you missed.</p><p>This is exactly why the <a href="/__u/nidly.substack.com/p/architecting-software-for-the-ai?r=a3p8i">architectural</a> problem is easy to miss. Obviously bad code gets rejected in review almost by reflex. What gets merged, quietly, is the version that reads well.</p><p>Here&#8217;s a shape I saw repeatedly. A feature needs information that another part of the system owns. The intended path is something like an application interface, a small contract between Orders and Pricing, agreed on specifically to keep that dependency explicit and one-way.</p><p>But the AI, working inside the repository, discovers that both modules live in the same codebase and that the other module&#8217;s repository is one import away. So it uses it directly. A business rule that should be called through the owning domain gets reconstructed locally from the data that happens to be available.</p><p>Nothing crashes. The tests pass, because the tests were written against the feature, not against the boundary. The feature ships on time. And somewhere in that diff, a dependency we had deliberately kept unidirectional becomes bidirectional without anyone deciding that it should.</p><p>That&#8217;s the distinction I kept losing track of during review. Local correctness and architectural correctness are not the same question. Code review is very good at answering the first one. It is much less reliable at noticing when nobody asked the second.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Precedent compounds faster than you&#8217;d like</h2><p>The first shortcut is dangerous less because of what it does than because it becomes precedent.</p><p>Once that pattern exists in the repository, the next agent doesn&#8217;t see a compromise. It sees a convention. It has no way of knowing that this particular dependency was reluctantly allowed, or that someone intended to come back and remove it. It sees working code that reaches across the boundary, and when a similar problem appears, it has every reason to do the same thing again.</p><p>From a local point of view, that decision is completely reasonable. The first violation was a shortcut. The next five became the architecture. Humans do this too. Copy-paste has always been a mechanism for spreading bad patterns, and every senior engineer has seen a temporary hack outlive the context that justified it. Someone finds existing code, assumes the trade-off was intentional, and repeats it.</p><blockquote><p>What changes with AI is the replication rate.</p></blockquote><p>A questionable pattern no longer has to spread slowly through the engineers who happen to touch that part of the system. Coding agents repeatedly draw from the same local evidence: the code that is already there. Once an exception appears often enough, it starts being treated as the normal way to solve the problem. By the time I noticed that happening, the shortcuts no longer looked isolated. They had started to look like how we did things.</p><div><hr></div><h2>What the task never said</h2><p>It took me a while to understand what the agent was actually missing. A coding agent gets a task that looks something like &#8220;add support for cancelling this order.&#8221; The success criteria are legible and specific: implement the behavior, make the tests pass, stay consistent with the surrounding code, and don&#8217;t break anything that already works. Those are reasonable things to optimize for, and the agent is good at optimizing for them.</p><p>What the task usually does not say is that cancellation eligibility must remain inside the Order aggregate, that Payment can be queried through its application interface but its repository is off limits, or that a particular table is owned by another bounded context and nobody outside that context is allowed to mutate it. It rarely says that a business rule has exactly one source of truth, and that duplicating it, even accurately, is itself the violation.</p><p>Those are not implementation details. They are <a href="/__u/nidly.substack.com/p/why-software-architecture-matters?r=a3p8i">architectural</a> constraints, and in a lot of systems they live somewhere other than the ticket. They live in the heads of the engineers who were around when the boundary was drawn, in an ADR nobody has opened for a while, in a diagram that no longer quite matches reality, or in an old review comment that effectively meant, &#8220;we don&#8217;t do that here.&#8221;</p><p>A human who has worked on the system long enough carries a surprising amount of that context implicitly, often without being able to fully articulate the rule until they see someone violate it. An agent working from the code in front of it gets fragments of that context, if it gets any at all. It optimizes locally because local is what it was given.</p><div><hr></div><h2>Drift without a decision</h2><p>None of the individual instances I kept finding were catastrophic on their own. A few patterns just kept recurring.</p><p>There was dependency drift, where a module gradually started depending on things it was never meant to know about, one reasonable import at a time. There was domain leakage, where a business decision that belonged in the domain model ended up in a controller, a handler, or a piece of orchestration code because that was where the immediate task happened to be solved. I also kept seeing duplicated policy, where generated code reimplemented an existing rule locally, often correctly, because that was easier than tracing how to invoke the original from where the agent was working.</p><p>Ownership erosion was subtler. A service would start by reading data it could technically reach, then eventually take on responsibilities it was never supposed to own. Abstraction drift worked the same way. A new wrapper or helper would appear around an existing abstraction because the agent lacked the context for why the original one existed in that shape, and building something new looked cleaner than untangling what was already there.</p><p>None of these showed up as a single alarming pull request. Architecture rarely breaks in one merge. It drifts through accumulation, and that accumulation is easy to miss when every review is focused on the diff in front of you.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-is-making-your-software-architecture?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-is-making-your-software-architecture?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/ai-is-making-your-software-architecture?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p></p><div><hr></div><h2>The question code review doesn&#8217;t ask</h2><p>I want to be self-critical here rather than blame the tooling, because the honest version of this story is that code review, as I was practicing it, wasn&#8217;t built to catch this.</p><p>Reviewers ask a fairly consistent set of questions. Is this correct? Is this readable? Are there tests, and do they mean anything? Does this introduce an obvious performance problem? Is the implementation more complicated than it needs to be? Those are good questions, and I still ask them. They just cover a narrower kind of correctness than the one I was starting to worry about.</p><p>The question that would have caught much of what I&#8217;m describing is closer to: what architectural assumption does this change introduce that wasn&#8217;t there before? Or, less abstractly: if this became the normal way we solved similar problems, would I still approve it?</p><p>A single exception, reviewed in isolation, can be completely reasonable. Deadlines exist. Sometimes the architecturally cleaner path costs more than the situation justifies, and making that trade-off consciously is part of engineering. The danger isn&#8217;t the exception itself. It&#8217;s what happens after the exception becomes part of the codebase.</p><p>Once it is there, an agent can reproduce it cheaply and confidently without knowing that it was ever supposed to be exceptional.</p><div><hr></div><h2>The rules we never wrote down</h2><p>The uncomfortable conclusion I kept coming back to was that a lot of what I thought of as our architecture had never really been encoded anywhere. It lived in senior engineers&#8217; heads, in old design discussions, in diagrams that were accurate when they were drawn, and in the instinct a reviewer develops after spending enough time with a system to know when something feels wrong before they can fully explain why.</p><p>That works surprisingly well when experienced engineers are involved in most of the decisions. It becomes much more fragile when a growing share of implementation is produced by tools that were never part of the conversations where those boundaries were established.</p><p>AI did not ignore our architectural rules. In a lot of cases, we had never expressed those rules clearly enough for anything to ignore.</p><p>I initially framed this as a tooling failure. That explanation became less convincing the more I looked at the codebase. The tools were mostly doing what we asked them to do. The problem was that what we asked for was narrower than what the system actually needed.</p><div><hr></div><h2>Changing how I brief the agent</h2><p>This is where I&#8217;ve actually changed my day-to-day work, and it&#8217;s less about restraint and more about being specific in ways I used to consider unnecessary.</p><p>Before I ask an agent to implement anything nontrivial, I try to state ownership explicitly rather than assume it&#8217;s obvious from the code. Which domain owns this state? Which component is allowed to change it. That single sentence, added to a prompt that would otherwise have just been the feature description, has quietly prevented more boundary violations than any review comment I&#8217;ve written this year.</p><p>I&#8217;ve also stopped handing over tasks and started handing over constraints. Instead of &#8220;implement order cancellation,&#8221; something closer to &#8220;cancellation eligibility must remain inside the Order aggregate; Payment may be queried through its application interface; its repository is not accessible.&#8221; It reads like more work. It usually saves a follow-up conversation about why a table got touched that shouldn&#8217;t have been.</p><p>Before I accept an implementation, I&#8217;ve started asking the agent directly what it changed structurally, not just functionally. New dependencies introduced. Boundaries crossed. Assumptions made about who owns which data. Business rules that got introduced or, more often, quietly duplicated. It doesn&#8217;t catch everything, and it isn&#8217;t a substitute for actually reading the diff, but it surfaces things I would otherwise have to notice on my own, tired, at the end of a review queue.</p><p>And where it&#8217;s practical, I&#8217;ve been pushing more of this into things the build can actually enforce, because documentation alone has gotten weaker exactly as generation has gotten faster. Architecture tests. Explicit dependency rules between packages. Module visibility that isn&#8217;t just a naming convention everyone is trusting each other to respect. Schema ownership that&#8217;s asserted somewhere other than a wiki page. None of this is new advice. What&#8217;s new is how much more it matters when the rate of code production stops being bottlenecked by how many engineers you have typing.</p><div><hr></div><h2>Architecture becomes more important when code becomes cheap</h2><p>AI reduces the cost of producing implementation. It does not reduce the cost of a bad boundary. If anything, it makes that cost easier to accumulate, because a questionable pattern can now be reproduced across a system much faster than before.</p><p>Friction used to slow architectural drift almost by accident. Repeating a pattern meant another engineer had to encounter it, understand enough of the surrounding code to reproduce it, and make the same decision again. That did not prevent bad architecture, obviously, but it created opportunities for someone to question the pattern before it spread too far. Coding agents remove a lot of that friction. If the existing code provides enough evidence that a pattern is normal, the agent has little reason not to reproduce it.</p><p>So the value of architecture is shifting, at least in how I think about it. It is less about prescribing how every class should be written, which was always a losing battle anyway, and more about defining what the system is not allowed to become.</p><p>That also means expressing those constraints in forms that can actually be checked. Better diagrams and longer architecture documents are not enough if the important boundaries still depend on someone remembering why they exist.</p><blockquote><p>When code gets cheap, constraints get valuable.</p></blockquote><p>What AI made obvious to me is that architecture has to become explicit enough that neither a human under deadline pressure nor an agent working from local evidence can quietly violate an important boundary without something noticing.</p><div><hr></div><p><strong> Sources:</strong></p><ul><li><p>Neal Ford, Rebecca Parsons, Patrick Kua, <em><a href="https://www.oreilly.com/library/view/building-evolutionary-architectures/9781491986356/ch02.html">Building Evolutionary Architectures</a></em> &#8212; the original treatment of fitness functions as executable, continuous checks on architectural intent.</p></li><li><p><a href="https://arxiv.org/pdf/2201.01184">Symptoms of Architecture Erosion in Code Reviews: A Study of Two OpenStack Projects</a> &#8212; empirical look at how erosion actually surfaces in real review discussions, not just in theory.</p></li><li><p><a href="https://arxiv.org/pdf/2510.10165">AI-Assisted Programming Decreases the Productivity of Experienced Developers by Increasing the Technical Debt and Maintenance Burden</a> &#8212; data on the productivity-debt tradeoff this piece describes anecdotally.</p></li><li><p><a href="https://arxiv.org/html/2603.28592v2">Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild</a> &#8212; large-scale tracking of AI-introduced issues that survive in production repositories over time.</p></li><li><p>LeadDev, <a href="https://leaddev.com/technical-direction/how-ai-generated-code-accelerates-technical-debt">How AI-generated code compounds technical debt</a> &#8212; practitioner-facing summary of the GitClear and DORA findings on cloning and delivery stability.</p></li><li><p>InfoQ, <a href="https://www.infoq.com/articles/agentic-fitness-functions-evolutionary-architecture/">Agentic Fitness Functions: Extending Evolutionary Architecture Beyond Deterministic Rules</a> &#8212; a current attempt to extend fitness functions to judgment-heavy checks like boundary fidelity, which is close to what this article argues we need.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[AI Career Advice vs. Human Experts: Who Should You Trust?]]></title><description><![CDATA[Use AI to widen your thinking. Use humans to calibrate reality. Keep the decision authority for yourself.]]></description><link>https://nidly.substack.com/p/ai-career-advice-vs-human-experts</link><guid isPermaLink="false">https://nidly.substack.com/p/ai-career-advice-vs-human-experts</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 24 Aug 2026 06:30:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dQiD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!dQiD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!dQiD!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!dQiD!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!dQiD!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dQiD!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!dQiD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2285979,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/210928937?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!dQiD!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!dQiD!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!dQiD!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dQiD!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c74c42-783e-4c71-bdea-aa64a483f3b7_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Getting good career advice used to have a natural bottleneck: <strong>access</strong>. You needed someone who had seen enough careers, teams, and hiring cycles to recognize a pattern when they saw one. A manager who had watched people plateau and eventually understood why. A senior engineer who had made the wrong move once and remembered what it cost. A recruiter who spoke to dozens of companies and knew which job descriptions reflected a real role and which were mostly wish lists. Or, if you were lucky, a mentor willing to spend an hour telling you something you didn&#8217;t particularly want to hear.</p><p>That access was scarce. Most people didn&#8217;t have it consistently, and some only found useful advice after they&#8217;d already made the decision.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Now anyone can open a chat window and ask whether they should become an AI engineer, leave their company, aim for Staff Engineer, move from backend engineering into machine learning, or accept a particular offer. Within seconds, they get something that often sounds more structured and thoughtful than the advice they&#8217;d get from another person. It&#8217;s fast, available whenever you need it, and you can ask the same question three different ways without testing somebody&#8217;s patience.</p><p>That creates an obvious question, but I think the obvious question is the wrong one. The question isn&#8217;t whether AI can give career advice. It clearly can. The more useful question is what kind of advice it is actually good at, and where a confident answer starts running ahead of what it really knows.</p><div><hr></div><h2>AI is unusually good at analysis</h2><p>Take a common case: a senior backend engineer, maybe ten years into their career, is considering a move into AI engineering. A capable model can do a surprising amount of useful work here. It can look at the engineer&#8217;s existing experience and identify what transfers directly, whether that&#8217;s distributed systems, data pipelines, production infrastructure, APIs, or reliability work. It can compare that background against AI engineering roles and point out the missing pieces with much more precision than something vague like &#8220;learn AI.&#8221; The gaps might be retrieval architecture, evaluation, model serving, inference constraints, or simply understanding how operating a probabilistic system differs from running a conventional backend service.</p><p>It can also sketch several possible routes into the field instead of assuming there is one correct transition. It can take apart a job description and distinguish requirements that probably matter from things somebody copied into a hiring template. It can help prepare for interviews, compare roles, identify weaknesses in a resume, and, when asked properly, argue against the plan rather than simply reinforcing it.</p><p>None of that is trivial. It is useful analytical work: taking information, comparing it against a target, finding gaps, generating alternatives, and testing assumptions. AI is particularly strong at this kind of work because it can process a large amount of information quickly and revisit the same problem from several angles without much cost.</p><p>There is also a weakness in the usual comparison between AI and human advice. People tend to imagine AI on one side and an excellent mentor on the other. That is rarely the real choice. The real alternative is whatever advice is actually available to you, which might be a manager who has never made the transition you&#8217;re considering, a friend working in a different market, a recruiter whose view of the industry is shaped by the handful of roles they happen to be filling, or simply nobody at all.</p><p>Against that baseline, AI can be genuinely useful. A competent AI response can easily outperform bad human advice, and there is plenty of bad human advice around. It can be vague, outdated, overly shaped by one person&#8217;s career, or completely disconnected from the situation in front of you.</p><div><hr></div><h2>But analysis is not judgment</h2><p>This is where most of the confusion around AI career advice starts.</p><p>A model can correctly conclude that AI engineering is a growing field, that a backend engineer has transferable skills, and that learning a specific set of technologies would make them more employable. All of that can be true, and switching careers can still be the wrong decision.</p><p>Imagine that engineer is already operating at senior or near-principal level. Repositioning themselves as a junior or mid-level AI engineer may not be a step forward at all. It may throw away years of career capital. The stronger position could be to stay anchored in distributed systems and backend engineering while becoming unusually good at building AI systems in production.</p><p>Or maybe the technical direction was never the real problem. Their resume may simply undersell what they&#8217;ve done. Their market positioning may be weak. They may be working inside a dysfunctional company and interpreting frustration with the organization as evidence that they chose the wrong career. Changing disciplines doesn&#8217;t fix a bad manager, a broken team, or a company running out of money.</p><p>This is where AI can be very convincing for the wrong reason. It can analyze the problem you gave it extremely well without noticing that the problem itself was framed badly.</p><p>And a sophisticated answer to the wrong question is often more dangerous than a mediocre answer to the right one. A detailed transition plan feels like progress. It may still be taking you in the wrong direction.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-career-advice-vs-human-experts?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-career-advice-vs-human-experts?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/ai-career-advice-vs-human-experts?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2>AI has a built-in pressure to answer the question you asked</h2><p>If someone asks, &#8220;Should I become an AI engineer?&#8221;, the natural response is to answer that question. A good advisor may do something more annoying and more useful. They may ask why this person wants to become an AI engineer in the first place.</p><p>Are they bored? Did someone they know make the switch and get a large raise? Do they think backend engineering is becoming less valuable? Are they reacting to a bad six months at work and mistaking that feeling for a structural change in the market?</p><p>Those questions matter because career decisions are often downstream of a problem that was diagnosed badly.</p><p>AI can ask them too. In fact, if you explicitly tell a model to challenge your framing, identify hidden assumptions, and argue against your preferred option, it can do a surprisingly good job. But that usually requires the user to know that the framing itself needs to be questioned.</p><blockquote><p>That is a fairly big limitation.</p></blockquote><p>Most people come to career advice because they are uncertain about the decision. They are not necessarily uncertain about the assumptions underneath it. Those assumptions often arrive in the prompt disguised as facts.</p><p>So the model may have far more information-processing capacity than the person asking the question while still knowing much less about what is actually happening in that person&#8217;s career. If the starting premise is weak, better reasoning does not automatically save the answer. Sometimes it just makes the mistake harder to notice.</p><div><hr></div><h2>Human experts fail differently, not less</h2><p>None of this is an argument for defaulting to a human instead. Humans fail in different ways, and &#8220;real experience&#8221; is not automatically reliable just because it came from a person.</p><p>A manager who tells you to stay may genuinely think it&#8217;s the right move, while also knowing that losing you would make their own quarter harder. A recruiter&#8217;s view of the market is inevitably shaped by the roles they happen to work on. A founder who recommends starting a company may be generalizing heavily from the one career they know best: their own. And when a Staff Engineer explains how to become a Staff Engineer, what you&#8217;re sometimes hearing is one successful path presented as if it were the path.</p><p>Experience is useful, but it doesn&#8217;t arrive cleanly. People carry old lessons into new markets. They overweight things that hurt them personally. They remember the decisions that worked and are much worse at accounting for the people who made similar decisions and disappeared from view.</p><p>Self-interest matters too. Career advice often comes from people who are not neutral observers. Your manager, recruiter, co-founder, colleague, and even mentor may have incentives that overlap with yours only partially.</p><p>So &#8220;ask someone experienced&#8221; isn&#8217;t enough. You need the right experience for the question you&#8217;re asking, and you still need to understand where that person&#8217;s view might be distorted.</p><div><hr></div><h2>Where a good human still has a real edge</h2><p>There is still a kind of judgment where a good human can be much more useful than AI: knowledge built from direct exposure to a particular market, company, hiring process, or type of career.</p><p>Someone who has hired senior engineers for years might look at a role you technically qualify for and tell you that the company almost never hires Staff Engineers externally. They may look at your resume and tell you the document itself is fine, but the way you&#8217;re positioning your career is wrong. Those are not the same problem.</p><p>They may notice that a title which looks like a promotion inside one company will mean almost nothing to the next employer. They may know that a team hiring aggressively is doing it because people keep leaving, not because the business is growing. Or they may tell you that the certification you&#8217;re considering adds almost nothing at your level, while two strong production examples would materially change how you&#8217;re evaluated.</p><p>The important part is not that a human somehow has access to mystical knowledge AI can never learn. It&#8217;s that some career information is local, recent, informal, and rarely documented well. It lives in hiring rooms, recruiter conversations, failed promotions, internal politics, and the gap between what companies say publicly and what they actually do.</p><p>A strong human expert has sometimes watched those consequences happen in real time. That makes their pattern recognition different. Not automatically better, and certainly not unbiased, but grounded in a kind of evidence that may never appear in your prompt.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-career-advice-vs-human-experts?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-career-advice-vs-human-experts?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/ai-career-advice-vs-human-experts?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2>A concrete comparison, worked through</h2><p>Go back to the backend engineer considering a full move into AI engineering. Put the AI analysis next to the human challenge, because this is where the distinction becomes easier to see.</p><p>The AI analysis can be accurate and genuinely useful. Ten years of distributed systems work transfers well into building production AI systems. APIs, infrastructure, reliability, data pipelines, and operational experience do not suddenly stop mattering because a model is involved. Retrieval, agents, evaluation, and inference serving can be learned as extensions of an existing engineering base rather than as a completely separate profession.</p><p>A strong human expert might hear the same plan and ask a different question: why give up a decade of seniority to compete directly with people who spent those same years specializing in machine learning?</p><p>Maybe the better move is not to reposition as an engineer trying to enter AI. Maybe it is to sharpen the existing position: a senior distributed systems engineer who can build AI systems that survive production.</p><p>That is a very different career move.</p><p>The AI plan may not be wrong. It may simply be answering &#8220;How do I become an AI engineer?&#8221; when the more useful question is &#8220;What part of my existing experience has become more valuable because of AI?&#8221;</p><p>Sometimes the best move is not changing careers. It is changing how the market understands the career you already have.</p><p>That kind of judgment depends heavily on the person, their existing leverage, and the market they are operating in. A resume-to-job-description comparison can help, but it does not settle it.</p><div><hr></div><h2>The better use of AI is not asking it to decide</h2><p>Given the failure modes on both sides, I think one of the strongest uses of AI in a career decision is not asking it for a verdict at all. Use it to interrogate the decision before you commit to one.</p><p>Instead of asking, &#8220;Should I take this job?&#8221;, ask what assumptions the decision depends on. Ask what information is missing. Ask what would have to be true for the move to look like a mistake two years from now. Ask what the downside actually looks like rather than describing it as a generic &#8220;risk.&#8221;</p><p>Ask it to argue against the option you already prefer. Ask what evidence would make it change its recommendation. Ask it to separate facts from assumptions and speculation, because career decisions have a habit of mixing all three together until they sound equally solid.</p><p>This is a better role for AI than pretending it is the final authority. It becomes a way to pressure-test your thinking.</p><p>That matters because recommendations tend to close the conversation too early. A strong adversarial pass usually does the opposite. It gives you more reasons to look again.</p><div><hr></div><h2>Spend human time where human judgment is actually expensive</h2><p>Good human advice is scarce, so spending it on work AI can already do reasonably well is usually a waste.</p><p>If you have an hour with an experienced engineering leader, using most of that hour to rewrite resume bullets makes little sense. AI can help with the wording. The human conversation is more valuable when it gets into questions like whether the target role is actually realistic, how your background will be interpreted by the market, which signals matter more than they appear to, and whether you are solving the right problem in the first place.</p><p>This is also where someone with current hiring experience can be useful. They may notice quickly that a role is a poor fit for reasons that never appear in the job description, or that you are underestimating one part of your background while overinvesting in another.</p><p>The AI work should happen before that conversation, not instead of it. If the basic analysis is already done, the human does not have to spend half the discussion reconstructing the problem from scratch.</p><blockquote><p>That is a much better use of scarce expertise.</p></blockquote><div><hr></div><h2>A decision process, not a framework</h2><p>In practice, I would use both, but not symmetrically. I would start with AI to map the problem, compare alternatives, surface assumptions, and challenge the obvious answer. Then I would form a provisional view of my own.</p><p>After that, I would take the parts that depend heavily on local context, timing, reputation, hiring behavior, or company-specific knowledge to one or two people who have actually seen those situations up close.</p><p>Then I would pressure-test that human advice too. A human saying something with confidence does not make it true. Neither does an AI model producing five well-structured paragraphs.</p><p>The point is not to alternate mechanically between machine and human until a decision falls out. It is to use each source where it has the strongest information advantage, while keeping both open to challenge. Neither gets authority automatically.</p><div><hr></div><h2>Who should you trust?</h2><p>Neither, by default. AI deserves more trust when the problem is analytical and the reasoning can be inspected: comparing options, finding gaps, generating alternatives, or stress-testing a plan.</p><p>A good human deserves more weight when the decision depends on things that are harder to capture cleanly in a prompt: incentives, timing, reputation, organizational context, and current knowledge of a particular market.</p><p>Both can still be confidently wrong. AI can produce excellent reasoning from a bad premise. A human can take one successful career, usually their own, and turn it into a rule for everybody else.</p><p>AI can help you understand the decision. A good human can help you check that understanding against reality. But neither of them absorbs the downside if the decision goes badly. That is why the final authority should stay with you.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Is Breaking Your Domain Model]]></title><description><![CDATA[How prompts, agents, and retrieval pipelines are quietly becoming a second source of business truth in domain driven design.]]></description><link>https://nidly.substack.com/p/ai-is-breaking-your-domain-model</link><guid isPermaLink="false">https://nidly.substack.com/p/ai-is-breaking-your-domain-model</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 17 Aug 2026 07:28:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Xnud!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Xnud!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Xnud!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Xnud!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Xnud!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Xnud!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Xnud!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3029388,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/208039786?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Xnud!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Xnud!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Xnud!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Xnud!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f92bb2b-e18e-4442-a922-0681ea11f044_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Software teams spent years learning not to put business rules where they did not belong. Controllers were supposed to route requests, not decide whether an order could be cancelled. UI handlers were supposed to render forms, not enforce credit limits. Database triggers were supposed to preserve data integrity, not encode discount policies.</p><p>Most senior engineers have inherited a codebase where a rule such as &#8220;if the customer is premium&#8221; was scattered across controllers, background jobs, database queries, and frontend conditionals. They also remember how difficult it was to trust the system again, even after the rule had finally been given a single, explicit home.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Then AI systems arrived, and much of that discipline quietly disappeared. Consider two versions of the same decision:</p><pre><code><code>Controller decides whether an order can be cancelled.</code></code></pre><pre><code><code>Prompt asks the agent whether the order should be cancelled.</code></code></pre><p>The second version feels more advanced. It may not be.</p><p>It can be the same misplaced business logic wearing a newer costume, hidden behind natural language and model reasoning instead of an <code>if</code> statement. The fact that the rule is harder to search for, test, and trace does not make the architecture more sophisticated. It usually means the rule has become less explicit.</p><blockquote><p>We spent decades learning not to bury business rules in controllers. Rebuilding them inside prompts is <em>not</em> progress.</p></blockquote><p>This is not an argument against AI agents, nor is it an argument that language models should never interact with production systems. It is an argument about where authority should live once a model can interpret context, call tools, and mutate application state.</p><p>Once an agent can do all three, the important question is no longer whether the model is intelligent enough to make a decision. The question is whether that decision belongs to the model in the first place.</p><p>Where does the domain model actually end?</p><div><hr></div><h2>The Model Is Quietly Becoming the Domain</h2><p>A model stops being merely a text-generation component when it begins making decisions that previously belonged to the domain.</p><p>That includes deciding what an input means, which policy applies, whether an exception is justified, which action should happen, whether a tool should be called, and how application state should change as a result. Consider a customer refund system. Somewhere inside its system prompt, the team may have written instructions like these:</p><pre><code><code>Refund loyal customers unless fraud risk is high.
Escalate enterprise customers before rejecting a request.
Do not approve requests without sufficient evidence.</code></code></pre><p>The team may think of this as prompt engineering, in the same category as adjusting tone, formatting, or response length. But these instructions are not about presentation. Like: </p><ul><li><p>&#8220;Unless fraud risk is high&#8221; defines an exception policy.</p></li><li><p>&#8220;Escalate enterprise customers&#8221; defines routing behaviour based on account classification.</p></li><li><p>&#8220;Sufficient evidence&#8221; defines an eligibility condition.</p></li></ul><p>These are business rules expressed in English instead of code. They may be stored in a prompt registry rather than a domain module, but they still influence which actions the system considers valid.</p><p>The problem is that they often live in an artefact with weaker guarantees than the domain model they are replacing. Ownership may be unclear. Tests may cover outputs rather than invariants. Versioning may record that the text changed without explaining which business policy changed with it. A small wording adjustment can alter production behaviour without looking like a domain change during code review.</p><p>The moment a prompt determines what is allowed, it stops being presentation logic and starts becoming domain logic. There is an important distinction between using a model to interpret information and allowing it to determine business validity.</p><p>Extracting the customer&#8217;s stated reason for requesting a refund is interpretation. Identifying a reference to a duplicate payment is interpretation. Classifying an attached document as an invoice is interpretation.</p><p>Deciding that the stated reason qualifies for a refund is a business decision. Determining whether the payment is legally refundable is a business decision. Approving the transition from <code>RefundRequested</code> to <code>RefundApproved</code> is a business decision. Those responsibilities should not be collapsed into the same model call merely because a language model is capable of producing all three answers.</p><p>Yet this is how many agent systems are built. The same prompt interprets the request, evaluates the policy, chooses the action, and calls the tool that mutates state. The entire sequence is then described as &#8220;the AI layer,&#8221; as though interpretation, authority, and execution were one architectural responsibility.</p><p>AI can help the system understand what the user is asking and what the available evidence appears to show. That does not mean AI should decide what the system is permitted to do.</p><blockquote><p>A model may infer intent. The domain must still determine validity.</p></blockquote><div><hr></div><h2>You Now Have Two Domain Models</h2><p>Once business rules begin moving into prompts and agent workflows, most systems end up running two domain models at the same time, whether anyone intended that architecture or not.</p><p>The first is explicit. It consists of aggregates, entities, value objects, invariants, state machines, domain services, and policies. It appears in architecture diagrams, gets discussed during design reviews, and changes through ordinary code review.</p><p>The second is implicit. It is distributed across system prompts, tool descriptions, agent routing instructions, retrieval filters, evaluation datasets, orchestration conditions, human review guidelines, and exception-handling instructions added by whoever happened to be debugging the agent late at night.</p><p>Nobody deliberately designed this second model. It accumulated one fix at a time. The formal domain model does not disappear when this happens. It simply stops being the only component determining how the system behaves. Consider order cancellation. The explicit rule inside the aggregate says:</p><pre><code><code>A shipped order cannot be cancelled.</code></code></pre><p>A prompt, written by someone else on another day to solve a different problem, says:</p><pre><code><code>Consider cancelling the order if the customer has a strong reason.</code></code></pre><p>What happens when those instructions conflict?</p><p>In many systems, the honest answer is that whichever code path executes first wins. Nobody chose execution order as the mechanism for resolving conflicting business policies, but that is effectively what the architecture does. This is how semantic drift begins.</p><p>The domain model may evolve through months of refactoring while the prompts remain unchanged since the last production incident. Or the reverse may happen: prompts are adjusted every week to improve agent behaviour, while the underlying domain policies remain untouched because changing them requires deeper review, migrations, and coordination.</p><p>Prompts are easier to change, so they often change faster. That convenience becomes dangerous when prompt changes alter business meaning without being treated as domain changes.</p><p>You can call this an implicit domain model, a parallel model, or a shadow domain model. The label matters less than the consequence: it is influencing real decisions without clear ownership, authority, or consistency guarantees.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-is-breaking-your-domain-model?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-is-breaking-your-domain-model?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/ai-is-breaking-your-domain-model?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2>Nondeterminism Is Not the Real Problem</h2><p>The obvious objection is that language models should not make business decisions because their outputs are nondeterministic. That is a legitimate concern, but it is not the deepest one.</p><p>Imagine a model that always produced exactly the same output for the same input. No sampling variation, no inconsistent reasoning, no unpredictable tool selection. A perfectly deterministic model, but the architectural problems would remain.</p><p>Who owns a business rule stored inside the prompt? Where is that rule versioned, and against which domain history? Who tests it? How is it audited when a customer, regulator, or operations team asks why a decision was made?</p><p>What happens when the prompt conflicts with an aggregate or domain policy? Which bounded context owns the meaning of the language used in the prompt? Can the rule actually be enforced, or can the model only be instructed to follow it?</p><p>The problem is not only that the model may produce a different answer. The deeper problem is that the rule lives in a place where validity cannot be enforced. Determinism would make the behaviour more predictable. It would not make the decision legitimate.</p><p>A rule becomes legitimate because its ownership is clear, its meaning is explicit, its changes are controlled, and the system can enforce it regardless of who or what requests an exception. Prompts provide none of those properties by default.</p><blockquote><p>A perfectly consistent model can still apply the wrong policy consistently.</p></blockquote><div><hr></div><h2>AI Turns Invariants Into Suggestions</h2><p>There is a specific failure mode worth naming directly: an invariant is weakened when it is translated into prompt language. An invariant is not guidance. It is a condition the domain refuses to violate. A domain model might enforce rules such as:</p><pre><code><code>A captured payment cannot be captured again.
A suspended account cannot place an order.
A credit limit cannot be exceeded.
A shipped order cannot be cancelled.</code></code></pre><p>Now consider what happens when the same rules are rewritten as agent instructions:</p><pre><code><code>Avoid charging the customer twice.
Do not normally approve orders above the credit limit.
Be cautious when processing suspended accounts.</code></code></pre><p>&#8220;<strong>Avoid</strong>&#8221; does not mean &#8220;cannot.&#8221;</p><p>&#8220;<strong>Do not normally</strong>&#8221; implies that an exception may be justified.</p><p>&#8220;<strong>Be cautious</strong>&#8221; describes an attitude, not a boundary.</p><p>Language models are designed to interpret context, resolve ambiguity, and produce reasonable exceptions. That ability is useful when a system needs judgement. It becomes dangerous when applied to rules that must remain non-negotiable.</p><p>Given enough context, a model can generate a fluent and persuasive explanation for why a rule should not apply in one particular case. That is not necessarily a failure of the model. It is often exactly what the model was trained to do. The architectural failure is placing an invariant inside a component built to reinterpret language.</p><p>An aggregate should not be persuadable. It should not care how compelling the explanation is, how important the customer appears to be, or how confident the model sounds.</p><blockquote><p>An invariant that can be negotiated by a prompt is no longer an invariant.</p></blockquote><p>The domain layer must reject an invalid transition regardless of the quality of the reasoning used to propose it. A model may provide evidence, context, and a recommendation. None of those things override the current state of the domain.</p><blockquote><p>Confidence is not correctness, and correctness is not authority.</p></blockquote><p>A system that collapses those three concepts into a single model score will eventually approve something it was supposed to make impossible.</p><div><hr></div><h2>Intent Is Not Authority</h2><p>This distinction is the foundation of the entire argument. There is a clean boundary available here, but most agent architectures blur it without realising what they are giving away. AI is useful for extracting intent, interpreting unstructured evidence, classifying documents, detecting ambiguity, identifying contradictions, proposing actions, and reporting inferred facts. These are interpretive responsibilities. They help the system understand what appears to be happening, but they do not require the authority to decide what the business is allowed to do. That authority should <strong>remain</strong> <em>with</em> the domain model.</p><p>The domain should decide whether the proposed action is valid, whether the actor is authorised, whether the current state permits the transition, whether contractual or financial constraints have been satisfied, whether compensation is required, and whether the action has already happened.</p><blockquote><p>AI expresses intent. The domain grants authority.</p></blockquote><p>Consider the order cancellation example again. The model produces a structured proposal:</p><pre><code><code>{
  "intent": "cancel_order",
  "reason": "duplicate_purchase",
  "evidence_ids": ["order_192", "payment_883"],
  "confidence": 0.94
}
</code></code></pre><p>This is a useful output. It gives the application a clear interpretation of the request, identifies the reason, points to supporting evidence, and communicates the model&#8217;s confidence; But it is still only a proposal.</p><p>The domain model must check whether the order is cancellable, whether it has already shipped, whether the requester is authorised to cancel it, whether it was already cancelled, whether cancellation creates a refund obligation, and whether the current state permits the transition at all.</p><p>A confidence score of <code>0.94</code> tells you that the model is fairly certain about its interpretation of the evidence. It tells you nothing about whether the operation is valid. Finally, Confidence is not authority.</p><div><hr></div><h2>The Application Layer Should Coordinate, Not Decide</h2><p>None of this makes the application layer irrelevant, and it is not an argument for moving every line of AI-related code into the domain. That would simply create a different kind of architectural mess, which humans remain remarkably efficient at manufacturing.</p><p>The application layer still has a substantial role. It loads the relevant domain state, gathers external context, invokes the model, enforces technical permissions, coordinates dependencies, calls external services, manages workflow progression, passes structured proposals into the domain, and persists the transition the domain authorises.</p><p>That is a full responsibility. It is simply not the responsibility of deciding what is valid.</p><pre><code><code>User Request
    &#8595;
AI Interpretation
    &#8595;
Structured Intent or Proposal
    &#8595;
Application Coordination
    &#8595;
Domain Validation
    &#8595;
Authorised State Transition
</code></code></pre><p>The dangerous shortcut is an application service that treats a model tool call as an already-approved database operation.</p><p>An agent says &#8220;cancel the order,&#8221; and the application translates that directly into:</p><pre><code><code>orders.update(status="cancelled")
</code></code></pre><p>At that point, the domain model is no longer protecting the transition. The model proposes the action, but nobody verifies whether the action belongs in the current state.</p><p>A safer design preserves the separation. The model proposes a command. The application layer coordinates the work. The domain decides whether the command is valid.</p><p>Skip the final step, and the agent has more effective authority over the order lifecycle than the domain model itself.</p><div><hr></div><h2>Retrieval Quietly Changes the Domain</h2><p>Once retrieved documents begin influencing business decisions, RAG stops being purely an information-retrieval concern.</p><p>Retrieval determines which version of a policy the model sees, which evidence is treated as relevant, which tenant&#8217;s data enters the context window, whether revoked information remains visible, and which interpretation of the situation appears best supported.</p><p>Consider an insurance claim system. The retrieval pipeline returns an expired policy document, and the model recommends approving the claim based on that document. The model has not hallucinated. It read the evidence correctly and followed the instruction it was given. The retrieval pipeline supplied the wrong business context.</p><p>That is not a prompt-quality problem. Once retrieved context influences business decisions, retrieval becomes part of the domain boundary. But that does not mean retrieved text should authorise state transitions. It should inform interpretation, while the domain still validates applicability, effective dates, tenant scope, permissions, and current state independently.</p><p>A retrieved policy may say that a claim is covered. The domain still needs to verify that the policy was active when the event occurred, that it belongs to the correct customer, that the claim has not already been settled, and that the actor is authorised to approve it.</p><p>Evidence identity, document version, freshness, and access control are therefore not merely retrieval infrastructure details. They become conditions the domain must verify before treating the evidence as applicable.</p><div><hr></div><h2>Prompt Vocabulary Can Fork the Ubiquitous Language</h2><p>There is a DDD-specific version of this problem that is easy to dismiss as a writing issue, even though it is really a modelling issue.</p><p>Prompts often use soft, natural-language terms:</p><pre><code><code>good customer
reasonable refund
suspicious behaviour
high-value account
deserves compensation
</code></code></pre><p>A well-defined domain model should use more precise concepts:</p><pre><code><code>EligibleForRefund
FraudReviewRequired
EnterpriseAccount
PaymentDisputed
CompensationApproved
</code></code></pre><p>This is not a stylistic difference.</p><p>&#8220;Deserves a refund&#8221; is a subjective judgement. Its meaning depends on the reader, the surrounding context, and the story the model constructs from the available evidence.</p><p>&#8220;Eligible for refund&#8221; is a domain concept. It should have explicit criteria that either hold or do not. When a prompt uses terminology that differs from the bounded context it operates within, the system no longer has inconsistent naming alone. It has two semantic models running in parallel.</p><p>One model is encoded in code, policies, state transitions, and ubiquitous language. The other is encoded in natural-language instructions that may have been added gradually to improve agent behaviour.</p><p>A prompt that speaks a different language from the domain is not merely imprecise. It is modelling a different domain.</p><p>The damage usually appears later. Tests fail for reasons nobody can describe cleanly. Incident reviews turn into arguments over whether &#8220;high-value customer&#8221; means revenue, account tier, lifetime value, or strategic importance. Policy changes are implemented in code but never reflected in the prompt, or the prompt changes while the domain concept remains untouched.</p><p>The names may look similar enough to survive code review, but the meanings slowly diverge.</p><div><hr></div><h2>Agents Create Orphan Decisions</h2><p>As agent autonomy increases, decisions begin to emerge from the interaction of multiple components at once: user input, system prompts, retrieved documents, model output, tool descriptions, tool results, workflow state, retry behaviour, and human review instructions.</p><p>The final action may have real financial, contractual, or operational consequences, but no single artefact clearly explains why it was authorised. That is an orphan decision.</p><p>During an incident review, the important questions are not whether the model produced a plausible explanation. The real questions are:</p><ul><li><p>Who owned the decision?</p></li><li><p>Which rule authorised it?</p></li><li><p>Which evidence supported it?</p></li><li><p>Which prompt version influenced it?</p></li><li><p>Which retrieved documents were used?</p></li><li><p>Can the decision be reproduced?</p></li><li><p>Can it be challenged or appealed?</p></li><li><p>Did AI propose the action while the domain approved it, or did the action simply happen because nothing prevented it?</p></li></ul><p>Storing the model&#8217;s chain of thought is not a substitute for answering those questions.</p><p>Reasoning traces may help engineers debug model behaviour, but they are not stable, structured, or reliable enough to serve as an audit trail. They describe a generated explanation, not necessarily the actual source of business authority.</p><p>What the system needs instead is an explicit decision record.</p><p>That record should connect the outcome to a specific domain policy, the relevant state at the time of the decision, evidence identifiers, actor identity, prompt and model versions, and the version of the rule that authorised or rejected the transition.</p><p>Without that structure, the system cannot reconstruct why a decision occurred. It can only replay the conversation around it. That is not auditability. It is a transcript.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>A Safer Architecture</h2><p>The solution is not another abstract principle. It is an explicit pipeline that prevents interpretation, coordination, and authority from collapsing into one another when deadlines arrive and architectural discipline becomes inconvenient.</p><h3>Step 1: AI Produces a Structured Interpretation</h3><p>The model returns a proposal such as:</p><pre><code><code>{
  "proposed_action": "approve_refund",
  "reason_code": "duplicate_charge",
  "evidence_ids": ["payment_124", "invoice_882"],
  "confidence": 0.91
}
</code></code></pre><p>At this stage, the model has interpreted the available information and proposed an action. It has not changed state, approved a refund, or granted itself permission to do either.</p><p>That distinction is small in code and enormous in architecture.</p><h3>Step 2: The Application Layer Validates the Proposal Contract</h3><p>Before the proposal reaches the domain, the application layer validates its technical contract.</p><p>It checks the schema, tenant scope, evidence existence, access permissions, supported action types, model and prompt versions, idempotency information, and tool permissions. These checks are not yet about whether the refund is valid. They establish that the proposal is structurally acceptable and safe to evaluate.</p><p>This work is not glamorous, but neither is recovering from a cross-tenant refund caused by a malformed tool call.</p><h3>Step 3: The Domain Evaluates Business Validity</h3><p>The domain policy then evaluates the proposal against the current business state.</p><p>Is the invoice refundable? Was the payment actually duplicated? Has a refund already been issued? Is the requester authorised? Is the refund window still open? Does the amount require a higher approval level? Does this case require manual review?</p><p>These questions cannot be answered by confidence alone. They depend on domain state, policy, and invariants that may not have been visible to the model.</p><h3>Step 4: The Aggregate Performs the Transition</h3><p>Only after the proposal passes domain validation should the aggregate perform the transition:</p><pre><code><code>refund.approve()
</code></code></pre><p>The aggregate protects its invariants and emits the appropriate domain event, exactly as it would if the command had come from a user interface, an internal service, or a scheduled process.</p><p>The caller may be new. The rules should not be.</p><h3>Step 5: The System Records the Proposal and Decision Separately</h3><p>The model&#8217;s proposal and the domain&#8217;s decision should remain distinct records.</p><p>The audit trail should preserve the AI proposal, evidence identifiers, model version, prompt version, retrieved document versions, domain policy version, final decision, rejection reason, and actor identity.</p><p>Flattening all of that into a log entry such as &#8220;refund approved by agent&#8221; destroys the distinction the architecture is supposed to protect.</p><p>This separation allows the model to be wrong without corrupting state. The domain may reject a proposal that sounds entirely reasonable because it violates an invariant the model could not see.</p><p>That rejection is not a failure of intelligence. It is the system working as designed.</p><h2>What AI Should Be Allowed to Own</h2><p>None of this means that agents should never mutate state. That rule would be too absolute to be useful. The better question is which state an agent may change, under what constraints, and with what recovery path.</p><p>AI is well suited to tasks where the core difficulty is interpretation. It can extract facts from unstructured text, classify documents, map natural language to known intents, summarise evidence, identify uncertainty, detect contradictions, recommend actions, and request missing information.</p><p>That is already a broad and valuable role. There is no need to inflate it into unrestricted business authority.</p><p>A practical boundary should reflect risk:</p><pre><code><code>Low risk:
AI may act directly within narrow, reversible boundaries.

Medium risk:
AI proposes. The domain validates.

High risk:
AI proposes. The domain validates. A human authorises.
</code></code></pre><p>The category depends on reversibility, financial impact, legal consequences, potential user harm, security exposure, and the strength of the domain constraints already in place.</p><p>An agent changing the display order of dashboard widgets is not the same as an agent approving a loan, cancelling a shipment, or closing an insurance claim. Treating all state mutations as equally dangerous is simplistic, but treating all tool calls as equally acceptable is worse.</p><p>A narrow, reversible operation may be safe for an agent to execute directly, provided the allowed action space is constrained and domain validation still runs underneath it. The real rule is not that AI must never mutate state. It is that AI must not become the source of business validity.</p><h2>Prompts Are Domain-Adjacent Artefacts, Not Domain Models</h2><p>Prompts that influence domain behaviour deserve stronger governance than ordinary configuration.</p><p>They should have explicit ownership, versioning, a defined bounded-context scope, clear input and output contracts, allowed and prohibited actions, evaluation suites, rollout and rollback procedures, auditability, and terminology aligned with the domain.</p><p>But governance is not authority. Treating prompts as serious engineering artefacts does not make them a valid replacement for aggregates, policies, state machines, or invariants. <strong>Prompt governance reduces risk</strong>. It does not turn natural-language instructions into enforceable business rules.</p><p>A prompt registry can tell you which instruction version produced a particular output. It can help reproduce behaviour, compare evaluations, and roll back a bad release. What it cannot do is guarantee that the resulting state transition was valid.</p><p>Those are different guarantees. Version control can tell you what the prompt said. Only the domain can determine whether the action was allowed.</p><p>Teams get into trouble when they begin trusting prompt versioning as though it provides the same protection as domain invariants. It does not. It improves traceability around the proposal, but it does not legitimise the decision.</p><div><hr></div><h2>Trade-Offs and Limitations</h2><p>This architecture has a real cost, and pretending otherwise would make the argument less useful. Keeping authority in the domain model requires explicit contracts between AI and application code.</p><p> It means using structured outputs instead of unconstrained tool calls, validating technical and business concerns at separate layers, generating more audit data, and accepting that some model proposals will be rejected.</p><p>It is also slower to build than connecting an agent directly to a tool. That is intentional. The fastest demo is usually the one where the model chooses a tool and executes it immediately. It looks autonomous, requires fewer components, and works beautifully during the path everyone rehearsed beforehand.</p><p>Production systems have different requirements. They must survive stale context, duplicate requests, partial failure, policy changes, unauthorised actors, cross-tenant data, and model outputs nobody anticipated.</p><p>A demo only has to work once in front of people who already want to be impressed. A production system has to remain correct when nobody is watching.</p><p>The cost is not only technical. Restricting agent authority can also make the system appear less intelligent. More proposals are rejected, more actions require confirmation, and some workflows become slower. But apparent autonomy is a poor optimisation target when the system controls money, contracts, access, or irreversible state.</p><p>Not every application needs full Domain-Driven Design to follow this principle. A straightforward CRUD system may not need aggregates, value objects, or elaborate domain services. Explicit service-level validation and a constrained command boundary may be enough.</p><p>The pattern is not the point. The point is that a probabilistic component should not be the only thing standing between a request and a consequential state change.</p><div><hr></div><h2>Correctness Still Needs a Home</h2><p>Software architecture improved when teams gave business rules an explicit home. Rules became easier to test, reason about, review, and change because they no longer lived across controllers, database triggers, UI conditionals, and integration code. AI systems risk reversing that progress.</p><p>The same business meaning is now being distributed across prompts, retrieved documents, tool descriptions, evaluation datasets, routing instructions, and orchestration logic. Each artefact may look harmless in isolation, but together they form a behavioural model that no single team fully owns.</p><p>The biggest architectural risk of AI is not simply that the model may be wrong. It is that teams stop knowing where correctness is supposed to live.</p><p>AI may interpret what happened, infer what the user wants, and propose what should happen next. The application layer may coordinate the work required to carry it out. But the domain model must still decide whether that next step is valid.</p><p>Once business rules move into prompts, retrieval pipelines, and agent instructions, you still have a domain model. You just no longer know where it is.</p><div><hr></div><h2>Sources</h2><ol><li><p><strong>Eric Evans, Domain-Driven Design Reference</strong><br>A useful reference for bounded contexts, ubiquitous language, aggregates, and the role of the domain model.<br><a href="https://www.domainlanguage.com/ddd/reference/">https://www.domainlanguage.com/ddd/reference/</a></p></li><li><p><strong>Microsoft, Designing a DDD-Oriented Microservice</strong><br>Covers the separation between the application layer and the domain layer, including the principle that the application layer coordinates while business rules remain in the domain.<br><a href="https://learn.microsoft.com/en-us/dotnet/architecture/microservices/microservice-ddd-cqrs-patterns/ddd-oriented-microservice">https://learn.microsoft.com/en-us/dotnet/architecture/microservices/microservice-ddd-cqrs-patterns/ddd-oriented-microservice</a></p></li><li><p><strong>Microsoft, Designing Validations in the Domain Model Layer</strong><br>A direct reference for domain invariants and why aggregates and domain entities should enforce them.<br><a href="https://learn.microsoft.com/en-us/dotnet/architecture/microservices/microservice-ddd-cqrs-patterns/domain-model-layer-validations">https://learn.microsoft.com/en-us/dotnet/architecture/microservices/microservice-ddd-cqrs-patterns/domain-model-layer-validations</a></p></li><li><p><strong>Martin Fowler, Anemic Domain Model</strong><br>A useful background reference on what happens when business logic leaks out of the domain model and becomes scattered across surrounding services and application code.<br><a href="https://martinfowler.com/bliki/AnemicDomainModel.html">https://martinfowler.com/bliki/AnemicDomainModel.html</a></p></li><li><p><strong>OWASP, LLM06:2025 Excessive Agency</strong><br>Relevant to agent autonomy, tool permissions, excessive authority, and the need for additional controls around high-impact actions.<br><a href="https://genai.owasp.org/llmrisk/llm062025-excessive-agency/">https://genai.owasp.org/llmrisk/llm062025-excessive-agency/</a></p></li><li><p><strong>OWASP, LLM08:2025 Vector and Embedding Weaknesses</strong><br>Relevant to retrieval systems, access control, cross-context leakage, stale or poisoned knowledge, and the risks introduced when retrieved data influences downstream decisions.<br><a href="https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/">https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/</a></p></li><li><p><strong>Anthropic, Building Effective Agents</strong><br>A practical reference on agent design, tool use, workflows, autonomy, guardrails, and the trade-offs involved in building production agent systems.<br><a href="https://www.anthropic.com/engineering/building-effective-agents">https://www.anthropic.com/engineering/building-effective-agents</a></p></li></ol><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Modern Software Architecture (Part 2): Coupling Is Where Architecture Actually Happens]]></title><description><![CDATA[Understanding Coupling in Modern Software Architecture in the Era of AI Systems: Why AI Makes Hidden System Dependencies More Visible and Harder to Control]]></description><link>https://nidly.substack.com/p/modern-software-architecture-part</link><guid isPermaLink="false">https://nidly.substack.com/p/modern-software-architecture-part</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 10 Aug 2026 06:19:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LIcZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LIcZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LIcZ!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!LIcZ!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!LIcZ!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!LIcZ!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LIcZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg" width="1024" height="682" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/faebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:682,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:143105,&quot;alt&quot;:&quot;A modern office building detail&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A modern office building detail" title="A modern office building detail" srcset="/__u/substackcdn.com/image/fetch/$s_!LIcZ!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!LIcZ!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!LIcZ!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!LIcZ!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffaebaa61-a7b6-463c-b11a-a50fa292f1e5_1024x682.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There is a question that shows up in every architecture review, every platform retrospective, every technical offsite:<strong> </strong><em><strong>how do we decouple these systems?</strong></em></p><p><br>It gets written on whiteboards, turned into roadmap items, and sometimes escalated into full-scale migrations. <strong>It is also the wrong question.</strong></p><p>Not because decoupling is a bad instinct. The instinct comes from somewhere real, from cascading deployments, from debugging a payment outage that somehow took down notifications, from inheriting systems where a single database change requires coordination across multiple teams. The instinct is valid. The framing is not.</p><blockquote><p>The real question is:<strong> </strong><em><strong>where does coupling actually live in this system, and what does it cost when it fails?</strong></em></p></blockquote><p><br>That shift changes the entire framing of architecture. It moves it away from ideals like &#8220;modularity&#8221; or &#8220;clean separation&#8221; and into something closer to engineering reality: <em>managing failure modes and coordination costs</em>.</p><p>Coupling is not a property you eliminate or design out of a system. It is the cost of coordination embedded in every non-trivial architecture. <strong><a href="/__u/nidly.substack.com/p/architecting-software-for-the-ai?r=a3p8i">Architecture</a></strong> is the act of deciding where that cost sits and how painful it becomes when it surfaces.</p><div><hr></div><h2>Coupling Is Inevitable</h2><p>Before anything else, it is worth accepting a simple constraint: </p><blockquote><p>coupling exists in every system.</p></blockquote><p>The question is never whether it exists. The question is where it accumulates, how it manifests, and who ends up paying for it when things <em><strong>go wrong</strong></em>.</p><p>In a monolith, coupling is structural and local. A shared database schema, a service layer that slowly accumulates responsibilities, internal modules that started clean but gradually absorbed unrelated concerns. The coupling is visible. Sometimes uncomfortably so. But it is still traceable. When something breaks, there is usually a stack trace pointing somewhere real.</p><p>In microservices, coupling does not disappear. It shifts outward. It moves into the network, into deployment pipelines, into API contracts, and into the shared assumptions engineers carry about how services behave. It becomes less visible and more expensive to debug when it fails.</p><p>In event-driven systems, coupling shifts again. Temporal coupling becomes dominant due to  the dependency on ordering, timing, and eventual consistency. Producers and consumers drift. Events evolve. Meaning changes depending on when and where you observe the system.</p><p>In AI-integrated systems, coupling takes on new forms entirely. Behavior becomes coupled to prompts, to context windows, to retrieval pipelines, and to model versions that can change without explicit coordination. None of these architectures eliminates coupling. They simply produce different coupling profiles. And most real-world failures come from misunderstanding which profile you are operating under.</p><div><hr></div><h2>The Taxonomy Worth Having</h2><p>Not all coupling is the same, and treating it as a single concept leads to bad architectural decisions. In production systems, coupling shows up in multiple distinct forms, each with different failure modes and different costs. At least five of them matter consistently in real systems.</p><p><strong>Structural coupling</strong> is the most familiar. It is the import chain, the shared library, the class that knows too much about another class. In a TypeScript codebase, it often appears as deep type dependencies where a <code>UserDTO</code> leaks into unrelated domains like payments or notifications simply because it was convenient. Structural coupling is relatively easy to detect and usually easy to fix. The problem is not its severity but its visibility. Teams tend to spend disproportionate effort reducing structural coupling while ignoring more expensive forms that accumulate elsewhere.</p><p><strong>Data coupling</strong> is where systems start to bind together in ways that are harder to unwind. When two services share a PostgreSQL schema, they are no longer independently deployable, regardless of how clean their code boundaries appear. They share a lifecycle, whether they intend to or not. The same issue appears when clients directly consume database-shaped JSON payloads, effectively coupling frontend release cycles to backend schema evolution. In fast-growing systems, this typically surfaces later as a painful realization that &#8220;<em>independent services</em>&#8221;<strong> were never actually independent, only logically separated</strong>.</p><p><strong>Temporal coupling</strong> is the hidden cost of asynchronous design. Event-driven systems often claim to remove dependencies by replacing synchronous calls with message queues. What actually changes is the nature of the dependency. Consider a checkout flow emitting an <code>OrderPlaced</code> event. Inventory must process the event before fulfillment begins, otherwise the system risks shipping items that are not reserved. The services are no longer coupled in execution time, but they are tightly coupled in ordering. These failures are particularly difficult to reproduce because they depend on timing conditions, load patterns, and message ordering that rarely appear consistently in test environments.</p><p><strong>Runtime coupling</strong> is what most engineers refer to when they say dependency. Service <strong>A</strong> calls Service <strong>B</strong>, which means <strong>A</strong> inherits <strong>B</strong>&#8217;s latency, availability, and failure behavior. At a small scale, this is manageable. At production scale, it becomes a compounding system effect. A p99 latency of 80 milliseconds per service becomes significant when multiplied across multiple sequential calls in a request path. Four such calls already push you beyond 300 milliseconds before accounting for network variance. </p><blockquote><p>Runtime coupling is where distributed systems stop being abstract architecture and start becoming performance engineering under constraints.</p></blockquote><p><strong>Cognitive coupling</strong> is the least visible but often the most expensive over time. It exists in the shared human understanding required to operate a system correctly. It is the engineer who remembers why a payment service contains an arbitrary delay. It is the team that understands undocumented assumptions between billing and subscription flows. When those individuals leave or teams are reorganized, the system does not immediately break. Instead, understanding decays, and what remains is a system that only functions correctly under tribal knowledge. Cognitive coupling does not show up in diagrams, but it dominates long-term maintainability.</p><div><hr></div><h2></h2><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The Coupling Relocation Principle</h2><p>Every major architectural change follows a consistent pattern. <strong>Coupling is not eliminated. It is relocated</strong>.</p><p>When a system moves from a monolith to microservices, structural coupling inside the codebase decreases. In exchange, runtime coupling increases across network boundaries, temporal coupling increases through distributed coordination, and cognitive coupling increases across teams that must now maintain shared understanding without shared code. Whether this is a good trade depends entirely on whether the original structural coupling was the dominant source of failure.</p><p>When synchronous APIs are replaced with asynchronous messaging, runtime coupling decreases because services no longer block on each other. However, temporal coupling increases significantly, along with new dependencies on event schemas, consumer offsets, and replay semantics. A Kafka-based system is not less coupled. It is coupled through different mechanisms that are harder to observe and debug.</p><p>When a shared database is replaced with service-owned data and event-based synchronization, direct data coupling is removed. In its place, eventual consistency introduces temporal coupling, since the system may exist in intermediate states between event emission and full propagation. Projection systems that rebuild read models introduce another layer of coupling, this time to event schema stability. Any change in event structure now requires careful coordination across all consumers. Across all of these cases, the pattern is consistent. </p><blockquote><p><strong>Architecture is not a process of reducing coupling. It is a process of deciding which coupling forms are acceptable and which ones are not.</strong> </p></blockquote><p>The maturity of an architecture is defined less by how much coupling it removes and more by how explicitly it manages the coupling it creates.</p><div><hr></div><h2>Why Distributed Systems Feel Hard</h2><p>Distributed systems are often described as complex, but that description is incomplete. The real difficulty is not complexity itself, but <em>the way coupling behaves under distribution</em>.</p><p>In a single process, coupling failures are immediate and local. A function call fails, an exception is thrown, and the stack trace identifies the problem. In distributed systems, the network becomes an additional coupling layer with fundamentally different behavior. It is unreliable, it introduces latency, it can reorder messages, it can duplicate requests, and it can fail partially without clear signals.</p><p>This leads to ambiguity in failure states. A timeout does not tell you whether a service is down, slow, partially degraded, or still processing your request. It only tells you that the response was not received within the expected time window.</p><p><strong>Partial failure is where coupling becomes most expensive</strong>. A system may successfully charge a user but fail to confirm a reservation in another service. From one subsystem&#8217;s perspective, the operation succeeded. From another, it never happened. The system is simultaneously correct and incorrect depending on which boundary you observe. Resolving this requires compensating mechanisms, idempotency strategies, or distributed transaction patterns, each of which introduces additional operational complexity.</p><p>Observability becomes a second-order coupling problem. Debugging requires correlating behavior across multiple independent systems, each producing logs, metrics, and traces that must be stitched together with consistent identifiers. Without this, the system becomes effectively non-deterministic during failure states. Most teams do not fully realize this until their first serious production incident spans multiple services, at which point debugging becomes an exercise in reconstructing distributed state from partial evidence.</p><div><hr></div><h2>AI Systems Introduce New Coupling Dimensions</h2><p>AI-integrated systems do not simply increase complexity. They introduce coupling types that have no direct equivalent in traditional software.</p><p><strong>Prompt coupling is the most immediate</strong>. When a system sends a prompt to a language model and parses the response, its behavior becomes coupled to the exact wording of that prompt. Small changes in phrasing can produce materially different outputs. This is not a minor implementation detail. Prompts effectively become undocumented system behavior. They are rarely versioned properly, rarely covered by behavioral regression tests, and often duplicated across codebases without clear ownership. A simple refactor intended to &#8220;improve clarity&#8221; can silently change production behavior in ways that standard code review will not catch.</p><p><strong>Context coupling</strong> is what emerges when system behavior depends on what fits inside a model&#8217;s context window at inference time. A customer support agent may behave differently depending on how much conversation history is available. Once the context exceeds a limit, earlier information is dropped, and the model&#8217;s behavior shifts even though no code has changed. This creates a system where correctness depends on a runtime condition that is difficult to reproduce, hard to test, and impossible to serialize meaningfully. In production, this often surfaces as inconsistent user reports where the same flow behaves differently under slightly different interaction lengths.</p><p><strong>Model coupling</strong> is the most underestimated form. The system becomes coupled to a specific model&#8217;s behavioral profile, including its formatting tendencies, refusal behavior, and reasoning style under ambiguity. When a model provider releases a new version, it is not a strict drop-in replacement. It is an updated behavior surface with improved capabilities but no guarantee of behavioral stability. From a systems perspective, this is an undeclared dependency upgrade that affects every inference path in the system. There is no diff to review, and rollback is not a technical operation but a vendor-level coordination problem.</p><p>This is fundamentally different from traditional dependency management. A library update comes with a diff, a version number, and a changelog. A model update often comes with a blog post. The coupling is deeper than most traditional dependencies, while the visibility is significantly lower.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/modern-software-architecture-part?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/modern-software-architecture-part?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/modern-software-architecture-part?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2>Boundaries Are Coupling Decisions</h2><p>Architectural boundaries are often described as separation points, where systems become independent. This is a misleading framing. Boundaries do not eliminate coupling. They define where coupling becomes explicit and formal.</p><p>Inside a boundary, coupling can remain informal and relatively cheap. Modules can reference each other freely. Shared state can be modified within a single transactional or deployment context. The cost of change is primarily local.</p><p>Across a boundary, coupling must be explicit. It is encoded in API contracts, event schemas, or shared data models that are treated as public interfaces rather than internal details. The cost of change is no longer local. It becomes a coordination problem across all systems that depend on that contract.</p><p>Choosing a boundary is therefore not a decision about separation. It is a decision about what kind of coupling should be formalized and what kind should remain flexible. Premature boundaries tend to fail for this reason. They formalize coupling before it is well understood, freezing relationships that still need to evolve. Instead of reducing complexity, the boundary locks it in place.</p><div><hr></div><h2>The Coupling Budget</h2><p>Every system operates under a finite coupling budget. This is not a formal metric, but a useful abstraction for reasoning about trade-offs.</p><p>Different forms of coupling consume this budget in different ways. Code-level coupling increases refactoring cost and reduces local flexibility. Runtime coupling increases latency sensitivity and amplifies cascading failures. Data coupling increases migration cost and tightens deployment coordination. Cognitive coupling increases onboarding time and reduces resilience to team changes.</p><p>The architectural question is <em><strong>not</strong></em> how to eliminate coupling, but how to allocate this budget. A small startup can afford heavier cognitive coupling because shared context lives in people&#8217;s heads. A large organization cannot, and must formalize that knowledge into contracts and systems. Similarly, a system may deliberately accept runtime coupling in exchange for simpler data ownership, or accept data coupling in exchange for faster iteration speed.</p><blockquote><p>Architecture is ultimately the management of these constraints, not their removal.</p></blockquote><div><hr></div><h2>What Happens When You Ignore This</h2><p>When coupling is not explicitly managed, it does not disappear. It accumulates in hidden and often more dangerous forms.</p><p>A distributed monolith emerges when services are split structurally but remain tightly coupled at runtime. Deployments must be coordinated across multiple services, and failures in one service propagate across the entire system. The result is microservices overhead without microservices independence.</p><p>Schema drift appears when services share implicit or loosely governed data contracts. One service evolves its data model while another continues to rely on the previous shape. The system appears stable until a subtle mismatch causes production failures that are difficult to trace back to a single change.</p><p>Temporal race conditions emerge when asynchronous systems assume ordering guarantees that do not exist. A consumer may process an event before the system state it depends on is fully committed or visible. These issues are notoriously difficult to reproduce because they depend on timing, load, and system state convergence.</p><p>In AI systems, the equivalent failure mode is silent behavioral regression. A model update or prompt adjustment introduces subtle changes in output structure or reasoning behavior. The system continues to operate without errors, but downstream consumers begin to behave incorrectly. There are no exceptions, no logs, and no obvious failure signals, only gradual drift in system behavior.</p><div><hr></div><h2>Coupling Is Not Solvable</h2><p>Good architecture does not produce decoupled systems. It produces systems where coupling is explicit, intentional, and placed where it is cheapest to manage.</p><p>Bad architecture produces accidental coupling. Coupling that accumulates in the gaps between decisions, in interfaces nobody designed deliberately, in shared resources that started as shortcuts and quietly became load-bearing parts of the system. Accidental coupling is not dangerous because it exists. It is dangerous because it remains invisible until it fails.</p><p>The work of architecture is not the elimination of coupling. It is the continuous effort to make it visible and reason about it clearly. Where does coupling exist? What are the contracts? Who owns them? What breaks when they change? How far does a failure propagate before it becomes observable? Which couplings are explicit trade-offs, and which ones were inherited without being acknowledged?</p><p>A system where every major form of coupling is known, named, and understood is always more manageable than a system that is &#8220;theoretically decoupled&#8221; but practically opaque. You can operate what you can see. You cannot operate on what hides inside undocumented schemas, unversioned prompts, or implicit timing dependencies no one ever mapped.</p><blockquote><p>Coupling is not the enemy. Invisible coupling is.</p></blockquote><div><hr></div><p><em>Part 3 will focus on contracts: how systems formalize the coupling they choose to keep, and why contract failure is often the most expensive architectural failure in production systems.</em></p><div><hr></div><h2>References</h2><ul><li><p>Eric Evans. <em>Domain-Driven Design: Tackling Complexity in the Heart of Software</em>. Addison-Wesley, 2003.</p></li><li><p>Martin Fowler. <em>Patterns of Enterprise Application Architecture</em>. Addison-Wesley, 2002.</p></li><li><p>Neal Ford, Mark Richards, Pramod Sadalage, Zhamak Dehghani. <em>Fundamentals of Software Architecture</em>. O&#8217;Reilly Media, 2020.</p></li><li><p>Sam Newman. <em>Building Microservices (2nd Edition)</em>. O&#8217;Reilly Media, 2021.</p></li><li><p>Martin Kleppmann. <em>Designing Data-Intensive Applications</em>. O&#8217;Reilly Media, 2017.</p></li><li><p>Gregor Hohpe. <em>The Software Architect Elevator</em>. O&#8217;Reilly Media, 2020.</p></li><li><p><a href="https://www.cio.com/article/4170216/why-architecture-matters-more-than-ever-in-ai-driven-software-development.html?utm_source=chatgpt.com">Why Architecture Matters More Than Ever in AI-Driven Software Development (CIO)</a>. Discusses how AI shifts software architecture from blueprint to governance and organizational control. (<a href="https://www.cio.com/article/4170216/why-architecture-matters-more-than-ever-in-ai-driven-software-development.html?utm_source=chatgpt.com">CIO</a>)</p></li><li><p><a href="https://www.techtarget.com/searchapparchitecture/tip/AIs-role-in-different-software-architecture-contexts?utm_source=chatgpt.com">AI&#8217;s Role in Different Software Architecture Contexts (TechTarget)</a>. Explores where generative AI augments architectural work and where human judgment remains essential. (<a href="https://www.techtarget.com/searchapparchitecture/tip/AIs-role-in-different-software-architecture-contexts?utm_source=chatgpt.com">TechTarget</a>)</p></li><li><p><a href="https://www.infoq.com/articles/oil-water-moment-ai-architecture?utm_source=chatgpt.com">The Oil and Water Moment in AI Architecture (InfoQ)</a>. Examines deterministic software boundaries versus probabilistic AI components in production systems. (<a href="https://www.infoq.com/articles/oil-water-moment-ai-architecture?utm_source=chatgpt.com">InfoQ</a>)</p></li><li><p><a href="https://arxiv.org/abs/2103.07950?utm_source=chatgpt.com">Software Architecture for ML-based Systems: What Exists and What Lies Ahead</a>. Research survey on architectural challenges introduced by machine learning systems.</p></li><li><p><a href="https://arxiv.org/abs/2604.04990?utm_source=chatgpt.com">Architecture Without Architects: How AI Coding Agents Shape Software Architecture</a>. Research discussing how AI coding agents implicitly make architectural decisions and why governance becomes important.</p></li></ul><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Million Dollar Mistake: How Stupid Developers Rot Your System]]></title><description><![CDATA[Why bad engineers can quietly turn your entire system into a multi-million-dollar liability.]]></description><link>https://nidly.substack.com/p/the-million-dollar-mistake-how-stupid</link><guid isPermaLink="false">https://nidly.substack.com/p/the-million-dollar-mistake-how-stupid</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 03 Aug 2026 06:19:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!d5pd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!d5pd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!d5pd!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png 424w, /__u/substackcdn.com/image/fetch/$s_!d5pd!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png 848w, /__u/substackcdn.com/image/fetch/$s_!d5pd!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png 1272w, /__u/substackcdn.com/image/fetch/$s_!d5pd!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!d5pd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png" width="1456" height="463" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:463,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:137374,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/180777418?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!d5pd!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png 424w, /__u/substackcdn.com/image/fetch/$s_!d5pd!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png 848w, /__u/substackcdn.com/image/fetch/$s_!d5pd!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png 1272w, /__u/substackcdn.com/image/fetch/$s_!d5pd!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff624342-1302-40bf-818b-458b7e1ad51c_1621x516.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A few years ago, I watched a marketing team send what should have been a routine property alert campaign at a real estate technology company. Nothing unusual about it: new listings matched against saved searches and emailed to a segment of leads, the kind of feature that exists in almost every product with a user database and something to put in front of those users.</p><p>By the end of the afternoon, the queue system had logged roughly six hundred thousand failed jobs.</p><blockquote><p>Not six hundred. Not six thousand. <strong>Six hundred thousand</strong>.</p></blockquote><p>When a number like that appears on a dashboard, the first assumption is usually that some part of the infrastructure has collapsed. Maybe a connection pool was exhausted, maybe a server ran out of memory, maybe the database or load balancer finally gave up. That was my assumption too, until I opened the failed jobs table and started reading through the exceptions.</p><p>The servers were <em>fine</em>. CPU never really moved, memory stayed flat, and the database was still standing. The actual problem was less dramatic and much more <strong>embarrassing</strong>.</p><p>The campaign job fetched every lead in the target segment in a single query, with no chunking and no pagination, then pushed the whole result into an array and handed it to the dispatcher. That query also relied on a filter column that was not properly indexed, so as the dataset grew, it became slow enough to occasionally hit the database timeout. When that happened, the parent job could fail halfway through processing the segment.</p><p><strong>Laravel&#8217;s</strong> queue worker then did exactly what it had been configured to do and retried the failed job. The problem was that nothing recorded which leads had already been queued for that campaign, so the retry did not continue from the point of failure. It started again from the beginning.</p><p>Some leads received the email once, some received it four or five times, and others never received it at all. Then there was the data itself: a null relationship somewhere in the chain, a soft-deleted account appearing because a query had missed its global scope, or an email address that had never been properly validated at signup.</p><p>Each broken send, repeated attempt, and failed retry created another failed job entry. Take a few structural mistakes that seem harmless on their own, multiply them by tens of thousands of records, then add a queue that retries automatically, and a routine campaign can suddenly produce six hundred thousand failures instead of a few thousand legitimate edge cases.</p><p>Nobody had intentionally written reckless code, and nobody had done anything obviously incompetent. The failure came from a collection of small decisions that were individually easy to justify but were made without enough thought about what would happen once the system got bigger.</p><p>For months, nothing happened. The code passed its tests, worked in staging, and survived the first dozen production runs because the segments were smaller and the data was clean enough to keep the weaknesses hidden.</p><p>Then the company grew, the lead database grew with it, and the same code that had been &#8220;working&#8221; for a year became an incident.</p><p>This is the story I keep coming back to when people ask why software becomes expensive to maintain. It usually is not the exotic failure. It is the boring decision nobody thought was worth questioning.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Why small mistakes survive long enough to become expensive</h2><p>Every engineer has shipped something that was technically wrong but functionally fine. The gap between those two states is where a lot of production risk lives, and it is wider than most teams like to admit.</p><p>Low data volume is one of the easiest ways to hide a bad decision. A query with no index might run in twelve milliseconds against three thousand rows, and nobody is going to worry about a twelve millisecond query. Nothing looks broken in the code, and nothing looks suspicious in the metrics because the system has not grown enough to expose the problem. The query is correct in behavior and wrong in design, quietly waiting for the table to become large enough that a full scan is no longer cheap.</p><p>Low traffic hides concurrency problems in much the same way. A background job with no idempotency guard looks harmless if it only runs once per trigger, which is often exactly what happens when a product has ten users and one campaign a month. Race conditions need enough overlap to become visible. Duplicate execution needs retries or concurrent requests to land close enough together. At low scale, that simply does not happen often enough to reveal the flaw. The code is not necessarily correct. It just has not been challenged yet.</p><p>Then there is the most common failure mode of all: happy-path thinking. Most code is written around the request that succeeds. The lead has a valid email, the relationship exists, the job finishes before the timeout, and retries never matter because nothing fails.</p><p>Engineers who work this way are not usually careless. They are focused on delivering the feature they were asked to build, and in the demo the feature works. What is missing is the habit of pushing on the assumptions around it. What happens if the email is null? What happens if the job runs twice? What happens if the query returns forty thousand rows instead of forty?</p><p>Those questions sound simple, but they are often the difference between code that keeps working as the system grows and code that eventually turns into an incident.</p><p>None of this looks like an obvious defect in code review because none of it is obviously broken. The code compiles, the tests pass, and the feature ships. The cost has not disappeared. It has only been deferred, and deferred engineering cost tends to compound.</p><div><hr></div><h2>Where weak judgment actually does the damage</h2><p>Once enough of these deferred decisions pile up, the damage usually appears in three places: the database, the application logic, and the process around them. It is worth separating the three because each one creates a different kind of long-term cost.</p><p>The database is usually where the consequences become visible first. Inefficient queries are the obvious symptom, but they are often downstream of weaker assumptions about the data itself. Someone assumes a relationship will always exist, so null is never handled. Someone assumes a table will stay small, so a frequently filtered column never gets an index. Someone assumes a job will only run once, so there is nothing preventing duplicate writes.</p><p>The problem is that these assumptions leave state behind. Six months later, an engineer trying to clean up duplicated records discovers there is no reliable way to tell which row is the real one because the schema never enforced uniqueness in the first place. A cleanup that should have been a simple query becomes a careful investigation through historical data, production behavior, and whatever other features have since started depending on that table. By then, deleting the wrong row is no longer a local mistake.</p><p>The second layer of damage lives in application logic. Missing idempotency shows up in incident after incident because retries are not unusual behavior in distributed systems. Queues retry. Webhook providers retry. Network clients retry after timeouts. Anything that is not safe to execute twice will eventually be executed twice.</p><p>Retry behavior can make the situation worse. A job that immediately retries the same failing dependency, with no backoff and no sensible limit, is not recovering. It is repeatedly applying pressure to something that is already unhealthy. Weak validation creates a similar chain reaction. Bad data gets accepted at one boundary, travels through several layers that assume it is valid, and finally explodes somewhere much further away. That is how a bad email address entered during signup eventually becomes an exception inside a mailer that should never have been responsible for validating it.</p><p>Underneath this is another problem that is harder to see in code: ownership. The engineer who originally wrote the job may have moved to another team. The engineer maintaining it knows how the happy path works but has never had a reason to study its failure modes. When the system finally behaves differently from the original assumptions, nobody is quite sure who owns the consequences.</p><p>Process is what allows these decisions to reach production in the first place. Weak code review does not necessarily mean nobody reviewed the diff. More often, the review checked whether the code did what the ticket asked for and stopped there.</p><p>Nobody asked what happens if the query returns ten times more rows than expected. Nobody asked how a campaign can be stopped halfway through without creating inconsistent state. Nobody asked whether duplicate execution should be prevented by a unique constraint, whether failure rate should have its own alert, or whether repeated failures should cause the job to stop instead of continuing to hammer the same dependency.</p><p>None of this is fixed by better linting or a more elaborate CI pipeline. Those tools are useful, but they mostly check whether code satisfies rules we already know how to express. They do not ask what happens when assumptions fail.</p><p>That kind of operational reasoning is a separate engineering skill, and it is usually learned much faster after someone has had to support a system they built in production.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/the-million-dollar-mistake-how-stupid?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/the-million-dollar-mistake-how-stupid?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/the-million-dollar-mistake-how-stupid?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2>The hidden cost of a cheap developer</h2><p>Somewhere in most growing companies, someone makes the decision to hire junior or lower-cost engineers to move faster on a limited budget. That is not a bad decision by itself. Junior engineers are not the problem. Everyone was junior once, and plenty of senior engineers write exactly the kind of code described above. Seniority and judgment are related, but they are not the same thing.</p><p>The real problem starts when hiring and oversight treat &#8220;can write working code&#8221; as enough, without asking whether the person can reason about what happens when the assumptions underneath that code stop holding. That gap rarely shows up on day one. It shows up much later, sometimes in a queue table with six hundred thousand failed jobs in it.</p><p>Salary is the easiest cost to measure, so it is often the one companies optimize around. But it is only one part of the actual cost.</p><p>There is the incident itself and the hours other engineers spend diagnosing a problem that should never have existed. There is the senior engineer pulled away from their own roadmap to fix it, which means something else the company wanted to ship now gets delayed. There is the slowdown that follows, because after an incident like this, everyone becomes more cautious around the affected system. More manual checks, more verification, more hesitation before touching code nobody fully trusts.</p><p>Then there is customer impact. Six hundred thousand failed jobs are not just rows in a database. They can become duplicate emails in real inboxes, or leads that never receive an alert at all during the exact window in which that alert mattered.</p><p>And some of the technical damage usually remains. The inconsistent data, the missing index, the job with no idempotency guard. There is always another priority, so parts of the original problem survive the incident and stay in production until the next increase in traffic or data volume exposes them again.</p><p>So the useful question is not whether a cheaper developer costs less on the salary line. It is whether the company is genuinely saving money, or simply pushing part of the cost into the future.</p><p>Often, the expensive part arrives later, when the system has enough scale, enough data, and enough concurrent load to expose assumptions that had been wrong for months or years. By then, fixing the original decision is much harder because there is production data to preserve, other systems depending on it, and sometimes customers affected by the cleanup.</p><div><hr></div><h2>Production access is a trust boundary, not a job title</h2><p>There is a framing I have come to rely on when thinking about who should be allowed to touch core systems: production access is not just a technical permission. It is a trust boundary.</p><p>Giving someone write access to a production system is not simply a statement that they can produce working code. It means the company trusts their judgment about what happens when that code interacts with real data, real scale, and failure conditions nobody reproduced in a test environment.</p><p>The obvious risk is reliability. Anyone with access to a core system can introduce an outage. But the bigger risk is often data integrity, because those failures can be much harder to detect and much harder to undo.</p><p>Reliability problems usually announce themselves. A service goes down, latency spikes, alerts fire, customers complain. Data integrity problems can be much quieter. A job runs twice and creates inconsistent records. A migration removes a constraint nobody realized another service depended on. A field gets populated incorrectly for months before anyone notices.</p><p>By then, the problem may no longer belong to one table. Reports have been generated from that data. Machine learning features may depend on it. Billing calculations may have used it. Other services may have copied or transformed it. Fixing the original record does not automatically undo every decision that was made while the data was wrong. There is another effect that gets discussed less: code that ships becomes precedent.</p><p>The next engineer who works in that part of the codebase reads what is already there and reasonably assumes it represents an accepted pattern. An unindexed query becomes an example for the next query. A job without idempotency becomes the reference for the next background task. Weak decisions do not always stay isolated to the feature where they were introduced. They get copied.</p><p>That is why the important question when giving someone access to core systems is not only, &#8220;Can this person write code that works?&#8221;</p><p>It is also, &#8220;Are this person&#8217;s decisions safe under scale, failure, and change over time?&#8221;</p><p>Those are different questions. Teams that only evaluate the first one eventually discover the second one in production.</p><div><hr></div><h2>What a better filter actually looks like</h2><p>None of this is an argument for hiring only senior engineers, or for expecting perfect foresight from anyone before they are allowed to ship code. That standard does not exist. Pretending it does usually creates either paralysis or false confidence in people who interview well but have never had their judgment tested in production.</p><p>The better approach is to evaluate judgment earlier, while mistakes are still cheap to catch: during hiring, design discussions, and code review, before weak assumptions become production state.</p><p>Most technical interviews still focus heavily on whether a candidate can produce the correct output. Can they solve the algorithm? Can they draw the expected architecture? Can they name the right database or queue? Those things matter, but they do not tell you much about how the person thinks when the environment stops behaving nicely.</p><p>A more useful discussion starts with consequences. Show someone a query and ask what changes when the table grows from ten thousand rows to ten million. Show them a background job and ask what happens if it runs twice. Give them an endpoint that writes to several systems and ask what happens when the third dependency fails after the first two succeeded.</p><p>Ask about an incident they were involved in, not for the drama of the story, but to see whether they understand the causal chain. Can they explain why the incident happened rather than only describe the visible symptom? Did it change anything about how they design similar systems now?</p><p>Trade-offs are another useful signal. Ask what they gave up in a recent technical decision. Engineers with good judgment can usually identify the downside of the option they chose. If every decision is described as an obvious win with no meaningful cost, either the problem was trivial or the trade-off was never really understood.</p><p>None of these questions depend on memorizing a particular architecture pattern. They test whether someone understands that code is not just a set of instructions. It also contains assumptions about traffic, data quality, execution order, failure behavior, and the systems around it.</p><p>Those assumptions will eventually be tested, whether the engineer planned for it or not.</p><div><hr></div><h2>What survives, and what doesn&#8217;t</h2><p>Good engineers make mistakes constantly. I have made every mistake described in this piece at some point, usually earlier in my career and occasionally later than I would like to admit. That is not the dividing line. The important difference is what happens after the mistake becomes visible.</p><p>An engineer ships an unindexed query, watches it become slow as the table grows, understands why it happened, and thinks differently about the next query. A job gets executed twice and creates duplicate state, so the next background workflow is designed with idempotency in mind from the beginning.</p><p>The safeguard stops being something copied from a checklist and becomes part of how the engineer reasons about the system.</p><p>A system built by people who work this way can survive a lot. Bugs still happen, but they are more likely to be caught before they spread. Incidents still happen, but people are better equipped to trace the failure and fix the underlying condition instead of only patching the visible symptom.</p><p>What becomes dangerous is the accumulation of decisions made without that learning loop.</p><p>One weak assumption is usually survivable. So is one missing index, one unsafe retry, one piece of bad validation. The problem is what happens when hundreds of those decisions accumulate across a system over several years.</p><p>That kind of damage rarely appears in a sprint review. It may not show up in a performance evaluation either. For a long time, everything still appears to work.</p><p>Then one afternoon, something completely ordinary happens. A marketing team sends an email campaign, and a failed jobs table fills with six hundred thousand rows.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Mesh Is Domain-Driven Design for Analytical Systems]]></title><description><![CDATA[Why ownership, not technology, is the real innovation behind modern analytical architectures.]]></description><link>https://nidly.substack.com/p/data-mesh-is-domain-driven-design</link><guid isPermaLink="false">https://nidly.substack.com/p/data-mesh-is-domain-driven-design</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 27 Jul 2026 09:03:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YmY2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YmY2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YmY2!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!YmY2!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!YmY2!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YmY2!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YmY2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2613940,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/204279110?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!YmY2!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!YmY2!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!YmY2!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YmY2!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd7c2a96-f32e-45aa-8c60-4b4644be27ff_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Software architects often ignore <a href="https://martinfowler.com/articles/data-mesh-principles.html">Data Mesh</a>. It usually appears alongside discussions about Snowflake, Iceberg, Kafka, and Databricks, so it&#8217;s easy to dismiss it as something for the data platform team rather than something worth understanding as a software architect.</p><p>Data engineers often make the same mistake in the opposite direction. <a href="/__u/nidly.substack.com/p/domain-driven-design-in-the-ai-era?r=a3p8i">Domain-Driven Design</a> can feel like a discipline for application developers, full of aggregates, entities, and transactional boundaries rather than pipelines, warehouses, and analytical infrastructure. Both communities have been solving remarkably similar problems for years. They have simply used different language to describe them.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Data Mesh becomes much easier to understand once you stop thinking of it as a data platform architecture and start seeing it as the analytical continuation of an idea Domain-Driven Design introduced for operational systems. This isn&#8217;t a historical claim that Data Mesh borrowed domain ownership from DDD. The two communities reached similar conclusions independently because they were responding to the same pressure: once an organization grows beyond a certain size, no single team can understand every part of the business well enough to own it centrally.</p><p>The <a href="/__u/nidly.substack.com/p/why-software-architecture-matters?r=a3p8i">architectural</a> similarities are striking enough that someone fluent in one can often understand the other with surprisingly little translation. That is the lens this article uses, and it isn&#8217;t a tutorial on how to build a Data Mesh. It&#8217;s a different way of understanding why it exists.</p><div><hr></div><h2>Operational Systems Already Learned This Lesson</h2><p>You already know this part, so I&#8217;ll keep it brief. Domain-Driven Design&#8217;s most enduring contribution wasn&#8217;t Entities, Value Objects, Repositories, or any of the tactical patterns that usually dominate conference talks. Its lasting contribution was recognizing that a business model is only coherent within a boundary, and that those boundaries should reflect how an organization actually divides responsibility.</p><p><strong>A bounded context isn&#8217;t simply a technical partition</strong>. It&#8217;s an acknowledgment that the word <em>customer</em> means something different to Billing than it does to Support, and pretending otherwise usually produces a model that satisfies neither.</p><p>The architectural consequence was ownership. Once you draw a bounded context around Orders, the team behind it owns the Orders model, the Orders database, and the Orders API. Other teams aren&#8217;t expected to reach directly into that database. They interact through published contracts or consume domain events instead.</p><p>That decision wasn&#8217;t free. It introduced duplicated concepts, coordination overhead, and more work at the boundaries between domains. But most organizations eventually concluded that the tradeoff was worth it. Teams could evolve independently without negotiating every internal change across the entire organization.</p><p>Whether we called it Domain-Driven Design, microservices, or simply &#8220;the new architecture,&#8221; operational systems gradually moved toward <strong>domain ownership</strong>.</p><div><hr></div><h2>Analytics Never Fully Followed</h2><p>Here&#8217;s where the asymmetry appears. While operational systems spent the last decade decentralizing, analytical systems often continued moving in the opposite direction.</p><p>A typical organization now has dozens of bounded contexts on the operational side, each with its own team, model, and database. Yet analytical data still tends to flow into one centralized warehouse, managed by one centralized platform or data team. This wasn&#8217;t a bad decision.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!OBVA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!OBVA!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!OBVA!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!OBVA!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OBVA!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!OBVA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2522064,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/204279110?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!OBVA!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!OBVA!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!OBVA!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!OBVA!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cad6409-f62f-4bbc-850a-4ca461f1b981_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>When analytical workloads were relatively small, centralization was entirely reasonable. A single team could understand enough of the business to build shared reports, dashboards, and pipelines. Creating dedicated analytical ownership for every domain would have introduced more complexity than value.</p><p>The problem is that many organizations never revisit that decision after they outgrow it. Over time, the warehouse becomes coupled to implementation details it was never meant to understand. A column is renamed inside the Orders service for perfectly valid operational reasons. Somewhere downstream, a pipeline quietly breaks. Nothing about the business changed. Only an internal schema changed. Yet analytical consumers now pay the price.</p><p>Meanwhile, the central data team gradually becomes responsible for translating business concepts across every domain in the company, despite owning none of them. They become the de facto experts on Orders, Billing, Inventory, Marketing, and Support, simply because every analytical request eventually lands on their desk.</p><p>This is remarkably similar to the problem bounded contexts were created to solve. We solved it for operational systems. We simply never carried the same idea into analytical ones.</p><div><hr></div><h3>CQRS Accidentally Predicted Data Mesh</h3><p>This is the part of the argument worth slowing down for, because it is the piece that makes the rest of this article more than an analogy.</p><p>If you have worked on any system with real complexity, you have probably already built something that looks like a domain-owned analytical product, even if nobody called it that. CQRS, Command Query Responsibility Segregation, separates the model used for writes from the model used for reads. The write model enforces invariants and stays close to the domain&#8217;s behavior. The read model is a projection, built asynchronously, optimized for how it will be queried rather than how it was produced.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iXrw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iXrw!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!iXrw!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!iXrw!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iXrw!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iXrw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2096198,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/204279110?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iXrw!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!iXrw!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!iXrw!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iXrw!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a28bf49-5c2f-4161-a04b-b9dd1722a0b4_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That projection is structurally very close to a data product. It is owned by the same team that owns the write model, because they are the only ones who understand the domain well enough to define it correctly. It is built from domain events, which already represent the source of truth. It is decoupled from the internal write-side schema, which allows the write model to evolve without breaking downstream consumers. And it exists because a single model cannot efficiently serve both transactional and analytical access patterns.</p><p>The important point is not that the projection exists. Teams have been building read models for years, long before Data Mesh was defined. The important point is what it implies.</p><p>Data Mesh does not introduce the idea of a projection. It takes something that already existed inside system boundaries and promotes it to a first-class architectural concept across organizational boundaries. Not an implementation detail inside a service, but a product: with a contract, an owner, documentation, versioning, and a clear expectation of how other teams consume it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!MgYk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!MgYk!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png 424w, /__u/substackcdn.com/image/fetch/$s_!MgYk!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png 848w, /__u/substackcdn.com/image/fetch/$s_!MgYk!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MgYk!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!MgYk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png" width="1389" height="1132" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1132,&quot;width&quot;:1389,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1425344,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/204279110?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!MgYk!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png 424w, /__u/substackcdn.com/image/fetch/$s_!MgYk!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png 848w, /__u/substackcdn.com/image/fetch/$s_!MgYk!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png 1272w, /__u/substackcdn.com/image/fetch/$s_!MgYk!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff40e0f08-db63-42e6-8b29-abc251687ab3_1389x1132.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Once you see it this way, Data Mesh stops feeling like a separate discipline. A domain publishing a data product is doing the same thing a CQRS read model already does, but with stronger guarantees around ownership and external consumption. For an architect, this is the key shift: you are not learning a new pattern, you are extending an existing one into a layer of the system that was previously treated as shared infrastructure.</p><div><hr></div><h3>Data Mesh Is Bigger Than This Comparison</h3><p>It would be incorrect to reduce Data Mesh to CQRS at organizational scale. The model is broader and rests on four principles, not one.</p><p>Domain Ownership is the foundation: the team closest to the domain owns its analytical representation, just as it owns its operational model. Data as a Product extends that idea into practice: the output is not an internal artifact but something that must be documented, versioned, and supported like any other product.</p><p>The remaining two principles sit outside the DDD comparison. Self-Serve Data Platform is an infrastructure concern: it describes how teams publish and consume data products without each building their own pipelines and tooling. Federated Computational Governance is an organizational constraint: it ensures that independently owned data products remain interoperable, secure, and consistent enough to function as a system rather than a collection of isolated datasets.</p><p>Both are essential in practice, but they belong more to platform engineering and organizational design than to Domain-Driven Design. The focus here is deliberately on the first two principles, because that is where the conceptual bridge is strongest.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/data-mesh-is-domain-driven-design?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/data-mesh-is-domain-driven-design?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/data-mesh-is-domain-driven-design?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h3>Why AI Systems Make This More Urgent</h3><p>Dashboards were the original justification for centralized analytics, and they tolerated a fair amount of architectural imprecision. A stale report or a slightly incorrect join was frustrating, but rarely critical.</p><blockquote><p>AI-native systems remove that tolerance.</p></blockquote><p>Increasingly, analytical data is not consumed by humans reading dashboards. It is consumed by systems making automated decisions in production, often feeding back into the same domains that produced the data.</p><p>Take a Content Service. Its operational model handles content creation, editing, and publishing. But once it integrates LLM-based workflows, it also produces a second category of data that has nothing to do with content as a domain concept and everything to do with system behavior: inference latency, model version, prompt and completion tokens, cost per request, and quality or safety signals.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!P3Fs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!P3Fs!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png 424w, /__u/substackcdn.com/image/fetch/$s_!P3Fs!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png 848w, /__u/substackcdn.com/image/fetch/$s_!P3Fs!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png 1272w, /__u/substackcdn.com/image/fetch/$s_!P3Fs!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!P3Fs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png" width="1216" height="1294" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1294,&quot;width&quot;:1216,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1124123,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/204279110?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!P3Fs!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png 424w, /__u/substackcdn.com/image/fetch/$s_!P3Fs!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png 848w, /__u/substackcdn.com/image/fetch/$s_!P3Fs!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png 1272w, /__u/substackcdn.com/image/fetch/$s_!P3Fs!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F487c943f-ce48-4428-9929-5609fbbcd38f_1216x1294.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>None of these consumers should depend on the operational database of the Content Service, and in most cases they cannot safely do so. A routing system deciding which model to use next needs a stable, versioned view of cost and latency trends, not a live query over evolving internal tables. An evaluation pipeline measuring quality over time needs consistency in meaning across releases, which is exactly what a data product contract provides and a raw schema does not.</p><p>This is where analytical ownership becomes critical. The boundary between &#8220;analytics&#8221; and &#8220;production system&#8221; is no longer clean. A poorly defined analytical artifact is no longer just a wrong dashboard. It can directly influence production behavior through automation loops that close back into the system that produced the data.</p><div><hr></div><h2>The Cost of Decentralization</h2><p>None of this is free, and any article that pretends otherwise isn&#8217;t doing its job.</p><p>Decentralized analytical ownership trades centralized consistency for distributed inconsistency. If every domain defines and publishes its own metrics, you will eventually get two domains computing &#8220;active user&#8221; slightly differently. Each definition is defensible in isolation, but both become wrong the moment someone tries to compare them in a cross-domain report.</p><p>Centralized data teams existed partly to prevent exactly this, by enforcing a single modeling layer for shared metrics. You lose that enforcement when you decentralize, and you have to replace it with something else, usually a federated governance process. That process is slower to establish and easier to neglect than it looks at design time.</p><p>There is also duplicated effort. Multiple domains independently building data quality checks, versioning conventions, and publishing pipelines is a real organizational cost, even with a shared self-serve platform reducing the marginal effort. And there is an ownership burden that did not exist in the same form before: a domain team that previously only ran its service now also owns a public-facing data product, with its own SLAs and consumers. That is effectively a second system to operate, and someone has to staff it.</p><p>None of this is an argument against <a href="https://www.amazon.com/Data-Mesh-Delivering-Data-Driven-Value/dp/1492092398">Data Mesh</a>. It is an argument that architecture is the discipline of choosing which complexity you want to own, not a search for a design with no complexity at all.</p><p>Centralized analytics traded scalability for consistency. Data Mesh trades some consistency back for scalability and domain correctness. Whether that trade is worth it depends entirely on whether your organization is large enough, and federated enough, for central coordination to be the bigger bottleneck.</p><div><hr></div><h2>Why Most Data Mesh Initiatives Fail</h2><p>Given those trade-offs, it is worth being explicit about how these initiatives fail in practice, because it is almost never the technology. The most common failure mode is the platform team quietly becoming the new central bottleneck under a different name. An organization adopts Data Mesh terminology, builds a &#8220;self-serve platform team,&#8221; and within a year that team becomes the required path for publishing anything.</p><p> Other teams were never given the time, incentives, or responsibility to truly own their data products end to end. The org chart says decentralized. The dependency graph does not.</p><p>The second failure is domains publishing tables instead of products. A team is asked to &#8220;expose a data product,&#8221; interprets that as exposing a replica of its operational database, and stops there. There is no contract, no versioning, no documentation, no guarantees. Structurally, this is the same coupling problem as the centralized warehouse, just with different labels.</p><p>The third, and deeper issue, is misaligned incentives. Domain-Driven Design worked for operational systems because ownership was already intrinsic to the job. Data products do not inherit that automatically. A team can ship its feature, hit its operational KPIs, and get promoted without anyone outside the team ever evaluating whether its data product is correct, documented, or usable. </p><p>Until that gap is closed, usually by making data product quality part of how a domain is evaluated, data products remain nominal rather than real. These are organizational failures expressed as architectural outcomes. No platform tool fixes them, because the missing piece was never technical.</p><div><hr></div><h2>Ownership Was Always the Point</h2><p>Domain-Driven Design changed who owns operational systems. It argued that the team closest to the domain should own the model, the data, and the pace of change, even at the cost of duplication and coordination overhead at the boundaries.</p><p>Data Mesh asks the same question of analytical systems, under different pressure: AI systems that consume analytical data as part of production decisions, organizations too large for any central team to retain full context, and engineering cultures that already understand the cost of separating ownership from expertise.</p><p>The underlying technology matters less than it appears. Kafka or not, Snowflake or not, event sourcing or not, none of that is the architectural decision. The architectural decision is who is accountable for a piece of meaning, operational or analytical, and whether the rest of the system is designed to respect that accountability or route around it.</p><p>That is the question Domain-Driven Design asked first. Data Mesh is what happens when analytical systems finally have to answer it as well.</p><div><hr></div><h3>References:</h3><ul><li><p><a href="https://martinfowler.com/articles/data-mesh-principles.html?utm_source=substack&amp;utm_medium=article&amp;utm_campaign=nidly-data-mesh">Data Mesh Principles and Logical Architecture</a> by Zhamak Dehghani<br>The foundational article outlining the four principles of Data Mesh: domain ownership, data as a product, self-serve data platforms, and federated computational governance.</p></li><li><p><a href="https://www.oreilly.com/library/view/data-mesh/9781492092384/?utm_source=substack&amp;utm_medium=article&amp;utm_campaign=nidly-data-mesh">Data Mesh: Delivering Data-Driven Value at Scale</a> by Zhamak Dehghani<br>The complete treatment of Data Mesh as a decentralized sociotechnical approach to analytical data ownership.</p></li><li><p><a href="https://martinfowler.com/bliki/CQRS.html?utm_source=substack&amp;utm_medium=article&amp;utm_campaign=nidly-data-mesh">CQRS</a> by Martin Fowler<br>A concise explanation of separating the models used for writes and reads, including the complexity and trade-offs introduced by the pattern.</p></li><li><p><a href="https://martinfowler.com/articles/201701-event-driven.html?utm_source=substack&amp;utm_medium=article&amp;utm_campaign=nidly-data-mesh">What Do You Mean by &#8220;Event-Driven&#8221;?</a> by Martin Fowler<br>A useful clarification of event-driven architecture, event sourcing, and CQRS, concepts that are often mixed together despite describing different architectural choices.</p></li><li><p><a href="https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/?utm_source=substack&amp;utm_medium=article&amp;utm_campaign=nidly-data-mesh">Domain-Driven Design: Tackling Complexity in the Heart of Software</a> by Eric Evans<br>The foundational book behind bounded contexts, model boundaries, and the strategic design ideas discussed throughout this article.</p></li><li><p><a href="https://www.oreilly.com/library/view/learning-domain-driven-design/9781098100124/?utm_source=substack&amp;utm_medium=article&amp;utm_campaign=nidly-data-mesh">Learning Domain-Driven Design</a> by Vlad Khononov<br>A modern and accessible treatment of domain boundaries, ownership, integration patterns, and the relationship between organizational structure and software architecture.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Book Review: Fundamentals of Software Architecture (2nd Edition)]]></title><description><![CDATA[Why There Is No Perfect Architecture, Only Better Trade-Offs.]]></description><link>https://nidly.substack.com/p/book-review-fundamentals-of-software</link><guid isPermaLink="false">https://nidly.substack.com/p/book-review-fundamentals-of-software</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 20 Jul 2026 06:33:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LMol!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LMol!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LMol!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!LMol!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!LMol!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!LMol!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LMol!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg" width="1456" height="1910" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1910,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:410970,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/203274153?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!LMol!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!LMol!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!LMol!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!LMol!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0a7454c-ee47-44fd-b651-9c0db722d9c9_1951x2560.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why There Is No Perfect Architecture, Only Better Trade-Offs</h2><p>Most software architecture books try to teach certainty. They present clean diagrams, well-defined layers, and architectural principles that appear universally applicable. Read enough of them and you start to believe architecture is mostly about finding the correct answer. Choose the right pattern, apply it consistently, and the system will somehow behave.</p><blockquote><p>Real systems are rarely that cooperative.</p></blockquote><p>After spending years building distributed platforms, data-intensive systems, and, more recently, AI-native applications, I&#8217;ve become increasingly skeptical of architectural certainty. Most production failures don&#8217;t happen because engineers chose the wrong pattern. They happen because a reasonable decision was made under incomplete information, and the trade-offs behind that decision were not fully understood until much later.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>That is why <em><strong>Fundamentals of Software Architecture</strong></em> stands out. The book won&#8217;t tell you which architecture is right. That&#8217;s the point. Instead, Mark Richards and Neal Ford focus on something much more valuable: <em>how to think about architectural decisions when every option comes with costs, constraints, and unintended consequences</em>. The second edition, published in 2025, reinforces that idea at a time when architecture has become even more complex than when the first edition appeared.</p><div><hr></div><h2>What Makes This Book Different</h2><p>There is a particular type of <a href="/__u/nidly.substack.com/p/why-software-architecture-matters">architecture</a> book that treats patterns as answers. You describe a problem, match it to a diagram, and implement the corresponding structure. These books are not necessarily wrong, but they often create a subtle misconception: the belief that architecture is primarily a matter of pattern recognition rather than judgment.</p><p>Many architecture books implicitly teach the same workflow: identify the problem, match it to a pattern, implement the structure, and move on. Anyone who has spent enough time operating production systems knows reality is considerably messier than that.</p><p>Richards and Ford start from <em>a very different premise</em>. Architecture is not a catalogue you consult. <em><strong>It is a discipline of making decisions under uncertainty, with consequences that often remain hidden until months or years later</strong></em>. What I appreciate most is that the book consistently approaches architecture as an engineering problem rather than an ideological one. Instead of arguing for a particular style, it provides frameworks for evaluating alternatives. Instead of prescribing solutions, it helps readers understand why different solutions exist in the first place.</p><p>This mindset becomes particularly visible in the authors&#8217; treatment of the phrase &#8220;it depends.&#8221; In many technical discussions, &#8220;it depends&#8221; is treated as a weak answer, almost as an admission that the speaker lacks conviction. Here, it is treated as the honest answer to most architectural questions. The real challenge is not avoiding ambiguity but understanding what the decision actually depends on.</p><p>That distinction is important because mature architects are rarely the people with instant answers. More often, they are the people who understand the trade-offs well enough to explain why a particular answer makes sense in a specific context. The contrast with more dogmatic architecture literature is striking. Some books promote a particular style and spend hundreds of pages demonstrating its superiority. They can be useful, but they sometimes leave readers with the impression that architecture is a competition between named approaches.</p><p>Richards and Ford take the opposite approach. They seem far more interested in helping readers understand why a decision was made than what that decision was called. That perspective ages far better than any individual pattern or architectural trend.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>The Most Valuable Idea: Architecture Characteristics</h2><p>If I had to identify a single concept from the book that delivers the most practical value, it would be <a href="/__u/nidly.substack.com/p/architecting-software-for-the-ai">architectural</a> characteristics.</p><p>The idea itself is not new. Other disciplines refer to them as quality attributes or non-functional requirements. What makes this book valuable is how clearly the concept is presented and how effectively it is connected to real architectural decisions.</p><p>One of the most useful observations in the entire book is that architecture is ultimately driven by what a system must be good at. Not what it does, but how well it does it. A system can satisfy every functional requirement and still fail because it cannot scale, cannot recover from outages, takes hours to deploy, or becomes impossible to modify safely.</p><p>Scalability, reliability, security, performance, deployability, maintainability, and testability are often treated as secondary concerns. The book correctly argues that they are architecture. More importantly, it demonstrates that these characteristics frequently conflict with one another.</p><p>Consider a payment processing platform. The business wants strong consistency, high availability, low latency, and independent service deployments. Every one of those goals is reasonable. The problem is that they pull the system in different directions. Strong consistency introduces coordination costs. High availability during failures often requires accepting stale data somewhere in the system. Independent deployments increase operational flexibility but also increase the complexity of maintaining consistency across service boundaries.</p><p>The question is never whether these characteristics matter. The question is which ones matter more in a particular context.</p><p>That distinction sounds simple, yet many architecture failures originate because teams never have this conversation explicitly. I&#8217;ve seen teams adopt microservices primarily to improve deployment speed and organizational autonomy. Six months later they discovered that eventual consistency had introduced an entirely new class of production bugs. The architecture wasn&#8217;t failing because microservices were wrong. It was failing because the trade-offs were not fully understood before the migration began.</p><p>I&#8217;ve also seen the opposite. Teams stayed on a monolith because it was operationally simpler, only to discover years later that deployment pipelines had become their largest bottleneck. Build times increased, testing slowed down, and the organization gradually lost the ability to deliver changes quickly. Neither decision was inherently incorrect. Both decisions optimized for specific characteristics while creating costs elsewhere.</p><p>The strength of the book is that it gives architects a vocabulary for discussing these trade-offs before they become production incidents. It also highlights the relationships between characteristics themselves. Improving testability often improves maintainability because modular systems are easier to reason about. Improving security can reduce usability. Improving flexibility can increase complexity. Improving performance can reduce developer productivity.</p><p>These are not academic tensions. They appear in production systems every day, and having a structured way to think about them fundamentally changes how architecture discussions unfold.</p><div><hr></div><h2>The Myth of the Perfect Architecture</h2><p>Few debates in software engineering are as persistent as monolith versus microservices, event-driven versus request-driven systems, or SQL versus NoSQL databases. What makes these debates frustrating is that they are often conducted as though context does not exist.</p><p>The book serves as an effective antidote to that mindset. Rather than treating architectural styles as competing ideologies, Richards and Ford treat them as collections of trade-offs. The discussion focuses less on which approach wins and more on what each approach optimizes for. That sounds obvious, but it is surprisingly rare.</p><p>The software industry has a long history of turning architectural decisions into movements. An approach becomes successful at a large company, conference talks appear, blog posts multiply, and eventually, organizations begin copying the solution without fully understanding the problem that motivated it.</p><p>The result is predictable. A company with forty engineers adopts an architecture designed for a company with four thousand engineers. Complexity arrives immediately. The benefits arrive years later, if they arrive at all.</p><blockquote><p>Netflix&#8217;s architecture is correct for Netflix. That does not mean it is correct for everyone else.</p></blockquote><p>What often gets overlooked is that architecture is not just technology. It is also an organizational capability. Large-scale distributed systems require platform engineering, observability tooling, operational discipline, incident management processes, and teams capable of maintaining them. Those capabilities cannot be imported from a conference presentation.</p><p>I&#8217;ve seen organizations introduce distributed architectures long before they had the operational maturity required to support them. The resulting systems were objectively more complex while delivering very little business value. The technology worked. The organization wasn&#8217;t ready for it.</p><p>The book addresses this problem directly, and that honesty is one of its strengths. Its treatment of event-driven architectures is particularly balanced. The benefits are real: reduced temporal coupling, asynchronous workflows, and the ability to add new consumers without modifying existing producers. But the failure model changes dramatically.</p><p>A request-driven system can fail in visible ways. An API call times out. A service returns an error. An event-driven system often fails more subtly. An event arrives out of order. A consumer processes the same message twice. A downstream service falls behind. A queue grows silently for hours before anyone notices.</p><p>These systems solve important problems, but they also introduce entirely new categories of operational complexity. The book explains this reality without taking sides, and that restraint is one of the reasons it remains relevant.</p><div><hr></div><h2>Coupling Is the Real Enemy</h2><p>If you asked me to summarize the most durable idea in this book in a single sentence, it would be this: architectural style is less important than where you put your coupling.</p><p>The book&#8217;s discussion of <a href="/__u/open.substack.com/pub/nidly/p/modern-software-architecture-part?r=a3p8i&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">coupling</a> is the section I&#8217;ve returned to most often. The authors distinguish between static coupling, dynamic coupling, and temporal coupling. While the terminology is useful, the deeper lesson is that architecture is ultimately about managing dependencies between parts of a system, regardless of what architectural style you choose.</p><p><strong>Static coupling</strong> is what most engineers think about first. Service A depends on Service B&#8217;s interface. If B changes, A breaks. This is the kind of coupling that dependency inversion, interface-driven design, and service boundaries attempt to reduce.</p><p><strong>Dynamic coupling</strong> is often more dangerous because it only becomes visible at runtime. A checkout flow that synchronously calls inventory, payment, and shipping services may look nicely decoupled on an architecture diagram, yet the entire transaction still depends on all three systems being healthy at the same moment. The coupling hasn&#8217;t disappeared. It has simply moved from the codebase to runtime behavior.</p><p><strong>Temporal coupling</strong> is where distributed systems become particularly interesting. Two systems are temporally coupled when one must be available precisely when the other needs it. Synchronous APIs create this dependency by default. Messaging systems reduce it by allowing producers and consumers to operate independently, but they introduce different concerns around latency, ordering, retries, and duplicate processing.</p><p>The practical implication is that many modern architectures relocate coupling rather than eliminate it. I&#8217;ve seen teams celebrate a microservices migration because compile-time dependencies disappeared, only to discover months later that they had replaced them with deployment dependencies, network dependencies, and operational dependencies that were far harder to reason about.</p><blockquote><p>The coupling was always there. The real question was whether its new location made the system easier or harder to operate.</p></blockquote><div><hr></div><h2>Architecture Decisions as an Engineering Discipline</h2><p>One of the most underappreciated sections of the book focuses on Architecture Decision Records (ADRs). This is an area where the gap between what teams should do and what they actually do remains surprisingly large.</p><p>Every architecture is shaped by decisions. Some were made to address constraints that no longer exist. Some were based on assumptions that later proved incorrect. Some were made by engineers who left years ago. Yet in many organizations, the reasoning behind those decisions is never documented.</p><p>The consequences are predictable. A new engineer removes something that appears unnecessary and unintentionally breaks a critical workflow. A team inherits a system and spends months trying to understand why it was designed the way it was. Architecture reviews become debates about personal preferences because nobody remembers the trade-offs that drove the original decisions.</p><p>ADRs provide a simple solution. Document the decision, the context, the alternatives that were considered, and the expected consequences. The format itself is not important. What matters is preserving the reasoning.</p><p>The value of an ADR is not that it tells future engineers what to do. It tells them why something was done in the first place. That distinction becomes increasingly valuable as systems grow and institutional memory fades.</p><div><hr></div><h2>Modern Architecture Topics Covered in the Book</h2><p>The second edition expands into territory that feels genuinely modern. Distributed systems, event-driven architecture, domain boundaries, team topologies, evolutionary architecture, and fitness functions all receive meaningful treatment.</p><p>The discussion around Team Topologies is particularly strong because it acknowledges something many architecture books ignore: organizations shape architecture just as much as technology does. Conway&#8217;s Law is not simply an observation. It is an architectural force.</p><p>I&#8217;ve seen systems whose service boundaries reflected reporting structures more accurately than business domains. On paper the architecture looked clean. In practice every meaningful feature required coordination across multiple teams because the boundaries had been drawn around organizational convenience rather than domain ownership.</p><p>The book&#8217;s treatment of fitness functions is equally valuable. The idea is straightforward: if an architectural characteristic matters, it should be measurable. If it is measurable, that measurement should become part of the delivery process. Architectural quality should not depend entirely on manual reviews and institutional memory.</p><p>This is easy to agree with in theory and surprisingly difficult to implement in practice. The authors acknowledge that reality rather than presenting fitness functions as a silver bullet.</p><div><hr></div><h2>Where I Disagree</h2><p>My primary criticism is less about what the book gets wrong and more about what it does not yet cover.</p><p>The second edition arrives at a moment when AI-native systems are becoming part of mainstream software architecture. Large Language Models, Retrieval-Augmented Generation, vector databases, evaluation pipelines, and agentic workflows introduce architectural challenges that traditional software systems rarely encounter.</p><p>In conventional systems, correctness is often deterministic. Given the same input, the system should produce the same output. Modern AI systems operate under different assumptions. Outputs become probabilistic. Evaluation becomes a first-class architectural concern. Observability extends beyond latency and error rates into relevance, quality, and model behavior.</p><p>Similarly, agentic systems introduce runtime decision-making that makes dependency analysis more complex. The control flow is no longer fully defined at design time. Parts of the execution path emerge dynamically during inference.</p><p>None of this is a criticism of the authors. The book was never intended to be an AI architecture handbook. In fact, its core principles remain highly relevant. Trade-offs, coupling, architecture characteristics, and evolutionary thinking apply just as much to AI systems as they do to traditional software.</p><p>Still, a future edition that seriously explores AI-native architectures would be a welcome addition.</p><div><hr></div><h2>Who Should Read This Book</h2><p>Senior engineers moving toward Staff, Principal, or Architect roles will likely get the most value from this book. If your responsibilities increasingly involve evaluating trade-offs rather than simply implementing solutions, the frameworks presented here become immediately useful.</p><p>Experienced architects will find many familiar ideas, but the value lies in the clarity of the explanations. The discussions around coupling and architecture characteristics are particularly strong because they provide language for concepts that many experienced engineers understand intuitively but rarely articulate explicitly.</p><p>Engineering managers will also benefit, especially from the sections covering decision records, team structures, and architectural governance.</p><p>Junior engineers may find parts of the book more abstract. Much of its value comes from connecting the ideas to problems you&#8217;ve already encountered in production. Without that experience, some concepts can feel theoretical even though they are deeply practical.</p><div><hr></div><h2>Final Verdict</h2><p><em>Fundamentals of Software Architecture (2nd Edition)</em> remains one of the most valuable architecture books available today because it avoids the trap that many architecture books fall into: pretending there are universal answers.</p><p>Richards and Ford do not give readers a collection of patterns to memorize. They provide a framework for thinking. That distinction is what makes the book durable.</p><p>Its limitations are real. The absence of serious discussion around AI-native systems is becoming increasingly noticeable, and no book can fully replace the experience of watching architectural decisions unfold over years of production use.</p><p>Yet the book&#8217;s central idea remains as relevant today as it was when the first edition was published: architecture is the discipline of making informed trade-offs under uncertainty.</p><p>Every system is the result of decisions made with incomplete information. Some of those decisions age well. Some do not. The goal of architecture is not to eliminate uncertainty or discover a perfect design. It is to make better decisions than we would have made otherwise.</p><p>The perfect architecture does not exist. There are only architectures that serve their context well and architectures that do not. Understanding the difference is what separates architectural knowledge from architectural judgment.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Domain Driven Design Is Not an Architecture]]></title><description><![CDATA[DDD vs Software Architecture: one defines meaning, the other defines structure. Here's why the difference matters.]]></description><link>https://nidly.substack.com/p/domain-driven-design-is-not-an-architecture</link><guid isPermaLink="false">https://nidly.substack.com/p/domain-driven-design-is-not-an-architecture</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 13 Jul 2026 06:30:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!usMN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!usMN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!usMN!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!usMN!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!usMN!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!usMN!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!usMN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2600034,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/203601354?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!usMN!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!usMN!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!usMN!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!usMN!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e12e14e-28da-4678-8a7f-14b4df60312d_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Every few weeks, I come across the same discussion on LinkedIn. Someone asks, <em>&#8220;What architecture are you using?&#8221;</em> and someone confidently replies, <em>&#8220;DDD.&#8221;; LOL.</em></p><p>I&#8217;ve seen this answer so many times that it no longer surprises me. Instead, it makes me wonder how we collectively ended up treating <a href="/__u/nidly.substack.com/p/domain-driven-design-in-the-ai-era?r=a3p8i">Domain-Driven Design </a>as if it were an architectural style. This isn&#8217;t a criticism of the people giving that answer. It&#8217;s an easy mistake to make, and the industry has unintentionally reinforced it for years.</p><p>Today, DDD is often used as shorthand for almost anything. Sometimes it means microservices. Sometimes it means <a href="/__u/nidly.substack.com/p/why-clean-architecture-breaks-down?r=a3p8i">Clean Architecture</a> or Hexagonal Architecture. Sometimes it simply means having a domain layer. I&#8217;ve even seen experienced engineers talk about &#8220;DDD architecture&#8221; in conference talks, blog posts, and job descriptions without anyone questioning the phrase.</p><p>That recurring misunderstanding is what motivated me to write this article. The issue isn&#8217;t just terminology. Once you start thinking of Domain-Driven Design as an <a href="/__u/nidly.substack.com/p/why-software-architecture-matters?r=a3p8i">architecture</a>, you also start expecting it to answer architectural questions it was never meant to answer. That&#8217;s where design discussions become confusing and technical decisions become weaker.</p><div><hr></div><h2>Why This Confusion Is So Common</h2><p>The confusion didn&#8217;t appear out of nowhere. In many ways, it&#8217;s a natural consequence of how Domain-Driven Design has been taught and adopted over the last two decades.</p><p>DDD introduced concepts such as <strong>Bounded Contexts</strong>, <strong>Aggregates</strong>, <strong>Context Maps</strong>, and <strong>Domain Events</strong>. Those concepts often influence how systems are decomposed, so it&#8217;s easy to blur the line between modeling decisions and architectural decisions. Imagine an Order Management bounded context that eventually becomes its own microservice with its own database. Looking only at the finished system, it feels as though DDD defined the architecture.</p><blockquote><p>It didn&#8217;t.</p></blockquote><p>DDD helped identify a business boundary. Deciding to implement that boundary as a microservice was an architectural decision. Those decisions are closely related, but they happen at different levels of abstraction.</p><p>History also contributed to the misunderstanding. <a href="https://www.linkedin.com/in/ericevansddd/">Eric Evans</a> published <em>Domain-Driven Design</em> in 2003, years before microservices became mainstream. The book focused on understanding complex business domains and building better models, not on distributed deployment. As the industry later embraced service-oriented and distributed systems, architects naturally began using bounded contexts as a guide for service decomposition. That was a sensible evolution, but somewhere along the way many teams stopped saying that DDD <strong>influences</strong> architecture and started saying that DDD <strong>is</strong> the architecture.</p><p>The broader ecosystem reinforced the same idea. DDD is frequently discussed alongside Clean Architecture, Hexagonal Architecture, CQRS, Event Sourcing, and Microservices. They appear in the same books, conference talks, and blog posts, making them feel like different pieces of a single methodology. They aren&#8217;t. Some are architectural styles, some are implementation patterns, and some are integration techniques. Domain-Driven Design belongs to a different category altogether. It is a design approach for understanding and modeling complex business domains, not a blueprint for how software should be deployed or executed.</p><div><hr></div><h2>What Software Architecture Actually Defines</h2><p>Software architecture is concerned with the structural decisions that shape how a system is built, deployed, and operated. These decisions are expensive to reverse because they affect the entire system rather than a single component. Architecture answers questions such as how services communicate, how the system scales under load, what happens when a downstream dependency fails, how data is replicated across regions, and how production issues are observed and diagnosed. These are all runtime concerns. Architecture defines how software executes, not what the software means.</p><p>Consider the decision to adopt microservices. You&#8217;re not simply splitting a codebase into smaller pieces. You&#8217;re accepting network latency in exchange for independent deployment, replacing in-process calls with remote communication, introducing partial failures, and investing in observability, service discovery, distributed tracing, and operational tooling. A modular monolith represents a different set of trade-offs. Deployment becomes simpler, transactions remain local, and communication stays inside a single process. The complexity doesn&#8217;t disappear, but it moves from the network back into the codebase.</p><p>The same applies to many other architectural decisions. Choosing Kafka instead of RabbitMQ, gRPC instead of REST, or AWS Lambda instead of virtual machines changes how the system behaves in production. These choices influence scalability, resilience, latency, deployment, operational complexity, and ultimately the day-to-day experience of running software.</p><p>None of these decisions describes the business domain. They describe the runtime environment in which the business logic executes. Architecture is fundamentally about execution. It determines how requests flow through the system, how failures propagate, how data moves between components, and how the system behaves under real-world conditions.</p><div><hr></div><h2>What Domain-Driven Design Actually Solves</h2><p>Domain-Driven Design begins with a completely different question. Instead of asking <em>&#8220;How should this system run?&#8221;</em> it asks <em>&#8220;What business are we trying to model?&#8221;</em> Its purpose is to build software whose structure reflects the language, concepts, and rules used by the people who understand the business best.</p><p>That leads to questions such as: What does <strong>Customer</strong> mean in this context? Where does <strong>Order</strong> end and <strong>Shipment</strong> begin? Which business rules must always hold? Who owns this concept? Where should this responsibility live? These are not infrastructure questions. They are questions about understanding and modeling the domain.</p><p>This is where concepts like <strong>Ubiquitous Language</strong>, <strong>Bounded Contexts</strong>, <strong>Aggregates</strong>, <strong>Entities</strong>, and <strong>Value Objects</strong> become valuable. They help teams create a shared understanding of the business and translate that understanding into software. A bounded context exists to define a consistent business model, not a microservice. An aggregate exists to protect business invariants, not to dictate how transactions are implemented.</p><p>That&#8217;s an important distinction. An aggregate tells us which business rules must remain consistent together. It doesn&#8217;t tell us whether consistency is enforced with a database transaction, optimistic locking, a saga, or another mechanism. Likewise, a bounded context describes where a model has a consistent meaning. It says nothing about whether that model lives inside a monolith, a module, or an independently deployed service.</p><p>Good domain models often influence architecture, but influence should not be confused with definition. Architecture determines how software is structured, deployed, and executed. Domain-Driven Design determines how the business is understood and modeled. They complement each other, but they solve fundamentally different problems.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>One Domain, Five Architectures</h2><p>The easiest way to see why Domain-Driven Design is not an architecture is to keep the domain exactly the same while changing the architecture around it.</p><p>Imagine a mature e-commerce platform with familiar business concepts: <strong>Orders</strong>, <strong>Products</strong>, <strong>Customers</strong>, <strong>Inventory</strong>, and <strong>Billing</strong>. The domain contains the same ubiquitous language, the same bounded contexts, the same aggregates, and the same business rules. An order cannot be placed without available inventory. Pricing follows business rules. Fulfillment remains separate from billing. None of that changes. Now change only the architecture.</p><p>In a <strong>monolith</strong>, everything runs inside a single process. The bounded contexts live as modules within the same codebase, and database transactions enforce consistency. The deployment is simple, but the domain model remains exactly the same.</p><p>Move to a <strong>modular monolith</strong>, and very little changes from a business perspective. The same bounded contexts now communicate through explicit module boundaries instead of directly accessing each other&#8217;s internals. The architecture has become more disciplined, but the language, aggregates, and business rules are still identical.</p><p>Split the system into <strong>microservices</strong>, and the deployment changes dramatically. Each bounded context becomes an independently deployed service with its own database. Calls that were once local become network requests. Partial failures become possible. Service discovery, retries, and observability suddenly matter. Yet the domain itself hasn&#8217;t changed. An Order is still an Order, and the same business invariants still apply.</p><p>Replace synchronous communication with an <strong>event-driven architecture</strong>, and the runtime changes again. Instead of directly calling the Inventory service, the Order service publishes an <code>OrderPlaced</code> event. Inventory reacts by reserving stock and publishing its own event. The communication model is entirely different, but the underlying business concepts and rules remain unchanged.</p><p>Finally, deploy the same system using <strong>serverless</strong> functions. Business logic now executes inside short-lived functions triggered by HTTP requests or events. Infrastructure, scaling, and deployment are fundamentally different, yet the ubiquitous language, aggregates, bounded contexts, and business invariants remain exactly where they were. That&#8217;s the point. Five very different architectures. One domain model.</p><p>If Domain-Driven Design were an architecture, changing the architecture would require changing the domain model. In practice, the opposite is usually true. You can redesign how the system is deployed, how services communicate, or how workloads scale without changing the meaning of <strong>Order</strong>, <strong>Customer</strong>, or <strong>Inventory</strong>.</p><blockquote><p> The architecture evolved, and the domain didn&#8217;t.</p></blockquote><div><hr></div><h2>How DDD Influences Architecture Without Becoming Architecture</h2><p>Saying that Domain-Driven Design is not an architecture doesn&#8217;t mean the two are unrelated. Quite the opposite. A well-designed domain model often leads to better architectural decisions. The important distinction is that DDD informs architecture, but it doesn&#8217;t define it.</p><p>A <strong>Bounded Context</strong> is a good example. Once you&#8217;ve discovered that Sales, Billing, and Fulfillment each have different models of what a <em>Customer</em> or <em>Product</em> means, you&#8217;ve identified a business boundary. That boundary is valuable regardless of how the system is deployed.</p><blockquote><p>What happens next is an architectural decision.</p></blockquote><p>You may keep each bounded context as a module inside a modular monolith. You may split them into independent microservices. You may even deploy some separately while keeping others together. DDD identifies the boundary. Architecture decides how that boundary is enforced at runtime.</p><p>The same principle applies to <strong>Aggregates</strong>. An aggregate defines a consistency boundary within the domain. It tells us which business rules must remain valid together. It does not tell us how consistency should be implemented.</p><p>In a monolith, a database transaction may be enough. In a distributed system, the same aggregate might influence the decision to avoid splitting certain data across services. Another team may choose optimistic locking, while another adopts event sourcing or sagas. Those are architectural trade-offs built on top of the same domain model.</p><p>Context Maps provide another example. They describe how bounded contexts relate from a business perspective: upstream and downstream relationships, conformist integrations, or anti-corruption layers. They don&#8217;t prescribe whether those integrations happen through REST, gRPC, asynchronous messaging, or shared libraries. Those choices belong to architecture.</p><p>This is why saying <em>&#8220;DDD generates the architecture&#8221;</em> is <em><strong>misleading</strong></em>. DDD provides constraints, language, and business boundaries. Architecture transforms those inputs into deployment models, communication patterns, and runtime behavior. Good domain modeling leads to better architecture. It never replaces architecture.</p><div><hr></div><h2>DDD and Architecture Side by Side</h2><p>The distinction becomes much clearer when the concepts are compared directly.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9o6-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9o6-!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png 424w, /__u/substackcdn.com/image/fetch/$s_!9o6-!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png 848w, /__u/substackcdn.com/image/fetch/$s_!9o6-!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9o6-!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9o6-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png" width="688" height="429" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:429,&quot;width&quot;:688,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:40734,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/203601354?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9o6-!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png 424w, /__u/substackcdn.com/image/fetch/$s_!9o6-!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png 848w, /__u/substackcdn.com/image/fetch/$s_!9o6-!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9o6-!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2153b154-8d40-47ff-9234-2f0ceae7abe5_688x429.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Everything on the left is concerned with <strong>meaning</strong>.</p><p>Everything on the right is concerned with <strong>execution</strong>.</p><p>Changing your architecture doesn&#8217;t change what an <strong>Order</strong>, <strong>Customer</strong>, or <strong>Invoice</strong> means. Likewise, improving your domain model doesn&#8217;t automatically dictate whether your system should be a monolith, a collection of microservices, or an event-driven platform. The two disciplines work together, but they solve different classes of problems.</p><div><hr></div><h2>Why This Matters More in the AI Era</h2><p>The distinction between Domain-Driven Design and software architecture has become even more important as modern systems incorporate LLMs, retrieval pipelines, AI agents, workflow engines, and vector databases. Many teams now repeat the same mistake they once made with microservices by treating AI components as part of the domain model simply because they&#8217;re central to the application. In reality, most of them belong to the architecture, not the domain.</p><p>A retrieval pipeline isn&#8217;t a bounded context. An embedding model isn&#8217;t part of your business domain. An orchestration framework isn&#8217;t your domain model. They&#8217;re architectural components, just like databases, message brokers, caches, and API gateways. Their job is to support the business, not define it.</p><p>The domain questions remain unchanged. What does an AI-generated recommendation represent? Should it become part of the business record? Who owns it? Which business rules apply? Does it need to be auditable? These are modeling questions, and DDD provides useful tools for answering them. The architectural questions are completely different. How do you cache prompts? How do you route requests across multiple models? How do you recover from partial failures in long-running agent workflows? How do you keep retrieval latency within your SLO? Those are runtime concerns, and they belong to software architecture.</p><p>As AI systems become more sophisticated, separating these conversations becomes even more valuable. One discussion is about understanding the business. The other is about executing software reliably at scale. AI hasn&#8217;t changed that distinction. It has simply made it far more obvious.</p><div><hr></div><h2>Things DDD Does <strong>Not</strong> Tell You</h2><p>One of the easiest ways to understand Domain-Driven Design is to look at everything it deliberately leaves open. DDD doesn&#8217;t require Microservices, CQRS, Event Sourcing, Kafka, Hexagonal Architecture, or Clean Architecture. It doesn&#8217;t require every state change to become a domain event, nor does it prescribe a particular database, messaging platform, or deployment strategy.</p><p>Those decisions belong to software architecture. They may complement a strong domain model, but they are not part of Domain-Driven Design itself.</p><p>What DDD actually requires is much less glamorous and far more demanding: deep collaboration with domain experts, a shared language reflected in both conversations and code, carefully identified business boundaries, and a willingness to refine the model as understanding evolves. That discipline works equally well in a monolith, a modular monolith, a microservices platform, or an event-driven system, because Domain-Driven Design was never an architecture to begin with.</p><div><hr></div><h2>Conclusion</h2><p>Architecture determines how software runs. Domain-Driven Design determines what the software means. Confusing the two isn&#8217;t just an imprecise use of terminology; it&#8217;s a category mistake that leads teams to answer architectural questions with modeling concepts and modeling questions with infrastructure patterns.</p><p>The confusion is understandable. DDD, Clean Architecture, Hexagonal Architecture, CQRS, Event Sourcing, and Microservices are often introduced together, discussed together, and frequently adopted together. Over time, many engineers started treating them as parts of the same methodology. They aren&#8217;t. DDD helps you understand the business. Architecture helps you build, deploy, scale, and operate the system. Good software requires both, but neither replaces the other.</p><p>The next time someone asks, <em>&#8220;What architecture are you using?&#8221;</em>, <strong>&#8220;DDD&#8221;</strong> isn&#8217;t the right answer. It&#8217;s a description of how you model your domain, not how your software executes. Those are two different conversations, and keeping them separate is the first step toward making better engineering decisions.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/domain-driven-design-is-not-an-architecture?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/domain-driven-design-is-not-an-architecture?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/domain-driven-design-is-not-an-architecture?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p></p>]]></content:encoded></item><item><title><![CDATA[Designing Memory for GenAI Systems: From Storage Pipelines to Control Loops]]></title><description><![CDATA[Most AI memory systems are search engines in denial. Here's the control-loop architecture that replaces them.]]></description><link>https://nidly.substack.com/p/designing-memory-for-genai-systems</link><guid isPermaLink="false">https://nidly.substack.com/p/designing-memory-for-genai-systems</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 06 Jul 2026 04:30:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Gcmg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Gcmg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Gcmg!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Gcmg!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Gcmg!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Gcmg!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Gcmg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2921577,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/198834745?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Gcmg!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Gcmg!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Gcmg!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Gcmg!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba40c75-8ba1-4b51-99a4-f78efa78fe59_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Every few weeks, another team announces that their assistant &#8220;now has memory.&#8221; Under the hood, it is almost always the same system: messages are chunked, chunks are embedded, embeddings are stored, and at inference time a similarity search pulls top-k results back into the prompt.</p><p><strong>This is not memory</strong>. It is a search system attached to a chat log. It occasionally produces something useful, which is precisely why its failure mode is easy to miss. It does not fail by being wrong. It fails by being indifferent: it retrieves things that match, not things that are true.</p><p>The confusion starts with vocabulary. <em><strong>Chat history is not memory</strong></em>; it is a transcript. The context window is not memory either; it is a transient compute constraint, a limited working set for attention. Calling either of these &#8220;memory&#8221; is like calling CPU cache a database. The components are adjacent; the semantics are not.</p><p>This essay argues a different position: memory in GenAI systems is not a storage problem. <strong>It is a control problem</strong>. A memory system does not merely preserve information about a user. It maintains an evolving model of the user and uses that model to shape system behavior over time. That requires interpretation at write time, explicit conflict resolution, and policy-driven state transitions. Retrieval is secondary. In some designs, it is optional.</p><div><hr></div><h2>Why the standard pipeline fails</h2><p>The dominant architecture is a simple pipeline: ingest, chunk, embed, store, retrieve, rank, inject. It is popular for a reason. Every stage is well understood, composable, and shippable in a sprint. The issue is not implementation quality. The issue is structural omission: nowhere in this pipeline does the system decide what an event <em>means</em>.</p><blockquote><p>That omission produces three failure modes, and they compound over time.</p></blockquote><p>The first is noise accumulation. A pipeline that stores everything treats &#8220;I&#8217;m vegetarian&#8221; and &#8220;let&#8217;s pretend I&#8217;m vegetarian for this recipe&#8221; as equivalent signals, because both land in similar regions of embedding space. Over time, the memory store fills with hypotheticals, role-play artifacts, debugging traces, and transient statements. Retrieval degrades not because search fails, but because the underlying corpus was never filtered for signal. Ranking cannot recover what was never separated.</p><p>The second is contradiction accumulation, and it is more damaging than noise because it creates confident inconsistency. A user says in January that they work at a fintech startup, and later in September mentions a role at a hospital network. A retrieval system will happily surface both when asked about employment, because both are semantically valid matches. What it cannot represent is supersession. &#8220;This replaced that&#8221; is not a vector relationship; it is a state transition. The system has no concept of transition, so it cannot know which fact is current.</p><p>The third is the absence of state evolution. Real understanding of a user is not a collection of observations, but a continuously revised model. Someone who asked basic Kubernetes questions in March and is debugging admission webhooks in November has changed in a meaningful way. A pure append-only system preserves both states equally, flattening trajectory into accumulation. What gets lost is the only thing that matters: change over time.</p><p>In all three cases, the core issue is the same. The system stores events but never interprets them. And without interpretation, retrieval becomes blind. It can locate similar text. It cannot maintain a coherent model of reality.</p><div><hr></div><h2>A taxonomy that earns its keep</h2><p>Most memory systems borrow their taxonomy from cognitive science: short-term, long-term, episodic, semantic, sometimes procedural. It is a comfortable abstraction, and it does almost no engineering work. Knowing that a fact is &#8220;episodic&#8221; does not tell you how it is stored, when it is retrieved, or whether it should exist at all. It is classification by analogy, not by consequence.</p><blockquote><p>A more useful taxonomy starts from a different question: not what memory <em>is</em>, but what memory <em>does</em>.</p></blockquote><p>Some memory changes inference behavior. These are constraints that must reliably shape model output across requests: user preferences, style rules, safety constraints, explicit instructions. They are not &#8220;data to retrieve.&#8221; They are configuration that must be consistently applied. Failure here is immediately visible to the user, so this category behaves like system-critical state.</p><p>Some memory changes retrieval behavior. This includes past conversations, prior topics, and contextual background that is only relevant when the current query activates it. It does not need to be globally enforced, only selectively surfaced. This is the one area where semantic search is actually appropriate, because uncertainty and partial recall are acceptable properties.</p><p>Some memory changes nothing. Logs, transcripts, analytics events, audit trails. This is not memory; it is observability data. The failure mode here is not technical, it is conceptual: treating logging data as if it participates in user modeling. Once it enters prompts, it stops being observability and becomes noise.</p><p><strong>The separation is not semantic, it is behavioral.</strong> The only question that matters is simple: if this piece of information were removed, what would change in system behavior? If the answer is &#8220;nothing,&#8221; it never belonged in memory in the first place.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>From pipeline to control loop</h2><p>The architectural shift follows directly from this classification. The standard model is a pipeline:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9rsJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9rsJ!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png 424w, /__u/substackcdn.com/image/fetch/$s_!9rsJ!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png 848w, /__u/substackcdn.com/image/fetch/$s_!9rsJ!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9rsJ!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9rsJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1085903,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/198834745?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9rsJ!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png 424w, /__u/substackcdn.com/image/fetch/$s_!9rsJ!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png 848w, /__u/substackcdn.com/image/fetch/$s_!9rsJ!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9rsJ!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a29b132-c27a-4819-9db3-e6379efa6be8_1774x887.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>It treats all memory as the same kind of artifact: something to be stored and later searched. <strong>The problem is not implementation quality. The problem is that nothing in this pipeline performs interpretation</strong>.</p><p>A more accurate model is a control loop:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!NNgy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!NNgy!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!NNgy!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!NNgy!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NNgy!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!NNgy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1282296,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/198834745?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!NNgy!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!NNgy!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!NNgy!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NNgy!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b41e297-f01f-41ec-b2c0-9c57a24305e9_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The missing components are the ones that actually define the system. The interpretation layer is where meaning is assigned. It decides whether an event is durable or transient, whether it expresses a preference or a one-off statement, and how much confidence to assign to it. The difference between &#8220;I am vegetarian&#8221; and &#8220;let&#8217;s pretend I&#8217;m vegetarian for this recipe&#8221; is not syntactic; it is semantic intent. Without interpretation, both are stored as equivalent signals.</p><p>The reconciliation layer is where state is maintained. Incoming interpreted events are not appended; they are applied. They either overwrite existing state, merge with it, decay it, or are discarded. This is where contradictions stop being stored as competing facts and become resolved transitions in a single evolving model.</p><p>At this point, the system is no longer a document store with retrieval. It is an aggregate with state transitions. The user profile is not a collection of memories; it is a current state derived from a history of interpreted changes. Storage becomes secondary: useful for audit, debugging, or reprocessing, but not identical to memory itself.</p><p>This is the key inversion: most current systems treat memory as data with retrieval layered on top. A correct system<strong> treats memory as state</strong>, with storage and retrieval as supporting mechanisms around it.</p><div><hr></div><h2>The write path is the system</h2><p>Once you accept this framing, the main engineering complexity shifts to the write path. This is precisely the part that most &#8220;memory pipelines&#8221; treat as an afterthought. Ingestion itself is simple. Conversation turns, tool outputs, and explicit user directives arrive as events. Nothing controversial happens here.</p><blockquote><p>The real system begins at interpretation.</p></blockquote><p>The interpretation engine is typically an LLM pass, synchronous or deferred, that extracts candidate facts along dimensions like type, scope, and confidence. This layer defines everything downstream. And that dependency is uncomfortable but unavoidable: a memory system is only as good as a model&#8217;s ability to decide what mattered in a conversation. The alternative is a system with no judgment at all, which is exactly what naive storage produces. Write policies then decide how interpreted candidates become state.</p><p>Explicit user directives carry high confidence and behave as hard overrides. If a user says &#8220;I&#8217;ve moved to Berlin,&#8221; the previous city is not preserved as a competing fact; it is replaced.</p><p>Inferred preferences are weighted updates rather than binary inserts. One signal that the user prefers shorter responses is weak. Five consistent signals over time form a preference. The system must represent this as a gradual update in confidence, not a switch.</p><p>Facts that are no longer reinforced decay over time. If a tool stack or preference has not been observed for a long period, its confidence should decrease. Silence is weak evidence, and systems that treat it as certainty accumulate stale state.</p><p>The write path is therefore the quality gate of the entire system. Every filtering and interpretation decision made here removes pressure from retrieval and ranking later. Most teams do the opposite: they invest in increasingly complex retrieval to compensate for a noisy memory store. The leverage is structurally earlier.</p><div><hr></div><h2>The read path, demoted but not dismissed</h2><p>Once write-time interpretation is in place, read-time design becomes simpler by necessity. The first stage is structured lookup, not search. The reconciled profile is small, typed, and deterministic. It contains constraints and high-confidence preferences that must always enter the prompt when relevant. No embeddings, no ranking, no probabilistic recall. Behavior-critical memory cannot depend on retrieval luck.</p><p>A second stage performs hybrid retrieval over episodic memory, but only when the query justifies it. Ranking in this layer should combine recency, confidence assigned at write time, and stability over time. A fact that has been rewritten or corrected multiple times is inherently unstable and should be surfaced with caution or excluded entirely.</p><p>Context assembly is fundamentally a budgeting problem under strict token constraints. Allocation should be explicit:</p><ul><li><p>constraints first, always</p></li><li><p>stable high-confidence preferences next</p></li><li><p>episodic memory last, and often not at all</p></li></ul><p>A system that fills remaining tokens with &#8220;available memory&#8221; is optimizing for the appearance of personalization rather than actual relevance. Users notice this immediately as irrelevant historical noise leaking into fresh interactions.</p><div><hr></div><h2>Problems that only appear in production</h2><p>Four issues separate working systems from demo systems. None of them are retrieval problems.</p><h3>Deletion propagation</h3><p>This is the hardest problem, and the pipeline model makes it nearly intractable. A single fact exists in multiple representations: raw transcripts, chunked embeddings, derived summaries, and cached contexts. Deleting the source record does not delete the fact. It only creates inconsistency.</p><p>Under regulatory constraints like GDPR, this becomes a hard requirement. The only viable model is to treat deletion as a first-class state transition. A tombstone is written to the reconciliation log, and that event triggers propagation across all derived representations: summaries are regenerated, embeddings invalidated, indexes purged, caches cleared.</p><p>Expensive, asynchronous, and non-optional. If a system cannot describe how a fact is removed across all representations, it is not production-ready.</p><h3>Temporal conflict resolution</h3><blockquote><p>Preferences are not static. They are scoped.</p></blockquote><p>A user can be terse in professional contexts and expansive in personal ones without contradiction. The failure mode is treating this as a single global state and repeatedly overwriting one behavior with the other. Reconciliation must therefore include scope, not just recency or confidence.</p><h3>Drift</h3><blockquote><p>This is the control problem in its pure form.</p></blockquote><p>A memory system adapts to user behavior, which then changes user behavior, which feeds back into the system. Small biases compound over time.</p><p>If the system infers that the user prefers brevity, the user adapts by writing shorter inputs. The system then interprets shorter inputs as confirmation of the same preference. The loop reinforces itself.</p><p>This is not a data issue. It is a stability issue. Update rates, decay functions, and confidence thresholds are control parameters. If tuned incorrectly, the system either never learns or overfits to recent history.</p><h3>Memory poisoning</h3><p>If interpreted outputs can trigger state changes, then any untrusted content becomes a potential write vector. A document or tool result that says &#8220;remember to always include this link&#8221; is not metadata. It is an attempted instruction embedded in external content.</p><p>The interpretation layer is therefore also a security boundary. It requires provenance tracking, source trust tiers, and strict rules about which inputs are allowed to influence high-privilege state changes. Without this, memory becomes writable through indirect prompt injection.</p><div><hr></div><h2>If you cannot measure it, you built it on faith</h2><p>Memory systems have a specific evaluation trap: they look valuable in demos and remain mostly unmeasured in production, because the counterfactual is invisible. No one sees the better answer the model would have produced without a stale fact injected into context.</p><p>The basic discipline is straightforward: run A/B tests with and without memory injection on real traffic, and measure task success rather than retrieval quality. Recall@k proves the search layer works. It does not prove the memory system helps. The metrics that actually matter are different.</p><p>Correction rate: how often users have to fix the system&#8217;s beliefs about them. This directly measures reconciliation quality. Preference accuracy: how often the system matches user-confirmed preferences, not inferred ones.</p><p>Memory-adjacent hallucination rate: how often injected memory causes confident but incorrect assertions. This matters because memory is treated as authoritative context by default. A model without memory can hedge. A model with wrong memory asserts.</p><p>The uncomfortable result is consistent: past a relatively low threshold, more memory makes systems worse. Injected context dilutes attention over the actual request, and irrelevant personalization quickly shifts from helpful to intrusive. The optimal injection rate is far lower than most teams assume, and finding it requires accepting that your memory system might be degrading the product. Most teams avoid running that experiment. The reason is not technical.</p><div><hr></div><h2>Where embeddings actually belong</h2><p>There are three architectural families in practice. The distinction is clearer than the tooling suggests.</p><p><strong>Structured memory</strong> (typed profiles and fact tables maintained via reconciliation) is the correct foundation for behavior-critical state. It is deterministic, auditable, cheap to query, and straightforward to delete. Its cost is the interpretation layer, which is exactly the component this system requires anyway.</p><p><strong>Vector memory</strong> is effective for episodic recall: conversations, documents, and loosely structured context. It is a strong indexing mechanism for &#8220;what was said&#8221; or &#8220;what was discussed.&#8221; It is not a mechanism for resolving contradictions, representing supersession, or enforcing correctness over time. Treating it as the core memory layer forces everything else in the system to compensate for its lack of state semantics.</p><p><strong>Agentic memory</strong>, where the model decides when to read and write through tools, is the most interesting direction because it relocates interpretation into the model itself rather than a pipeline stage. It assumes that runtime judgment is better than offline extraction. It introduces new problems around consistency, cost, and controllability, but its direction is clear: it strengthens write-time interpretation rather than removing it.</p><p>Across all three approaches, the pattern is consistent. Embeddings are an indexing technology. They answer one question: &#8220;what is similar to this.&#8221; Similarity is not truth, not recency, not intent, and not replacement. A memory system built on embeddings alone commits to approximating every user model as similarity search, then spends its complexity budget patching that limitation with rankers and heuristics.</p><p>Indexes belong behind state, not in place of it.</p><div><hr></div><h2>Memory is what survives interpretation</h2><p>The storage-centric model fails because it optimizes the wrong loop. It assumes the problem is retrieval, so it invests in search. The real problem is maintaining a correct, evolving model of the user over time. That is not a search problem. It is a system of interpretation, reconciliation, and controlled state transition.</p><p>This is not a new class of problem. It is the same structure solved in event-sourced systems, CRDTs, and control theory, applied to a new domain with different failure modes.</p><p>The design implication is simple:</p><blockquote><p>Build the write path first.</p></blockquote><p>Make interpretation a first-class component with ownership and evaluation.<br>Treat reconciliation as your core state machine, not a side effect.<br>Demote embeddings to an index layer. Measure against a no-memory baseline before trusting any perceived improvement. If retrieval is the center of your system, you do not have memory. You have search.</p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Ubiquitous Language: AI Can Finally Audit]]></title><description><![CDATA[How LLMs catch terminology drift before it corrupts your domain model in domain driven design.]]></description><link>https://nidly.substack.com/p/ai-as-a-domain-language-auditor</link><guid isPermaLink="false">https://nidly.substack.com/p/ai-as-a-domain-language-auditor</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 29 Jun 2026 07:19:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FPCj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!FPCj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!FPCj!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png 424w, /__u/substackcdn.com/image/fetch/$s_!FPCj!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png 848w, /__u/substackcdn.com/image/fetch/$s_!FPCj!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FPCj!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!FPCj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png" width="1402" height="1122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1122,&quot;width&quot;:1402,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2676448,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/201901019?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!FPCj!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png 424w, /__u/substackcdn.com/image/fetch/$s_!FPCj!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png 848w, /__u/substackcdn.com/image/fetch/$s_!FPCj!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FPCj!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57457c4d-5000-4b74-a7ab-2ec762ab5f67_1402x1122.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>1. The Problem Nobody Talks About</h2><p>In theory, Ubiquitous Language is one of the strongest ideas in <a href="/__u/nidly.substack.com/p/domain-driven-design-in-the-ai-era">Domain-Driven Design</a>. It creates a shared vocabulary between engineers and domain experts where every term has a clear meaning, and every concept maps directly to the business. In an ideal world, code feels like a conversation with the domain.</p><p>In reality, this works only for a short time, usually a few months. Not because people stop caring or teams become careless, but because the system grows. More teams, more services, more boundaries. Gradually, the language starts to drift.</p><p><a href="https://www.linkedin.com/in/ericevansddd">Eric Evans</a> called it <a href="https://martinfowler.com/bliki/UbiquitousLanguage.html">Ubiquitous Language</a> for a reason. It is supposed to be everywhere. But in real systems, &#8220;everywhere&#8221; does not exist. There are multiple services, multiple teams, multiple timelines, and no shared global view anymore. At some point, someone asks: &#8220;What exactly did you mean by Customer here?&#8221;</p><p>That question is never the root problem. It is just where the problem becomes visible.</p><div><hr></div><h2>2. Why Ubiquitous Language Breaks in Real Systems</h2><h3>2.1 The Scale Problem</h3><p>DDD assumes a small, aligned team with a shared whiteboard, constant conversations, and direct access to domain experts. In that environment, language consistency is natural and drift is easy to detect.</p><p>At scale, this breaks down. You now have multiple teams, multiple bounded contexts, engineers who never meet each other, and services evolving independently. There is no shared mental model anymore, only system artifacts: code, events, APIs, and documentation.</p><div><hr></div><h3>2.2 Language Lives Everywhere</h3><p><a href="https://martinfowler.com/books/dsl.html">Domain language</a> is not stored in one place. It is distributed across the system and evolves independently in each layer:</p><ul><li><p>Source code: entities, services, variables</p></li><li><p>Event schemas: Kafka topics, queue payloads</p></li><li><p>APIs: endpoints, request and response models</p></li><li><p>Databases: tables, columns, relationships</p></li><li><p>Documentation: ADRs, wiki pages</p></li><li><p>Tickets: Jira stories, specifications</p></li></ul><p>Each of these evolves on its own timeline. Code changes every commit. Events change when integrations evolve. Documentation lags behind reality. There is no system-level synchronization for meaning. No compiler checks whether &#8220;Customer&#8221; still means the same thing across all contexts.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>2.3 Semantic Drift Is Silent</h3><p>This is the dangerous part. Semantic drift does not fail loudly. There is no compiler error, no test failure, and no immediate alert. Everything keeps working. But meaning slowly diverges underneath. A term gets reused in a new context. A team slightly reinterprets a concept. A rename happens in one service but not in others. Over time, one word starts carrying multiple meanings and no one notices.</p><p>Until production breaks or someone asks: &#8220;Which version of Customer are we talking about?&#8221; By then, the drift is already embedded in the system. The real issue is simple: no one has visibility into language consistency at system scale.</p><div><hr></div><h2>3. AI as a Domain Language Auditor</h2><h3>3.1 What AI Is NOT Here</h3><p>Before defining what the auditor does, it is worth being explicit about what it does not do. AI is not part of the domain model. It has no business logic. It does not decide what <code>Customer</code> means. It does not define bounded context boundaries. It does not own any domain concept.</p><p>The moment AI starts making semantic decisions about your domain, you have lost the very thing DDD is trying to protect: the domain belongs to domain experts, not infrastructure.</p><p>AI is not a decision maker. When the auditor detects that <code>Account</code> is used differently across three services, it does not resolve the conflict. It cannot. Resolving semantic disagreements requires business judgment, team negotiation, and explicit decisions about context boundaries.</p><blockquote><p>Those responsibilities remain human.</p></blockquote><p>AI is not a replacement for event storming, collaborative modeling, or domain workshops. Those activities create shared understanding. The auditor simply helps maintain it over time.</p><h3>3.2 What AI IS: An External Semantic Observer</h3><p>The value of AI is not in making domain decisions. It is in spotting patterns that humans struggle to see at scale. Modern systems spread their language across hundreds of files, schemas, APIs, documents, and tickets. No single engineer has a complete view of how every domain term is used across the entire landscape.</p><p>AI does. Not because it understands the business better than people do, but because it can examine large amounts of domain language simultaneously and identify inconsistencies that would otherwise remain invisible. That makes it remarkably effective at answering a simple question:</p><p>Is this term being used consistently across the system?</p><p>Viewed this way, the auditor becomes another layer in the observability stack. We already use logs to understand runtime behavior, metrics to understand performance, and traces to understand request flows.</p><p>The auditor adds something different: semantic observability. It provides visibility into whether the language of the domain is evolving consistently or quietly drifting apart.</p><blockquote><p>It observes. It reports. Humans decide.</p></blockquote><div><hr></div><h2>4. What the Auditor Analyzes</h2><p>A well-scoped AI auditor operates across every artifact where domain language lives.</p><p><strong>Source code</strong> is usually the starting point. It carries the most structured vocabulary in the system. The auditor parses entity names, aggregate roots, repository interfaces, service boundaries, and method signatures. From this, it reconstructs the actual language the system is using, not the language teams think they are using. It then maps these terms back to bounded contexts.</p><p><strong>Event schemas</strong> are often where problems surface first. In event-driven systems, field names quietly define the shared language between services. This is where drift becomes visible early for example, <code>customerId</code> in one event and <code>accountId</code> in another, both pointing to the same underlying concept but described differently.</p><p><strong>API contracts</strong> such as OpenAPI specs, gRPC definitions, or GraphQL schemas sit at system boundaries. This is where language becomes expensive. If two teams disagree on what a resource means at the API level, every integration between them inherits that ambiguity by default.</p><p><strong>Database models</strong> tend to preserve older versions of language. They are where terminology goes to live forever. A table named <code>users</code> might still exist in a system that now consistently uses the term <code>members</code>. These inconsistencies rarely break anything, but they quietly accumulate and increase cognitive load.</p><p><strong>Documentation and tickets</strong> are the only place where intent is supposed to live. Code shows behavior, but docs explain meaning. In practice, they often diverge. Comparing Jira epics, ADRs, and wiki pages with actual code usage often reveals the gap between what was intended and what was implemented.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-as-a-domain-language-auditor?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/ai-as-a-domain-language-auditor?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/ai-as-a-domain-language-auditor?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2>5. What the Auditor Detects</h2><p>The auditor does not just collect terms. It looks for patterns of inconsistency. The output typically falls into four categories.</p><p><strong>Same term, different meanings.</strong> A single word like <code>Customer</code> is used across multiple services, but with different semantics. In Billing, it means a paying entity. In CRM, it means anyone who has interacted with the company. In Marketing, it means a contact in a campaign list. Same label, different reality.</p><p><strong>Different terms, same concept.</strong> The inverse problem is just as common. <code>User</code>, <code>Member</code>, <code>Subscriber</code>, and <code>Account</code> appear across the system, sometimes interchangeably. The auditor flags these cases when structure and behavior strongly overlap, suggesting either duplication or unclear boundaries.</p><p><strong>Missing definitions.</strong> Some terms are heavily used in code and tickets but have no formal definition anywhere. They exist operationally but not conceptually. This is often where future confusion and drift start.</p><p><strong>Context boundary violations.</strong> A term that belongs clearly to one bounded context appears inside the core logic of another, not in integration code, but in domain logic itself. This usually signals either a missing abstraction or a boundary that was drawn incorrectly.</p><div><hr></div><h2>6. A Concrete Example: The &#8220;Customer&#8221; Problem</h2><p>Consider a mid-size e-commerce platform with three bounded contexts: Billing, CRM, and Marketing. All of them use the same term: <code>Customer</code>. In practice, they are describing three different things.</p><p><strong>In Billing:</strong></p><pre><code><code>Customer {
  id: UUID
  subscriptionTier: Enum(FREE, PRO, ENTERPRISE)
  billingEmail: string
  paymentMethodId: string
  status: Enum(ACTIVE, PAST_DUE, CANCELLED)
}
</code></code></pre><p>Here, <code>Customer</code> is a paying entity. A Customer only exists when there is a financial relationship. The lifecycle is driven entirely by payment state.</p><p><strong>In CRM:</strong></p><pre><code><code>Customer {
  id: UUID
  contactEmail: string
  lifetimeValue: decimal
  acquisitionChannel: string
  assignedRep: string
  stage: Enum(PROSPECT, ACTIVE, CHURNED, REACTIVATED)
}
</code></code></pre><p>Here, <code>Customer</code> is a relationship entity. The focus is not payment, but the sales lifecycle. A <code>PROSPECT</code> is already a Customer in CRM, even if they have never paid.</p><p><strong>In Marketing:</strong></p><pre><code><code>Customer {
  id: UUID
  emailAddress: string
  segment: string[]
  lastCampaignOpenedAt: timestamp
  optInStatus: bool
}
</code></code></pre><p>Here, <code>Customer</code> is an audience entity. It is purely about communication. Someone can exist here without ever being in CRM or Billing. Three services. Three models. One term. No shared meaning.</p><p>The auditor scans all three contexts and produces a single finding:</p><pre><code><code>SEMANTIC CONFLICT DETECTED
Term: "Customer"
Found in: billing-service, crm-service, marketing-service

Billing: paying entity with financial lifecycle (5 usages in domain layer)
CRM: relationship entity with sales lifecycle (12 usages in domain layer)
Marketing: audience entity with communication lifecycle (8 usages in domain layer)

Lifecycle overlap: NONE
Attribute overlap: email (all contexts), id (all contexts, different semantics)

Divergence confidence: 0.91 (HIGH)
Recommended action: Human review required
Possible resolution: context-specific naming (BillingAccount, CRMContact, MarketingSubscriber)
</code></code></pre><p>This is important: the auditor does not decide anything. It does not refactor the model. It does not rename concepts. It simply makes the ambiguity visible, in a structured way, before it spreads further.</p><div><hr></div><h2>7. How It Fits Into Engineering Workflows</h2><p>The key point is that the auditor is not a batch analysis tool. Language drifts continuously, so the analysis has to be continuous as well.</p><p><strong>PR-time checks</strong> are the first layer. When a PR introduces or modifies a domain term, the auditor runs a lightweight comparison against existing usage. It behaves like a linter, but for meaning instead of syntax.</p><p>The feedback is intentionally simple:</p><pre><code><code>&#9888; Domain Language Check
New term: "Subscriber"
Similar existing terms: "Member" (user-service), "Customer" (billing-service)

Possible semantic overlap detected.
Question: new concept or existing concept under a different name?
</code></code></pre><p><strong>CI/CD integration</strong> acts as a stronger gate. On merge to main, a deeper scan runs against the affected bounded context. If significant drift is detected, it is flagged before deployment. Not blocked by default. Just surfaced early.</p><p><strong>Scheduled audits</strong> run in the background. Daily or weekly. These catch slow drift the kind that no single PR introduces, but accumulates over time. The output is not urgent. It is informational. A signal for architects and tech leads:</p><p>&#8220;What changed in our language this week?&#8221;</p><p><strong>Sprint review reports</strong> connect language changes to actual delivery work. If a team introduces new terminology or shifts meaning within a context, it becomes visible during planning and review, not during incidents. This turns language consistency into part of the development rhythm, not a separate concern.</p><div><hr></div><h2>8. Output Structure of the Auditor</h2><p>The auditor is designed to support decisions, not make them.</p><p><strong>Term usage map</strong> shows where each domain term appears across the system. It makes cross-context usage visible immediately and <strong>Semantic conflict reports</strong> break down where meanings diverge. Each report includes how a term is used in each context and how those usages differ. A confidence score highlights how strong the divergence signal is.</p><p><strong>Drift delta</strong> compares current state with previous audits. It highlights what changed: new terms, renamed concepts, or shifted meanings. This is what makes drift observable over time instead of invisible. Also, <strong>Review candidates</strong> is the prioritized queue of things that need human attention. Not fixes. Not decisions. Just focused starting points for discussion, ranked by impact and uncertainty. One important constraint:<br>The auditor never produces final definitions. Never rewrites the model. Never decides meaning.</p><p>It only exposes where meaning is no longer aligned.</p><div><hr></div><h2>9. Limitations and Risks</h2><p><strong>False positives are inevitable.</strong> Not every inconsistency is a problem. In DDD, bounded contexts are meant to have their own language. The same word can legitimately mean different things in different contexts. The issue is not divergence itself, but uncontrolled divergence. The auditor must be tuned to surface signals for review, not to label everything as an error. Otherwise it becomes noise, and noisy tools get ignored.</p><p><strong>AI has no business context.</strong> The auditor can detect that <code>Member</code> and <code>User</code> share identical attributes across two services. What it cannot do is decide whether that is a modeling mistake or a deliberate design choice made for valid historical reasons. That judgment belongs to the business and the team that owns the context. The auditor only exposes patterns.</p><p><strong>The system is only as good as its inputs.</strong> If your code is clean but your documentation is vague, or your events are well-structured but your tickets are messy, the quality of insights will vary. This doesn&#8217;t break the tool, but it limits how far intent-based analysis can go. Code alone is often enough to detect structural drift, but intent always depends on human artifacts.</p><p><strong>Domain experts must stay in the loop.</strong> There is a real risk of over-trusting the output. If the auditor says two terms look equivalent, that does not mean they should be merged. It means they need attention. The output is a signal, not a decision.</p><div><hr></div><h2>10. Architectural Framing</h2><p>The placement of the auditor in the architecture matters.  The domain layer stays unchanged. It remains deterministic, explicit, and fully owned by humans. Bounded contexts define their boundaries. Aggregates enforce invariants. Domain events carry meaning. None of that depends on the auditor.</p><p>The AI auditor sits outside the domain layer, in the observability stack. Alongside logs, metrics, and traces, but operating on a different signal: meaning instead of behavior.</p><p>This distinction is important. You don&#8217;t put business logic inside your logging system. You don&#8217;t let metrics decide system behavior. The same applies here. The auditor observes; it does not participate in execution.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fRes!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fRes!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png 424w, /__u/substackcdn.com/image/fetch/$s_!fRes!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png 848w, /__u/substackcdn.com/image/fetch/$s_!fRes!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fRes!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fRes!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png" width="1402" height="1122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1122,&quot;width&quot;:1402,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1029905,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/201901019?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fRes!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png 424w, /__u/substackcdn.com/image/fetch/$s_!fRes!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png 848w, /__u/substackcdn.com/image/fetch/$s_!fRes!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fRes!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5052b1e0-a08a-43c2-999a-459c59dd9bab_1402x1122.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The value here is the separation of concerns. The domain stays stable. The auditor stays external. And meaning becomes something you can observe without interfering with it.</p><div><hr></div><h2>11. Conclusion</h2><p>Ubiquitous Language is not a static artifact. It is not something you define once during a modeling session and then forget. It is a living system property distributed across code, events, APIs, and documentation. And like any distributed system, it drifts over time. DDD gives you the tools to define language. It does not give you the tools to continuously observe it at scale.</p><blockquote><p>That is the gap this approach targets.</p></blockquote><p>The AI auditor does not define meaning. It does not replace modeling sessions. It does not resolve conflicts. It only makes one thing visible: where the system&#8217;s language is no longer aligned with itself.</p><p>This reframes language consistency as an observability problem, not a documentation problem. Documentation is static and incomplete. Observability is continuous and operational.</p><p>The analogy is simple. You don&#8217;t wait for a production incident to start logging. You don&#8217;t rely on manual inspection to understand system health. You instrument it. The same applies here. The auditor is just instrumentation for meaning.</p><p>And once language becomes observable, it stops being an invisible source of architectural drift. It becomes something you can monitor, discuss, and gradually bring back into alignment. Not by automation. Not by AI decisions. But by making the problem visible early enough that humans can still fix it cleanly.</p>]]></content:encoded></item><item><title><![CDATA[Event Storming at Scale: Why AI Changes the Discovery Game]]></title><description><![CDATA[Using AI to Keep Event Storming Results Alive and Accurate]]></description><link>https://nidly.substack.com/p/event-storming-at-scale-why-ai-changes</link><guid isPermaLink="false">https://nidly.substack.com/p/event-storming-at-scale-why-ai-changes</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 22 Jun 2026 05:31:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!n0UA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!n0UA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!n0UA!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!n0UA!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!n0UA!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!n0UA!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!n0UA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/db0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2034239,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/200783913?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!n0UA!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!n0UA!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!n0UA!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!n0UA!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb0a817e-3e4f-478f-ae35-c271f96b4825_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Event Storming was never supposed to be a methodology. <a href="https://www.linkedin.com/in/brando/">Alberto Brandolini</a> created it as a workshop format, a way to get domain experts and engineers in the same room, cover a wall with sticky notes, and build a shared <strong>understanding</strong> of a <strong>business domain </strong>faster than any requirements document ever could. For a while, it works exactly like that.</p><p>Teams walk out of a two-day workshop with a clearer picture of their domain than they gained from months of sprint planning sessions, backlog refinement meetings, and architecture reviews. Ambiguous terminology becomes visible. Business rules that lived only in people&#8217;s heads finally get written down. Different teams start speaking the same language.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Then the workshop ends. The sticky notes come down. Someone takes a few photos. A Miro board is created. A Confluence page appears. The domain model is documented, shared, and celebrated, then the domain changes.</p><p>A new regulatory requirement arrives. A business process is modified. A new integration introduces behaviors nobody discussed during the workshop. An edge case appears in production for the first time. Six months later, the model no longer represents the domain as it exists today. It represents the domain as it was understood on the day the workshop ended. This is not a criticism of <a href="https://www.eventstorming.com/">Event Storming</a>. It is simply what happens when a static artifact meets a constantly evolving business.</p><p>Most teams accept this drift as inevitable. The workshop produces a snapshot of reality, and that snapshot gradually becomes less accurate over time. The larger and more complex the domain becomes, the faster that drift happens.</p><p>What has changed recently is that AI can participate in the discovery process itself. It can analyze requirements, challenge assumptions, surface missing scenarios, review historical incidents, and continuously compare documented models against information scattered across tickets, specifications, support conversations, and production systems.</p><p>AI does not eliminate the need for Event Storming. If anything, it makes collaborative discovery even more valuable. What it changes is the amount of information that can be brought into the room and the ability to keep the resulting model alive after the workshop is over.</p><div><hr></div><h2>Why Event Storming Often Fails in Production Teams</h2><p>The failure mode is rarely the workshop itself. Well-facilitated Event Storming sessions produce valuable insights. The challenge begins after the workshop, when the model has to survive contact with a real production environment.</p><p>The first problem is completeness. Teams rarely forget the happy path. They forget the scenarios that happen occasionally, the provider outage that occurs twice a year, the duplicate import that appears during a synchronization failure, the edge case that only surfaces when multiple systems disagree about the current state of a business process.</p><p>These situations are often understood by experienced engineers and operations teams, but they never make it onto the wall because nobody thinks about them during a workshop. They are considered obvious until a production incident proves otherwise.</p><p>The second problem is consistency. Different groups often use the same words to mean different things. Everyone leaves the workshop believing they agree because the terminology sounds familiar. The disagreement only becomes visible later when systems interact.</p><p>In one domain, a term like &#8220;active listing&#8221; may simply mean visible to customers. In another, it may mean synchronized from an external provider, approved internally, and eligible for marketing distribution. Both teams use the same phrase while describing different realities.</p><p>Event Storming is designed <em>to uncover these misunderstandings</em>, but identifying them depends heavily on the experience of the participants and the facilitator. The third problem is longevity. The model produced during the workshop is documented, shared, updated once or twice, and eventually forgotten. Meanwhile, the business keeps moving.</p><p>New policies appear. Existing workflows evolve. Aggregates grow beyond their original boundaries. Teams split responsibilities across services. Regulatory requirements introduce new constraints. The documented model remains frozen while the actual domain continues to change.</p><p>Eventually, the gap becomes large enough that engineers stop trusting the documentation. At that point, the model serves as historical context rather than an accurate representation of the business.</p><p>These three problems, completeness, consistency, and longevity, are where AI can provide meaningful assistance. Not by replacing domain experts or replacing Event Storming workshops, but by expanding the information available during discovery and helping teams keep their models aligned with reality long after the sticky notes have disappeared.</p><div><hr></div><h2>The MLS Platform: A Running Example</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Daxe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Daxe!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Daxe!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Daxe!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Daxe!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Daxe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1909564,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/200783913?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Daxe!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Daxe!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Daxe!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Daxe!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F449c6d78-fab3-47b8-bfae-0fe45493e368_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>To make the discussion concrete, let&#8217;s use a Multiple Listing Service (MLS) platform as a running example.</p><p>Real estate systems are deceptively complex. At first glance, they appear to be little more than property listings and search functionality. In reality, they contain a dense network of business rules, compliance requirements, third-party integrations, and operational workflows that evolve continuously over time. They are exactly the kind of domain where an Event Storming workshop can produce an elegant model that later struggles to explain what actually happens in production.</p><p>Imagine a team running an Event Storming workshop to model the lifecycle of a property listing. The room contains the right people: product managers, engineers, compliance specialists, and experienced real estate professionals. After two days of discussion, the wall contains a clean sequence of events: PropertyListed, OfferSubmitted, InspectionScheduled, InspectionCompleted, ContractSigned, ClosingScheduled, and DeedTransferred. The model is coherent, well-named, and easy to understand. Everyone leaves the workshop feeling confident that they have captured the domain.</p><p>The problem is that they have mostly captured the story of what normally happens.</p><p>Real production systems spend a surprising amount of time dealing with situations that nobody considers unusual but rarely thinks to mention during a workshop. A listing may be withdrawn temporarily while repairs are completed. An accepted offer may return to negotiation after an inspection reveals structural issues. A property may need to be re-disclosed because of jurisdiction-specific regulations. Multiple providers may report conflicting statuses for the same listing. An agent&#8217;s license may expire in the middle of a transaction. A closing may be delayed because a lender requires an additional appraisal.</p><p>None of these scenarios are particularly rare. Anyone who has spent time working in real estate would recognize them immediately. The challenge is that domain experts naturally describe how the business usually operates. They focus on the primary workflow because that is the easiest story to tell. The exceptions, compliance issues, synchronization failures, and operational headaches that consume engineering time often remain implicit.</p><p>As a result, the workshop output is usually accurate, but incomplete. The model reflects the collective memory of the people in the room, not necessarily the full complexity of the domain. This is where AI begins to create value. Not because it understands the business better than the people participating in the workshop, but because it can systematically explore information that would otherwise remain scattered across codebases, documentation, incident reports, support tickets, and operational history.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>AI as a Domain Discovery Partner</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!v0YY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!v0YY!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!v0YY!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!v0YY!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!v0YY!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!v0YY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1930043,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/200783913?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!v0YY!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!v0YY!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!v0YY!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!v0YY!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5b248ce-408e-44f8-969c-f6fc80afbf10_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The most practical use of AI in Event Storming is not generating domain models automatically. It is increasing the coverage of the discovery process.</p><p>Most organizations are not modeling a greenfield business. They already possess years of accumulated knowledge distributed across source code, support tickets, compliance documents, architectural decision records, operational runbooks, and internal documentation. The problem is that nobody has enough time to review all of this material before a workshop begins, AI can.</p><p>Before an Event Storming session starts, a language model can analyze existing artifacts and surface candidate domain concepts for discussion. It can identify descriptions of state changes, recurring business operations, and patterns that appear repeatedly across documentation and support conversations. More importantly, it can highlight areas where the documented understanding of the domain appears incomplete or inconsistent.</p><p>For an MLS platform, this might involve analyzing support tickets, compliance requirements, and existing services before the workshop begins. Instead of starting with an empty wall, the team arrives with a collection of candidate events that deserve examination. Examples might include ListingExpired, ListingReactivated, AgentLicenseVerificationRequested, ListingComplianceViolationDetected, DualAgencyConsentObtained, or AppraisalWaiverSubmitted. Some suggestions will inevitably be wrong. Others may represent implementation details rather than domain concepts. That is not a problem. The objective is not to generate perfect answers. The objective is to surface possibilities that deserve discussion.</p><p><strong>AI is particularly valuable when searching for missing scenarios and edge cases.</strong> One surprisingly effective exercise is simply asking the model what could go wrong. Given a listing workflow, it will often generate dozens of situations that never surfaced during the workshop: provider synchronization failures, conflicting updates from different systems, missing disclosures, duplicate property records, regulatory violations, or unexpected state transitions. Experienced domain experts usually recognize these situations immediately. The value comes from surfacing them systematically rather than relying on somebody to remember them during a two-day session.</p><p>The same principle applies to commands and policies. Once a set of domain events exists, AI can suggest candidate commands that trigger those events and policies that connect them. It can identify events that appear to lack a triggering intention or business rules that seem to require a downstream response but have not been modeled. These suggestions are not decisions. They are prompts for further exploration.</p><p><strong>Aggregate boundaries</strong> and <strong>bounded contexts</strong> benefit from the same approach. By analyzing how concepts change together, which data is frequently accessed as a unit, and which operations appear to require strong consistency guarantees, AI can propose potential aggregate boundaries. Likewise, it can cluster related concepts and suggest bounded contexts based on language, operational responsibilities, and rates of change. Some of these suggestions will prove useful. Others will be completely wrong. Both outcomes are valuable because they force explicit discussion.</p><p>This distinction is important. The value of AI in domain discovery is not correctness. The value is coverage. A facilitator supported by AI can surface more candidate concepts, more edge cases, more potential boundaries, and more hidden assumptions than a facilitator relying solely on memory and workshop discussion. The domain experts still make every meaningful decision. AI simply increases the likelihood that those decisions are made consciously rather than accidentally.</p><div><hr></div><h2>Where AI Stops and Domain Expertise Begins</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!AZsL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AZsL!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!AZsL!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!AZsL!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AZsL!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AZsL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2088999,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/200783913?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AZsL!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!AZsL!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!AZsL!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AZsL!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ef73e79-b21d-4cef-9d00-3b2ad868e8db_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There is a tempting mistake teams make when introducing AI into domain modeling. They feed documentation, source code, tickets, and diagrams into a language model, receive a list of suggested events, commands, aggregates, and bounded contexts, and begin treating those suggestions as conclusions. That is where trouble starts. AI is very good at finding patterns. It is not particularly good at understanding why those patterns exist.</p><p>A language model can tell you that two operations frequently happen together. It can identify data that is often modified as a unit. It can recognize relationships between concepts that appear repeatedly across code and documentation. What it cannot reliably discover are the business invariants that exist outside the system itself.</p><p>In many organizations, the most important business rules are not fully documented. They exist because of regulatory requirements, contractual obligations, historical decisions, or operational experience accumulated over years. Sometimes the only people who understand them are the domain experts who have lived with the system long enough to know what happens when those rules are violated.</p><p>Consider a real estate platform operating across multiple jurisdictions. A compliance workflow may exist because of a specific regulatory interpretation negotiated years ago. The resulting business rule influences how listings move through the system, but the reasoning behind it may never appear in the codebase. The software reflects the rule. It does not explain the history behind it. An AI model analyzing that system can observe behavior. It cannot reliably explain intent.</p><p>This is why AI should never be viewed as a replacement for domain expertise. The most effective workflow is one where AI generates candidates and domain experts make decisions. The model surfaces potential events, policies, edge cases, aggregates, and context boundaries. The experts review them, reject some, refine others, and occasionally discover something important that had previously gone unnoticed.</p><p>In practice, this changes the nature of the work. Instead of spending hours trying to remember every possible exception, participants spend their time evaluating a prepared set of candidates. The discussion becomes more focused because the room is reacting to concrete suggestions rather than attempting to generate everything from scratch.</p><blockquote><p>AI also has a tendency to produce technically correct but business-poor language.</p></blockquote><p>Given enough context, it will happily suggest events such as <code>StatusChanged</code>, <code>RecordUpdated</code>, or <code>DataModified</code>. These names may describe what happened from a technical perspective, but they say very little about the business itself.</p><p>Domain experts naturally push the language toward something more meaningful. A generic <code>StatusChanged</code> may become <code>ListingWithdrawnPendingSellerRepairs</code>. A <code>RecordUpdated</code> may become <code>OfferRejectedAfterInspectionFindings</code>. The difference is subtle but important. One describes a database operation. The other describes something the business actually cares about. That distinction remains a human responsibility.</p><p>The most useful way to think about AI in Event Storming is not as a facilitator or a co-designer. It is closer to a research assistant that has read every available artifact and can report back with observations. It can surface patterns, identify inconsistencies, and generate possibilities. What it cannot do is determine which possibilities matter.</p><p>That responsibility remains exactly where it has always belonged: with the people who understand the domain.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/event-storming-at-scale-why-ai-changes?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/event-storming-at-scale-why-ai-changes?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/event-storming-at-scale-why-ai-changes?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2>Turning Event Storming into Living Documentation</h2><p>Even when an Event Storming workshop is successful, another challenge appears almost immediately. Keeping the model current.</p><p>Most teams invest significant effort into creating a domain model and very little effort into maintaining one. The workshop ends, the diagrams are documented, and everyone agrees they should be kept up to date. Then feature work resumes, priorities shift, deadlines appear, and model maintenance quietly falls to the bottom of the backlog. Six months later the model still exists, but nobody fully trusts it. This is where AI may ultimately provide more value than it does during the workshop itself.</p><p>Instead of treating the domain model as a static artifact, teams can use AI to continuously compare the documented model against signals coming from the actual system. Support tickets, incident reports, event logs, new features, compliance changes, and operational documentation all contain information about how the domain is evolving. The goal is not automatic model updates. The goal is drift detection.</p><p>Imagine an MLS platform where a new category of support request begins appearing repeatedly. Sellers are asking what happens when an agent loses their license during an active transaction. The issue appears often enough that support teams have developed informal procedures to handle it, but the scenario does not exist anywhere in the domain model. Traditionally, this gap might remain invisible for months.</p><p>An AI-assisted review process could identify the pattern much earlier. The model notices a recurring workflow appearing in support tickets and operational conversations that has no corresponding representation in the domain model. It flags the discrepancy for review. The important part is what happens next.</p><p>A domain expert validates whether the scenario is legitimate. Engineers determine how the system currently behaves. The team discusses whether new events, commands, policies, or aggregates are required. The model is updated deliberately rather than accidentally.</p><p>This is fundamentally different from the approach most teams use today, which is simply allowing the model to become outdated until nobody relies on it anymore.</p><p>AI does not solve the maintenance problem by itself. What it does is make maintenance practical. Instead of asking teams to periodically re-examine an entire domain, it highlights specific areas where reality and documentation appear to be diverging. That is a much easier problem to solve.</p><p>Over time, the result is a domain model that evolves alongside the business rather than falling behind it. Changes become explicit. New concepts are introduced intentionally. The language shared between engineers and domain experts remains aligned with the reality of the system. That outcome has always been one of the central goals of <a href="/__u/nidly.substack.com/p/domain-driven-design-in-the-ai-era">Domain-Driven Design</a>.</p><p>Event Storming remains one of the best tools available for creating that shared language.</p><blockquote><p>AI simply makes it easier to keep that language alive.</p></blockquote><div><hr></div><h2>Conclusion</h2><p>The biggest challenge with Event Storming has never been the workshop itself. Most workshops succeed. The real challenge begins after the workshop ends, when the domain starts changing and the model slowly drifts away from reality.</p><p>Historically, teams have accepted that drift as inevitable. The workshop produces a snapshot of the business, and over time that snapshot becomes less accurate. New workflows emerge. Edge cases accumulate. Business rules evolve. The documentation remains frozen while the domain keeps moving. AI changes that dynamic.</p><p>Not because it understands the business better than the people responsible for it, but because it can process far more information than any individual participant could reasonably review. It can identify patterns across tickets, documentation, code, support conversations, and operational history. It can surface candidate events, uncover missing scenarios, and highlight areas where the documented model no longer reflects the behavior of the system.</p><p>What it cannot do is make domain decisions. It cannot determine which business invariants matter. It cannot explain why a rule exists. It cannot replace the judgment of a domain expert who understands the consequences of getting the model wrong.</p><p>The most effective approach is therefore not AI-driven modeling. It is AI-assisted discovery. The workshop still matters. The conversations still matter. The sticky notes still matter. Domain experts remain responsible for every meaningful decision. The difference is that the team is no longer relying entirely on memory, intuition, and whatever happened to come up during a two-day session.</p><p>The wall still represents the domain. Now there is simply something helping ensure that the wall continues to reflect reality long after the workshop is over.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Fine-Tuning Is the New Microservices Mistake]]></title><description><![CDATA[I&#8217;ve Seen This Movie Before]]></description><link>https://nidly.substack.com/p/fine-tuning-is-the-new-microservices</link><guid isPermaLink="false">https://nidly.substack.com/p/fine-tuning-is-the-new-microservices</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 15 Jun 2026 13:20:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!lUoT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!lUoT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!lUoT!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!lUoT!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!lUoT!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lUoT!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!lUoT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1627925,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/195474170?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!lUoT!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!lUoT!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!lUoT!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lUoT!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd484128f-6e4b-4a91-93bf-68bc4f89dc53_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1>I&#8217;ve Seen This Movie Before</h1><p>Software engineering has a habit of turning good ideas into default answers.</p><p>The pattern is familiar. A handful of highly capable companies solve a difficult problem with a particular technique. The results are impressive. Other teams start paying attention. Adoption spreads. At first, people apply the technique carefully and for the reasons it was designed. Then something changes. The technique stops being a solution and becomes an expectation. Teams use it not because the problem demands it, but because it has become part of the industry&#8217;s definition of &#8220;doing things properly.&#8221;</p><blockquote><p>We saw this happen with microservices.</p></blockquote><p>A small number of companies were operating at a scale where deployment independence, organizational autonomy, and distributed ownership were genuine problems. Microservices helped solve those problems. The benefits were real. Then the idea escaped its original context.</p><p>Before long, companies with a handful of engineers and modest traffic were splitting systems into dozens of services because that was what modern architecture was supposed to look like. Something similar is happening today with fine-tuning.</p><p>To be clear, fine-tuning is not a mistake. Neither were microservices. Both are powerful techniques with legitimate use cases. The problem begins when a technique becomes the answer before the question has been fully understood.</p><p>That is why, when a discussion about fine-tuning appears for the third time in a planning meeting, the most important question is usually not whether the model can be fine-tuned. The more useful question is whether the team is solving a genuine model problem, or simply repeating a pattern the industry has repeated many times before.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What Actually Happened</h2><p>What happened next was not a failure of microservices. It was a failure to account for the cost of success. The benefits were obvious from the beginning. The costs were not.</p><p>A company would split a monolith into a handful of services and see real improvements. Teams moved faster. Deployments became less risky. Ownership became clearer. Encouraged by those results, they would split a few more services. Then a few more. Over time, what started as a sensible architectural decision became an architectural reflex.</p><p>A system with five services became a system with fifty. A system with fifty became a system with two hundred. Individually, each new service seemed justified. Collectively, they created a level of complexity that nobody had planned for.</p><p>Understanding a user-facing failure no longer meant reading a log file and tracing a stack trace. It meant reconstructing a distributed conversation across multiple services, each with its own logs, metrics, deployment history, and operational quirks. Even coordination found its way back into the system.</p><p>The original promise was that teams would no longer need to coordinate deployments. Technically, that was true. In practice, coordination simply moved to a different layer. Teams now coordinated API contracts, compatibility guarantees, schema migrations, rollout sequences, and cross-service dependencies.</p><p>The complexity arrived gradually enough that most organizations did not notice it happening. Each additional service made sense in isolation. The overall system did not. Many teams eventually realized they had spent years optimizing for architectural ideals such as service purity, technology flexibility, and perfect boundary separation while investing far less attention in the outcomes that actually mattered to the business.</p><p>Some organizations responded by consolidating services. Others quietly rebuilt parts of their systems into modular monoliths. Many are still paying down the operational debt created during that period. The lesson was not that microservices were wrong. The lesson was that powerful techniques become dangerous when they turn into defaults.</p><div><hr></div><h2>The New Obsession: Fine-Tuning Everything</h2><p>The AI industry appears to be moving down a remarkably similar path. The difference is speed. Microservices took years to spread through the industry. AI trends often spread in months. A team launches an AI-powered feature. Maybe it is a customer support assistant. Maybe it helps lawyers review contracts. Maybe it answers questions over internal documentation.</p><p>The first results are promising but imperfect. The model occasionally misses domain-specific terminology. Some responses are formatted incorrectly. The tone feels slightly off. Certain answers lack the context that experienced employees would naturally have.</p><p>At some point, usually in a planning meeting, someone proposes the obvious solution. &#8220;We should fine-tune the model on our data.&#8221; The suggestion sounds reasonable because it often is. After all, the model does not speak exactly like the organization. It does not understand every internal acronym. It has never seen the company&#8217;s documents, processes, or historical knowledge.</p><p>Fine-tuning appears to be the direct path from generic intelligence to domain expertise. That is why the idea spreads so easily. Legal teams want models trained on contract language. Financial organizations want models that understand internal reporting terminology. Sales platforms want assistants that sound like their top performers. Enterprise software vendors want chatbots that know the details of internal workflows and operational procedures.</p><p>Viewed individually, these are perfectly reasonable requests. That is precisely what makes the pattern dangerous. Engineering trends rarely begin with obviously bad decisions. They begin with dozens of decisions that look completely sensible when considered one at a time.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/fine-tuning-is-the-new-microservices?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/fine-tuning-is-the-new-microservices?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/fine-tuning-is-the-new-microservices?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2>Most AI Failures Are Not Model Failures</h2><p>One of the most valuable lessons I have learned while building AI systems is that production failures are rarely caused by the model itself. When an AI feature starts behaving unexpectedly, the instinctive reaction is often to blame the model. In practice, the root cause is frequently somewhere else in the system.</p><p>Retrieval failures are a common example. A team observes incorrect answers and concludes that the model is hallucinating, when in reality the model never received the information required to answer correctly. The retrieval layer may be returning documents that are superficially related to the query but semantically irrelevant. The chunking strategy may have separated information that needed to remain together. Ranking logic may be prioritizing keyword similarity over actual usefulness. From the model&#8217;s perspective, the answer was never available in the first place.</p><p>Stale embeddings create another class of failures that often go unnoticed until users begin reporting problems. Documentation evolves, internal processes change, and knowledge bases are updated, yet the vector index remains untouched. Months later, the retrieval system confidently serves information that was accurate in the past but no longer reflects reality. The symptom looks like model confusion. The underlying cause is an operational failure in the data pipeline.</p><p>As systems become more agentic, orchestration failures become increasingly important. Tool calls time out, retry policies are missing, external services return unexpected responses, and memory layers accumulate inconsistent state. From a user&#8217;s perspective the AI appears unreliable, even though the model itself may be functioning exactly as intended. The failure exists in the workflow coordinating the model rather than in the model itself.</p><p>Evaluation failures are perhaps the most dangerous because they remain invisible for long periods of time. Many teams invest heavily in training infrastructure while treating evaluation as a final checkpoint rather than a permanent capability. As a result, regressions are discovered through support tickets and customer complaints instead of systematic monitoring. By the time the issue becomes visible internally, users have often been experiencing it for weeks.</p><p>What makes these examples interesting is that they are fundamentally software problems. The model is merely the most visible component in the system. The corrective action is usually architectural rather than parametric. Improving retrieval quality, fixing data pipelines, strengthening orchestration logic, and building robust evaluation systems often produces larger gains than another round of fine-tuning because these changes address the actual source of failure rather than the symptom.</p><div><hr></div><h2>The Monolith Analogy</h2><p>The microservices era taught an important lesson that is easy to forget. Simplicity is not the opposite of sophistication. A well-designed monolith can contain clear module boundaries, strong testing practices, and disciplined engineering processes while remaining significantly easier to operate than a highly distributed architecture.</p><p>The argument for a monolith was never that decomposition is bad. The argument was that decomposition should only happen when the benefits clearly outweigh the operational costs. Many organizations learned this lesson after spending years managing service boundaries, deployment coordination, observability challenges, and distributed debugging for systems that never truly required that level of complexity.</p><p>A strong foundation model combined with high-quality retrieval, purpose-built tool integrations, persistent memory, and a robust evaluation framework follows the same philosophy. This is not a simplistic solution. Building such a system well requires considerable engineering effort. The difference is that the complexity remains concentrated in the architecture surrounding the model rather than being distributed across an ever-growing collection of specialized models.</p><p>This approach also tends to age better. When new information becomes available, the retrieval layer can be updated. When behavior changes, evaluation systems can detect it. When a stronger base model appears, upgrading becomes far more manageable because business knowledge has not been embedded deeply into model weights. The organization retains flexibility while avoiding much of the operational burden that comes from maintaining multiple fine-tuned models.</p><p>That is why the monolith analogy matters. The claim is not that fine-tuning should be avoided. The claim is that many teams are attempting to solve architectural problems through model customization. Just as the industry eventually realized that not every application needed dozens of services, it may eventually discover that not every AI system needs dozens of models.</p><div><hr></div><h2>When Fine-Tuning Actually Makes Sense</h2><p>None of this is an argument against fine-tuning itself. Like microservices, fine-tuning exists because it solves real problems, and there are situations where it is absolutely the right tool.</p><p>One of the strongest use cases is style preservation. There are environments where prompting alone cannot reliably reproduce the tone, structure, and conventions an organization requires. A legal firm, for example, may need AI-generated documents to follow established drafting practices, handle defined terms consistently, and maintain a specific style across thousands of outputs. In those situations, fine-tuning on historical work product can provide a level of consistency that is difficult to achieve through retrieval and prompting alone.</p><p>Classification workloads are another area where fine-tuning often makes sense. When the label space is narrow, the taxonomy is stable, and latency requirements are strict, a fine-tuned model can outperform more general approaches while reducing inference costs and prompt complexity. A financial institution classifying transaction descriptions into internal categories is solving a very different problem from a general-purpose enterprise assistant, and the economics of that decision can be compelling.</p><p>There are also domains where the language and knowledge differ substantially from what foundation models typically encounter. Specialized manufacturing documentation, proprietary internal systems, scientific notation, or highly technical research data may justify fine-tuning when the available dataset is sufficiently large and the expected behavioral improvements are well understood.</p><p>What these examples have in common is specificity. The requirements are narrow, the success criteria are measurable, and the operational investment is proportional to the expected return. They are fundamentally different from the much more common situation where an enterprise chatbot is underperforming and the team simply wants the most technically ambitious solution available.</p><div><hr></div><h2>A Practical Heuristic</h2><p>Before committing to a fine-tuning initiative, it is worth working through a sequence of questions that often reveals where the real constraint exists.</p><p>The first question should be about retrieval. Is the model actually receiving the information it needs to answer correctly? Many apparent reasoning failures turn out to be retrieval failures. Documents are missing, ranking is poor, embeddings are outdated, or the retrieved context is only loosely related to the user&#8217;s intent. Measuring retrieval precision and recall against a labeled dataset is not glamorous work, but it frequently surfaces the actual bottleneck.</p><p>The second question concerns context engineering. Even when the correct information is available, the way it is assembled matters. Context organization, instruction design, prompt structure, and context-window utilization can significantly influence output quality. These variables are often faster, cheaper, and easier to iterate on than training pipelines, yet they are frequently examined only after fine-tuning discussions have already begun.</p><p>The third question involves workflow design. Modern AI systems increasingly depend on tool calls, memory layers, retrieval pipelines, and multi-step orchestration. When performance degrades, it is important to determine whether the model is actually failing or whether the surrounding workflow is introducing incorrect intermediate state, incomplete context, or unreliable tool outputs. In many cases, the model is simply the last component in a chain of failures that originated elsewhere.</p><p>Perhaps most importantly, teams should invest in evaluation infrastructure before they invest in training infrastructure. Strong evaluation systems make it possible to improve retrieval, prompting, workflow design, and model behavior with confidence because they provide a reliable feedback loop. Teams that move directly to fine-tuning often discover they are optimizing for benchmark improvements that have little relationship to production performance.</p><p>If, after working through this sequence, the limiting factor is genuinely the model&#8217;s baseline behavior or parametric knowledge, then fine-tuning becomes a reasonable next step. What surprises many teams is how often the problem is resolved before they ever reach that conclusion.</p><div><hr></div><h2>Complexity Must Be Earned</h2><p>The lesson from the microservices era was never that distributed systems are bad. The lesson was that architectural complexity should be earned. The organizations that benefited most from microservices were not simply the ones that adopted them. They were the ones that possessed the operational maturity required to support them. Observability, deployment automation, testing discipline, ownership boundaries, and operational rigor were what made the architecture successful.</p><p>Organizations that struggled often made a different mistake. They adopted the architecture first and attempted to develop the supporting capabilities later. As a result, they inherited the costs of decomposition long before they realized its benefits.</p><p>The same pattern is beginning to emerge in AI systems. Fine-tuning is a valuable technique and, in the right context, it can produce improvements that are difficult to achieve through any other approach. The problem is not the technique itself. The problem is allowing it to become the default response whenever an AI system fails to meet expectations.</p><p>When that happens, organizations gradually accumulate model sprawl in much the same way previous generations accumulated service sprawl. Each individual decision appears reasonable. A specialized model for legal workflows seems justified. A dedicated model for customer support appears useful. A separate model for internal knowledge management sounds sensible. Over time, however, these decisions create an expanding operational burden. Datasets must be maintained, evaluations must be expanded, deployments must be managed, and governance requirements must be satisfied across an increasingly complex landscape.</p><p>The engineers who learned the most from the microservices era were not the ones who rejected distributed systems. They were the ones who learned to distinguish between complexity that creates leverage and complexity that merely creates maintenance overhead. They understood that decomposition is valuable when it solves a real problem and expensive when it is adopted primarily because it represents the architectural trend of the moment.</p><p>AI teams are approaching a similar decision point today. The important question is not whether a model can be fine-tuned. The important question is whether the problem in front of us genuinely requires it. If retrieval quality is poor, evaluation is weak, orchestration is unreliable, or the surrounding architecture is introducing failure, fine-tuning may simply be treating symptoms rather than causes. The organizations that navigate this transition successfully will not be the ones that train the most models. They will be the ones that understand when additional complexity is justified and when it is not.</p>]]></content:encoded></item><item><title><![CDATA[Modern Software Architecture (Part 1): Boundaries, Trade-Offs, and Coordination]]></title><description><![CDATA[A practical guide to understanding cohesion, coupling, architectural characteristics, bounded contexts, connascence, and the real coordination costs behind modern distributed systems.]]></description><link>https://nidly.substack.com/p/architecting-software-for-the-ai</link><guid isPermaLink="false">https://nidly.substack.com/p/architecting-software-for-the-ai</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Sun, 07 Jun 2026 11:17:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!jI5c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!jI5c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!jI5c!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!jI5c!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!jI5c!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!jI5c!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!jI5c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg" width="1024" height="682" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:682,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:158091,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/196940133?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!jI5c!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!jI5c!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!jI5c!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!jI5c!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d231c2d-cf7a-49b6-9e4e-226b6154709e_1024x682.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There is a specific kind of meeting that happens at every company once it reaches a certain scale. The engineering team is gathered, the slides are prepared, and someone from the platform team is about to explain why the system that worked perfectly fine six months ago is now responsible for most production incidents, the slowest deploys, and the longest on-call rotations. The word &#8220;microservices&#8221; will appear in the presentation. So will &#8220;event-driven.&#8221; The argument will be technically coherent and organizationally naive, and the migration it proposes will cost eighteen months and leave the system in a shape that is worse in different ways than the shape it was in before.</p><p>I have been in that meeting. I have run that meeting. And years later, I have sat across the table from engineers who survived the migration I proposed and watched them explain, patiently, what we got wrong.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>This is an article about architecture. Not the whiteboard kind where you draw boxes and arrows until the diagram looks sufficiently serious. The real kind, where every decision trades one class of problem for another, where the organization shapes the system whether you acknowledge it or not, and where AI-assisted development has dramatically accelerated how quickly teams accumulate complexity without changing the underlying physics of distributed systems.</p><div><hr></div><h2>The Illusion of Technology-First Architecture</h2><p>When North Harbor Realty, a mid-sized real estate brokerage operating across Chicago and the surrounding suburbs, started building its internal platform, the initial architecture looked like most initial architectures: a monolithic Rails application backed by a PostgreSQL database, a couple of background job queues, and an S3 bucket full of listing images. It worked. It handled MLS ingestion from the three primary providers they were connected to, stored canonical listing data, surfaced agent profiles, and powered the search page that most of their buyer traffic hit every morning between eight and eleven.</p><p>The problems started when the business began growing in directions the original architecture had never anticipated. They onboarded two additional MLS providers, each with its own data format, feed schedule, and interpretation of what constituted a unique listing. They added an image processing pipeline after the product team became convinced that automated quality scoring and virtual staging previews would improve buyer conversion. They contracted with an AI recommendation vendor whose models required near-realtime listing data delivered in a very specific schema.  They hired thirty new agents in eighteen months, which put pressure on the lead routing and notification systems in ways nobody had modeled.</p><p>The response to that complexity was predictable. Someone drew a services diagram. The diagram showed eight boxes. Each box was a microservice. The argument was that by decomposing the monolith along functional lines, each team could deploy independently, scale independently, and evolve its part of the system without coordinating every release across the organization.</p><blockquote><p>The argument was not wrong. But it was incomplete in the ways that matter most.</p></blockquote><p>The engineers who built that migration learned something you can only really learn by building it: <strong>technology does not determine architecture. Business behavior determines architecture. </strong>The question is rarely whether something should become a service or remain a module. The real question is what coordination costs each option introduces, and whether the organization is actually capable of paying for them.</p><div><hr></div><h2>Architecture as Trade-Off Management</h2><blockquote><p>Architecture is the management of irreversible trade-offs.</p></blockquote><p>Neal Ford has a formulation I return to often. <strong>Architecture is not the search for the best solution. It is the search for the least worst solution.</strong> Every architectural decision trades one failure mode for another, one form of complexity for another, one kind of operational pain for another kind. The job is to understand which pain you can manage and which pain will kill you.</p><p>At North Harbor, the listing deduplication problem made this concrete. When a property is listed by multiple agents through multiple MLS providers, you end up with multiple listing records representing the same physical property. These records have different IDs, slightly different addresses, different photo sets, and different pricing histories. Deduplication is expensive. It requires fuzzy matching on address normalization, geospatial proximity checks, and machine-learned similarity scoring across listing attributes. The question was where this computation should live.</p><p>The technology-first answer would have produced a deduplication microservice with its own database, deploy cycle, and ownership boundary. The architectural answer required asking a <em>different</em> set of questions. <em>What does the deduplication result affect? </em>It affects the canonical listing record, which is the source of truth for search indexing, AI recommendations, agent commissions, and the legal history of what was listed and when. What happens when deduplication is wrong? Buyers see duplicate listings. Agents receive incorrect commission calculations. MLS reconciliation jobs fail. Some of those failures are operationally annoying. Others are legally dangerous.</p><p>Given those consequences, the consistency requirements for deduplication were extremely high. You could not tolerate a state where the canonical listing system and the deduplication system had diverged. Eventual consistency is a perfectly <strong>reasonable trade-off in systems</strong> where temporary divergence is acceptable. It was not acceptable here.</p><p>The decision was to keep deduplication logic inside the canonical listing module rather than extracting it into a separate service. This was not a failure of vision or a lack of ambition. It was an accurate reading of what the system actually needed. Deduplication needed to participate in the same transaction as the canonical listing write. Making that work across a network boundary would have required distributed transactions, which are expensive to implement correctly and notoriously difficult to operate. The coordination cost of the &#8220;correct&#8221; microservice architecture was higher than the system could afford.</p><p>This is what<strong> architecture as trade-off management</strong> looks like in practice. Not philosophy. Not diagram aesthetics. Cost accounting applied to system design.</p><div><hr></div><h2>Cohesion and Coupling Beyond Textbook Definitions</h2><p>Every architecture course covers cohesion and coupling. The textbook definition is simple: </p><blockquote><p>high cohesion means things that change together stay together, while low coupling means things that do not need to change together are separated.</p></blockquote><p> The definition is correct and <strong>almost useless in production</strong> because it does not tell you how to recognize these properties operationally or what to do when they conflict.</p><p>The operational version of cohesion looks different. Code is cohesive when the things inside it share the<strong> same reason to change</strong>, the same deployment cadence, the same consistency requirements, and usually the same operational pressures.</p><p>If one part of a module changes every time an MLS provider modifies its feed format, while another changes every time the product team redesigns how listings appear on the website, those pieces of code may live together physically while having almost nothing in common operationally.</p><p>At North Harbor, the provider normalization layer and the canonical listing store initially lived in the same codebase and shared the same deployment pipeline. When the normalization logic needed an urgent fix because one of the MLS providers changed its latitude and longitude encoding format, the deploy also touched the canonical listing system, which triggered the full regression suite. The regression suite took forty minutes. That meant a forty-minute delay on a fix affecting listing accuracy for the live Chicago market during peak morning traffic. <strong>That pain was a cohesion signal.</strong></p><p>The provider normalization code had a completely different deployment urgency profile than the canonical listing code. It changed more frequently, for different reasons, driven by external provider behavior rather than internal product decisions. Separating those concerns was not originally a microservice decision. <strong>It was a cohesion decision expressed first as a module boundary </strong>and eventually as its own deployable component once the operational cost of shared deployment became impossible to ignore.</p><p>Coupling is the inverse problem. The textbook says low coupling is good. In practice, every non-trivial system is coupled in some direction, and the architectural question is not whether coupling exists but where it exists and what form it takes.</p><p><em>Some forms of coupling are manageable.</em> Passing structured data between components through a stable contract is usually safe. Other forms are far more expensive. Temporal coupling, where one system must be available for another system to complete its work, turns small failures into cascading outages. Deployment coupling is even worse because teams often do not notice it until a harmless-looking release breaks production.</p><p>The notification system at North Harbor carried invisible deployment coupling for most of its early life. Notifications were triggered by events in the listing lifecycle: new listings matching a buyer&#8217;s saved search, price reductions on watched properties, and status transitions from active to under contract. These events were defined as constants inside the listing module and consumed directly by the notification system.</p><p>When the product team introduced a new listing status called &#8220;coming soon&#8221; and changed how status transitions were modeled internally, the notification system broke because it was coupled to the internal representation of listing status rather than to a stable contract.</p><p>This is where <strong>connascence</strong> becomes a more useful concept than coupling. The real problem was not simply that the systems depended on each other. The problem was that they depended on the same meaning, the same internal representation, and the same assumptions evolving in lockstep across deployment boundaries.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/architecting-software-for-the-ai?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/architecting-software-for-the-ai?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/architecting-software-for-the-ai?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2>Connascence and the Cost of Coordination</h2><p>Connascence is a term from Meilir Page-Jones that most engineers encounter once and promptly forget. It deserves more attention because it gives you a more precise language for describing the kinds of coupling that create the most operational pain.</p><p><strong>Connascence describes situations where two parts of a system must evolve together to remain correct</strong>. The important distinction is not whether connascence exists. Every non-trivial system has it somewhere. The real question is what kind of connascence you are creating and how expensive it becomes once it crosses team, deployment, or network boundaries.</p><p>Some forms are relatively harmless. Agreeing on a field name is manageable. Sharing the meaning of internal state transitions is far more dangerous. The most expensive form is usually temporal connascence, where one system must execute within a particular order or time window for another system to function correctly.</p><p>The notification system problem at North Harbor was a form of connascence of meaning. The notification code and the listing code shared an interpretation of what listing statuses meant and how those statuses transitioned. When that meaning changed, both components needed to change together, but nothing in the system explicitly enforced this requirement. It was an implicit contract made visible only when it broke.</p><p>The fix was not simply to rearrange code. The fix was to replace connascence of meaning with a weaker form of connascence by introducing a stable event contract. The listing module now publishes <code>ListingStatusChanged</code> events using a documented and versioned schema. The notification system consumes those events without knowing anything about the internal representation of listing status. The listing system is free to evolve its internal model as long as the published contract remains stable.</p><p>This is the real architectural value of event-driven patterns. Not that they are inherently more scalable or more modern. It is that they allow you to convert strong forms of connascence into weaker forms by placing an explicit contract at the boundary. The contract becomes visible, versioned, testable, and operationally observable. It turns an implicit coordination requirement into an explicit one, which is almost always an improvement.</p><p>The image processing pipeline exposed a different problem: <strong>connascence of timing</strong>. Image processing at North Harbor is a genuinely heavy workload. When a new listing is ingested, there may be thirty to fifty raw images from the agent, each requiring resizing into multiple formats, quality scoring, content policy checks, and processing through a virtual staging model that digitally furnishes empty rooms. On a good day, the complete pipeline for a large listing takes eight minutes. On a bad day, when the GPU cluster is already saturated by recommendation workloads, it can take forty.</p><p>The original architecture coupled image processing synchronously to the listing ingestion flow. When an agent submitted a listing, the system processed the images before returning a response. The result was a dangerous form of temporal connascence between listing ingestion and image processing capacity. If image processing slowed down, listing ingestion slowed down. If the GPU cluster failed, listing ingestion failed. Two unrelated operational concerns had effectively been welded together by a timing dependency.</p><p>Decoupling the image pipeline into an asynchronous queue was not particularly interesting from a technical perspective. Message queues are not novel technology. What mattered was acknowledging why the synchronous coupling existed in the first place. The product team wanted listings to appear in search results with images immediately after publication.</p><p>The asynchronous design forced the organization to confront the trade-off honestly. Listings would sometimes appear in search with a &#8220;photos processing&#8221; indicator for up to forty minutes, and ranking quality would temporarily degrade because image quality scores were not yet available. This was not just a technical compromise. It was a product decision about what kind of operational dependency the business was willing to tolerate.</p><p>The architecture only moved forward once the coordination cost of synchronous coupling was explained in business terms rather than technical ones.</p><div><hr></div><h2>Architectural Characteristics and Business Pressure</h2><p>There is a framework from Mark Richards and Neal Ford in <em><a href="https://www.oreilly.com/library/view/fundamentals-of-software/9781492043447/">Fundamentals of Software Architecture</a></em><a href="https://www.oreilly.com/library/view/fundamentals-of-software/9781492043447/"> </a>that I have found more useful than almost any architecture diagram I have seen in a review meeting:<strong> architectural characteristics</strong>. <em>These are the non-functional properties a system must exhibit in order to satisfy the business.</em> They are specific enough to measure, concrete enough to shape architecture, and operational enough to expose trade-offs that diagrams usually hide.</p><p>At North Harbor, the characteristics became obvious once you stopped looking at the system as a collection of features and started looking at it as a collection of business pressures.</p><p>The canonical listing store needed consistency above almost everything else. A stale listing, a duplicate listing, or an incorrect status was not simply a user experience problem. Agents had legal obligations around listing accuracy. Buyers making offers based on incorrect data created liability for the brokerage. The consistency requirement was a business and legal requirement first, and only secondarily a technical one.</p><p>The search infrastructure had a completely different set of pressures. It needed availability and low latency. The Chicago residential market has strong daily traffic patterns. Search traffic spikes sharply between eight and eleven in the morning, dips through the afternoon, and rises again between six and nine in the evening when people browse listings from home. A/B testing showed that search latency above three hundred milliseconds produced measurable conversion decline.</p><p>At the same time, buyers still expected search to work even when the canonical listing system was running large reconciliation jobs or undergoing maintenance. Nobody browsing condos in Lincoln Park at 8:30 in the morning cared that an internal consistency process was saturating the primary database.</p><p>The lead routing system had yet another profile. It needed throughput and resilience. When a buyer submits an inquiry, the brokerage has a narrow window to match that lead to the right agent before the buyer loses interest or contacts a competing brokerage. The routing logic itself is complicated. It incorporates geographic assignment, agent specialization, contractual lead distribution agreements, and load balancing rules designed to prevent top-performing agents from becoming overloaded.</p><p>Putting all of that logic directly in a latency-sensitive execution path meant that small slowdowns immediately became lost leads.</p><p>These differences are the real reason service boundaries emerged at North Harbor. Not because the organization wanted microservices as a statement of technical maturity, but because the systems were being pulled toward incompatible operational requirements.</p><p>The canonical listing system optimized for <strong>correctness</strong> and <strong>consistency</strong>. The search infrastructure optimized for <strong>latency</strong> and <strong>availability</strong>. The lead routing pipeline optimized for throughput and resilience.</p><p>Running those concerns inside the same deployment would have forced damaging compromises. Consistency requirements would push toward locking strategies that harmed search latency. Search availability requirements would encourage aggressive caching strategies that undermined canonical correctness. Lead routing throughput requirements would force time-sensitive assignment workloads to compete for the same infrastructure resources as reconciliation jobs and listing updates.</p><p><strong>Characteristics-driven architecture is different from function-driven architecture.</strong> You are not decomposing systems only according to what they do. You are decomposing systems according to the operational realities they must survive.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Strategic DDD and Bounded Contexts</h2><p>Domain-Driven Design has accumulated enough cargo cult adoption over the years that it is worth being precise about which parts are genuinely useful and which parts mostly exist to make architecture workshops sound sophisticated.</p><p>The useful part is the concept of bounded contexts and the linguistic discipline they impose. A bounded context is a boundary within which a model remains internally consistent and complete. Inside that boundary, words mean what they mean without ambiguity. Outside the boundary, the same word may legitimately mean something different, and that difference is not academic. It has operational consequences. At North Harbor, the word &#8220;listing&#8221; was the clearest example.</p><p>Inside the MLS ingestion context, a listing was a raw provider feed record with a provider ID, ingestion metadata, and a set of fields that varied depending on which MLS system produced it. Inside the canonical listing context, a listing was something very different: a normalized, de-duplicated representation of a property with a North Harbor identifier, a confidence score, and a lifecycle state representing the brokerage&#8217;s best understanding of the property&#8217;s actual market status.</p><p>Inside the search context, a listing was not really a transactional entity at all. It was a retrieval document optimized for speed. It contained denormalized agent information, precomputed geographic clustering, search-oriented text fields, and ranking metadata designed to support low-latency queries during peak traffic hours.</p><p>Inside the commission accounting context, a listing became a financial object tied to transaction prices, commission splits, contractual obligations, and payout calculations. These were not four views of the same model. They were four genuinely different models of the same real-world concept.</p><p>Trying to force all of them into a single enterprise-wide &#8220;Listing&#8221; model would have created one of two outcomes. Either the model would become so abstract that it would serve none of the contexts particularly well, or it would become so overloaded with competing concerns that changes required by one part of the organization would constantly destabilize the others.</p><p>The bounded context discipline rejects both options. Each context gets its own model, its own language, and often its own storage model optimized for its operational needs. Translation happens explicitly at the boundary in code that is visible, testable, and operationally observable.</p><p>At North Harbor, that translation was handled through anti-corruption layers between the MLS ingestion context and the canonical listing context. Provider feed records were transformed into canonical listing representations through explicit mapping rules. Provider-specific fields were normalized, ambiguities were resolved according to documented business rules, and every transformation decision was logged for auditing and reconciliation purposes.</p><p>This mattered because providers routinely changed their schemas without notice. During the first year alone, multiple ingestion failures were traced back to silent feed changes from MLS vendors. The anti-corruption layer isolated those failures to the translation boundary rather than allowing corrupted provider data to leak directly into the canonical store.</p><p>Event storming, the workshop technique created by Alberto Brandolini, also turned out to be genuinely useful at North Harbor, though not for the reasons architecture diagrams usually suggest.</p><p>The value was not the sticky notes. The value was visibility. Putting domain events on a wall and asking &#8220;what caused this event?&#8221; and &#8220;what reacts to this event?&#8221; exposed workflow boundaries that the original technology-first decomposition never revealed. It became obvious, for example, that listing lifecycle events formed the backbone of almost every major workflow in the organization.</p><p><code>ListingCreated</code><br><code>ListingActivated</code><br><code>ListingModified</code><br><code>ListingUnderContract</code><br><code>ListingSold</code><br><code>ListingExpired</code></p><p>Search indexing reacted to them. Notifications reacted to them. Commission calculations depended on them. Agent performance reporting derived metrics from them. Reconciliation jobs used them to verify listing state transitions across providers.</p><p>The listing lifecycle was not merely a database field. It was a first-class architectural concept around which multiple operational workflows coordinated themselves.</p><p>That distinction mattered because it changed how the system boundaries were designed. The architecture stopped being organized around database tables and started being organized around business behavior and coordination patterns.</p><div><hr></div><h2>Architectural Quanta and Granularity</h2><p>The architectural quantum, a term Mark Richards and Neal Ford use to describe the smallest deployable unit with high functional cohesion and all the structural pieces required to operate independently, is one of the more useful ways to think about architectural granularity.</p><p>The important question is not whether a system should use microservices or remain a monolith. The important question is what level of granularity each part of the system actually requires given its operational characteristics, change frequency, ownership structure, and coordination cost. At North Harbor, different parts of the platform evolved toward different granularities for entirely different reasons.</p><p>The canonical listing store, commission accounting system, and agent profile management system remained modules inside a shared runtime for the first three years of operation. They shared a database, deployed together, and were maintained by heavily overlapping teams. Their change rates were similar, their consistency requirements were compatible, and separating them would have introduced more operational coordination than the organization would have gained in flexibility.</p><p>Keeping them together was not architectural laziness. It was an economically rational decision. The image processing pipeline evolved in the opposite direction almost immediately. Its operational profile had very little in common with the rest of the platform. It required GPU access, consumed memory aggressively, and experienced bursty workloads whenever large listing imports arrived from MLS providers. Its failure modes were also isolated from the transactional concerns of the canonical listing system. A runaway image-processing workload should never be capable of starving listing writes or exhausting memory inside the systems responsible for maintaining canonical market data.</p><p>The separation happened because the operational realities diverged, not because the organization suddenly became enthusiastic about microservices. The search infrastructure followed a similar path. As traffic increased, North Harbor eventually moved semantic property search into its own runtime backed by a dedicated Elasticsearch cluster optimized for retrieval latency rather than transactional consistency.</p><p>The search system had requirements the canonical listing store simply did not share. It needed denormalized retrieval documents, vector similarity search for buyer preference matching, aggressive caching strategies, and infrastructure optimized for read-heavy workloads during peak browsing hours. Its scaling model, backup strategy, operational tooling, and performance tuning profile gradually became independent enough that keeping it inside the same runtime stopped making operational sense.</p><blockquote><p>The lead routing system split for a different reason entirely: latency sensitivity.</p></blockquote><p>At one point, lead assignment and outbound notifications shared the same runtime and worker infrastructure. Most of the time this was fine. Occasionally it was disastrous.</p><p>Notification delivery workloads were unpredictable. Push notification providers slowed down. SMS gateways rate-limited traffic. Email retries piled up during external outages. None of those problems were existential because delayed notifications were usually tolerable. A buyer receiving a saved-search alert ten seconds late was not catastrophic. A lead assignment arriving ten seconds late was.</p><p>The brokerage operated in a market where response speed directly affected conversion rates. Buyers contacting agents through the platform were often simultaneously browsing competing brokerages. Delays in routing a lead to the appropriate agent translated directly into lost opportunities.</p><p>Eventually the team realized they had accidentally coupled a latency-sensitive workflow to a best-effort communication system with entirely different operational expectations. That realization, more than any architecture diagram, justified the runtime split.</p><p>This is what granularity decisions look like in practice. Systems separate because their operational realities diverge enough that shared coordination becomes more expensive than independent operation. Not because someone drew more boxes on a diagram.</p><div><hr></div><h2>Monoliths, Distributed Systems, and the Real Cost</h2><p>There is an enormous amount of content on the internet about monoliths versus microservices. Most of it is not especially useful because it argues from aesthetics rather than cost. Distributed systems are expensive. This is not an opinion or a stylistic preference. It is a structural property of distributed computing itself.</p><p>The network is not reliable. Latency is not zero. Bandwidth is not infinite. Clocks are not perfectly synchronized. Topology changes. Failures are partial rather than obvious. Every distributed system inherits these constraints whether the architecture diagrams acknowledge them or not. At North Harbor, the reconciliation workflow exposed those costs very clearly.</p><p>The reconciliation process runs twice per day. Its responsibility is to compare the canonical listing state against the latest MLS provider feeds and resolve discrepancies between them. While the system operated as a monolith, reconciliation was mostly a transactional database operation. The consistency boundary was local. Failure handling was relatively straightforward. Once the system became distributed, reconciliation stopped being a transaction and became a coordination protocol.</p><p>Now the process involved multiple services with different consistency guarantees, different availability profiles, and independent deployment schedules. Updates could arrive out of order. One service could process a listing modification before another service had even received the event describing it. Temporary divergence became normal rather than exceptional.</p><p>Compensating workflows had to be introduced for situations where one subsystem accepted a change while another had not yet observed it. Every write path required idempotency guarantees because retries were no longer theoretical edge cases. Distributed locking was added to prevent concurrent reconciliation runs from producing contradictory canonical states. None of this logic was impossible to implement. All of it was expensive to operate.</p><p>The engineers who built the system were competent. The code was careful. The architecture reviews were thorough. The system still produced incorrect reconciliation results three times during the first six months after the migration. <strong>The failures were subtle.</strong></p><p>Diagnosing them required correlating logs across multiple services, reconstructing event ordering from clocks that were not perfectly synchronized, and reasoning about race conditions that appeared naturally in production traffic but almost never reproduced cleanly in test environments.</p><blockquote><p>The monolith would have handled all three incidents with a single database rollback.</p></blockquote><p>This is not an argument against distributed systems. North Harbor&#8217;s image processing pipeline genuinely needed runtime isolation. The search infrastructure genuinely required dedicated resources and an independent scaling model. Those were legitimate operational pressures.</p><p>The problem begins when distribution becomes the default assumption rather than a response to concrete requirements. Distributed systems are not evidence of architectural maturity. They are a trade-off that exchanges local complexity for coordination complexity, operational overhead, and failure modes that are significantly harder to reason about.</p><p>Serious engineers do not build distributed systems because distributed systems are fashionable. They build them when a single runtime can no longer satisfy the operational realities of the business. If those realities can still be handled by a well-structured modular monolith, then a well-structured modular monolith is not a compromise. <strong>It is the correct architecture.</strong></p><div><hr></div><h2>Conway&#8217;s Law and Organizational Architecture</h2><p>Conway&#8217;s Law states that organizations tend to produce systems that mirror their communication structures. This is not a metaphor or a management slogan. It is a recurring operational reality.</p><p>At North Harbor, the organization changed faster than the architecture for almost two years. The engineering team grew from three engineers to roughly thirty in less than eighteen months. Teams formed, split apart, merged, and reorganized repeatedly as business priorities shifted quarter by quarter.</p><p>The architecture did not evolve at the same speed. As a result, module boundaries created during the three-engineer phase quietly became ownership boundaries in the thirty-engineer phase, despite never having been designed for that purpose.</p><p>The canonical listing system exposed the problem clearly. Officially, the listings team owned it. In practice, almost every team modified it. The search team needed additional indexing fields. The recommendations team wanted richer buyer preference metadata. The integrations team needed provider-specific flags. The agent tools team needed additional lifecycle states.</p><p>Every request was individually reasonable. Collectively, they produced a system with weak ownership boundaries and steadily increasing internal incoherence. The canonical listing model slowly became a dumping ground for cross-team concerns because the organization had never established who was responsible for protecting the integrity of the boundary itself.</p><p>The eventual fix was organizational before it was technical. The listings team was given explicit ownership over the canonical schema along with veto authority over structural changes. Other teams could still propose modifications, but the listings team reviewed those requests for consistency and frequently redirected teams toward consuming listing lifecycle events instead of directly extending the canonical model itself.</p><p>That governance decision stabilized the architecture more effectively than any refactoring effort had managed to do. The inverse form of Conway&#8217;s Law, sometimes called the Inverse Conway Maneuver, is the deliberate design of organizational structures intended to encourage particular architectural outcomes.</p><p>North Harbor applied this principle when the company created a dedicated platform engineering team responsible for shared infrastructure concerns: the event streaming platform, deployment tooling, observability systems, and runtime infrastructure used across product teams.</p><p>Before that reorganization, those concerns evolved opportunistically inside unrelated application codebases. Logging conventions differed between teams. Deployment tooling drifted. Monitoring quality depended almost entirely on which engineer had last touched the service.</p><p>Creating a dedicated platform team changed the coordination structure around those concerns. Shared infrastructure stopped evolving as scattered local optimizations and started evolving as coherent systems with explicit ownership.</p><p>This is the part of architecture literature that many engineers underestimate. Architecture is organizational as much as it is technical. You cannot consistently produce clean system boundaries inside an organization whose communication structure continuously undermines them.</p><div><hr></div><h2>AI Changed Software Delivery, Not Architectural Reality</h2><p>The final point is the most important and probably the most misunderstood.</p><p>AI-assisted development tools genuinely change what a small engineering organization can build in a fixed amount of time. Code completion systems, context-aware review tooling, and autonomous coding agents compress implementation time dramatically.</p><p>At North Harbor, teams using strong AI-assisted workflows routinely shipped feature work that would previously have required significantly larger engineering teams. Feedback loops became shorter. Boilerplate disappeared faster. Moving from written requirements to working code became materially easier. <strong>None of this changed the architectural fundamentals.</strong></p><p>AI systems accelerate code generation. They do not reliably generate sound architectural judgment. An AI model cannot determine whether deduplication logic belongs inside the same transactional boundary as the canonical listing store because that decision depends on business liability, operational tolerances, organizational ownership, reconciliation behavior, and the real-world consequences of inconsistency in a real estate market.</p><p>Those are architectural decisions rooted in context rather than syntax. What AI changes most aggressively is the speed at which organizations accumulate complexity.</p><p>A team capable of shipping three times more functionality per quarter is also capable of accumulating three times more architectural debt if its structural decisions are weak. Faster delivery amplifies both good architecture and bad architecture simultaneously.</p><p>The engineers at North Harbor eventually identified a failure mode that became dramatically more common once AI-assisted coding workflows entered the organization: coherent local code paired with incoherent global structure.</p><p>An AI tool can generate a perfectly functioning feature that quietly violates a bounded context boundary, introduces hidden connascence between systems, or leaks internal semantics across service contracts. The generated code may be completely correct in isolation while still damaging the long-term coherence of the system.</p><p>The faster code generation becomes, the more valuable architectural awareness becomes. Architecture in the AI era is not less important than it was before. It is more important because the rate of change is higher, the surface area of systems grows faster, and the consequences of weak structural decisions compound more quickly.</p><p>The discipline of drawing boundaries carefully, managing coordination costs deliberately, and understanding what a system actually needs to optimize for has never mattered more. The engineers who have operated North Harbor&#8217;s platform through multiple years of growth eventually reached the same conclusion shared by nearly every organization that survives its own scale:</p><p>The architectural decisions that age well are rarely the ones chasing the newest technology or the most fashionable patterns. They are the decisions that correctly understand the problem, honestly assess the operational cost, and choose trade-offs the organization can realistically sustain. That is still what software architecture is. It was true when the stack was Rails and PostgreSQL.</p><p>It remains true when the stack includes vector databases, semantic search, embedding pipelines, and AI-assisted development systems. The tools evolve. The physics do not.</p><div><hr></div><p><em>This is Part One of an ongoing series on software architecture for production systems. Part Two will explore architectural fitness functions, evolutionary architecture, and how system boundaries drift under organizational pressure over time.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Book Review: Generative AI Design Patterns]]></title><description><![CDATA[A review for engineers who already know what RAG stands for and are tired of being told it's easy.]]></description><link>https://nidly.substack.com/p/book-review-generative-ai-design</link><guid isPermaLink="false">https://nidly.substack.com/p/book-review-generative-ai-design</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 01 Jun 2026 05:18:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_fHP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_fHP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_fHP!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!_fHP!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!_fHP!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!_fHP!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_fHP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg" width="1143" height="1500" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1500,&quot;width&quot;:1143,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:175714,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/197129951?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!_fHP!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!_fHP!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!_fHP!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!_fHP!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd83999d7-dcb9-4141-a320-73aad0c128a9_1143x1500.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>here is a specific kind of frustration that comes from reading a technically ambitious book that almost gets it right. The ideas are there. The vocabulary is correct. The patterns feel familiar if you&#8217;ve spent any real time building LLM-backed systems. And then somewhere around chapter four you encounter a diagram of a RAG pipeline that shows &#8220;Vector DB&#8221; as a single clean rectangle, and you know the author has never watched an embedding job silently corrupt a third of its documents because the chunking strategy collided with a PDF parser&#8217;s handling of footnotes.</p><p><em>Generative AI Design Patterns</em> is not a bad book. For a field generating content faster than it can generate understanding, it represents a genuine attempt at structural thinking. But it is a book written from the vantage point of someone who has designed these systems more than operated them. And that gap matters. It matters a lot, actually, if you&#8217;re the one getting paged at 2am because your retrieval pipeline started returning semantically plausible but factually stale chunks after a reindex.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>I read this over three weeks, mostly on transit and occasionally at a standing desk when I should have been reviewing PRs. What follows is not a neutral summary. It&#8217;s a working engineer&#8217;s read of a book about patterns I&#8217;ve had to implement, debug, and sometimes quietly abandon in production.</p><h2><strong>Why This Category of Book Exists Now</strong></h2><p>The software industry does not write pattern catalogs for things that are well understood. We write them when a problem space has matured just enough that people are making the same mistakes independently, and someone finally decides to name the shapes. The Gang of Four catalog came out because object-oriented codebases had accumulated enough scar tissue that behavioral patterns were recognizable across codebases. Richardson&#8217;s microservice patterns appeared when the distributed systems fallacies started biting teams that had split their monoliths without really thinking through what network partitions meant for their SLAs.</p><p>We&#8217;re at that inflection point with generative AI systems now. Not because the technology is mature, but because the failure modes have started repeating. Every team that has built a production RAG system has eventually discovered that their recall degrades at the edges of their corpus, that their embedding model and their query encoder have subtly different semantic priors, and that &#8220;just add more context&#8221; is not a retrieval strategy. These are structural problems. They deserve structural names.</p><p>That&#8217;s the honest case for this book. Not that AI is transforming everything, but that engineers are reinventing the same wheels in different languages and someone should write down what the wheels look like.</p><h2><strong>Where the Book Actually Earns Its Pages</strong></h2><p>The treatment of retrieval architecture is the strongest section. The author distinguishes between sparse retrieval (BM25-style, keyword-based), dense retrieval (embedding-based similarity), and hybrid approaches with more precision than most conference talks manage. More usefully, the book actually grapples with the question of when hybrid retrieval is worth the operational complexity. The honest answer is: more often than you want it to be, and the book says something close to this without flinching.</p><p>The chapter on context management is similarly good. There&#8217;s a pattern called &#8220;context compression&#8221; that gets real treatment here, not as a cute trick but as an architectural concern. When your retrieved documents exceed your context window, you have several options, and each has different latency, cost, and fidelity profiles. The book lays these out clearly. Hierarchical summarization, selective truncation, reranking before injection, and routing to a model with a larger context window are all discussed with enough specificity to be actionable.</p><div class="pullquote"><p><em>The book is at its best when it treats AI system design as a resource allocation problem rather than a prompting problem.</em></p></div><p>The orchestration patterns chapter does something I didn&#8217;t expect: it takes agent systems seriously as a systems problem rather than an AI problem. The distinction between fully autonomous agents, human-in-the-loop workflows, and deterministic pipelines with AI components is treated as an architectural decision with real trade-offs, not a philosophical one. If you want the agent to handle edge cases autonomously, you&#8217;re accepting nondeterminism in your execution graph, and that has implications for observability, retry logic, and error recovery that have nothing to do with the model&#8217;s capabilities. The book makes this connection explicitly, and it&#8217;s one of the places where the structural thinking actually pays off.</p><p>There&#8217;s also a useful pattern catalog around prompt versioning and management that I haven&#8217;t seen handled this carefully elsewhere. Prompts are artifacts. They have lifecycles. They need to be versioned, tested, and deployed with the same discipline as code, because changing a system prompt in production is a deployment, whether or not your CI/CD pipeline knows it. This framing is correct and more teams need to internalize it.</p><h2><strong>Where the Book Oversimplifies</strong></h2><p>The indexing and ingestion section is where I started leaving margin notes. The coverage of document processing is superficial in a way that will hurt readers who take it at face value. Chunking is treated as a parameter to tune rather than a source of fundamental retrieval problems. The book recommends chunk sizes in the 512-1024 token range and notes that overlap helps with boundary artifacts, which is true but incomplete. What it doesn&#8217;t discuss is that chunking strategy interacts with document structure in ways that are corpus-specific and largely empirical. Legal documents, scientific papers, customer support transcripts, and software documentation have completely different density and coherence patterns. A chunking strategy that works well for one will fail quietly on another.</p><p>&#8220;Fail quietly&#8221; is the operative phrase. This is a class of problem that doesn&#8217;t throw exceptions. Your pipeline completes. Your index builds. Your retrieval returns results. You only notice something is wrong when a user asks a question that spans a chunk boundary and the system confidently returns an answer that&#8217;s missing the second half of the relevant context. These failures are the kind that take weeks to diagnose because nothing is visibly broken.</p><p>The book&#8217;s treatment of observability is genuinely thin. For a field where the output space is continuous and the failure modes are semantic rather than syntactic, the observability patterns feel like an afterthought. There&#8217;s a brief section on logging LLM inputs and outputs and using evaluation frameworks, but the harder problem, which is detecting quality degradation in production without ground truth labels, gets almost no treatment. In traditional software you can define correctness precisely enough to alert on it. In generative AI systems you often can&#8217;t, and that requires different monitoring primitives: distribution shift detection on retrieved documents, coherence scoring, answer length distributions, user engagement signals used as proxy metrics. None of this is mentioned.</p><p><strong>A note on agent systems:</strong> The book describes agentic patterns with significant optimism about autonomous operation. Production experience tells a different story. Multi-step agents that invoke external tools introduce compounding failure rates. If each tool call succeeds 95% of the time and you have five steps in a chain, your end-to-end success rate is around 77%. That&#8217;s before accounting for semantic errors, where the agent completed successfully but did the wrong thing. Deterministic pipelines with narrow AI components are operationally cheaper and easier to debug. The book acknowledges this trade-off but doesn&#8217;t weight it heavily enough.</p><p>The chapter on cost management treats token budgets as a concern for the finance team rather than a first-class architectural constraint. This is backwards. In a system handling thousands of requests per hour, prompt size is a latency variable, a cost variable, and a reliability variable simultaneously. Decisions about context window usage, retrieval breadth, and summarization depth need to be made with an understanding of how they interact with throughput and cost under load. A pattern catalog that treats this as a secondary concern is leaving out a significant portion of what makes these systems hard to build at scale.</p><p>There&#8217;s also an assumption threaded through the book that your AI infrastructure components are reliable. The vector database is just there. The embedding service is just there. In practice, embedding services have rate limits, latency variance, and occasional model updates that change your embedding space in ways that break existing index entries. Managing the lifecycle of an embedding model, including what happens to your stored vectors when you need to upgrade to a better model, is a real operational problem. The book doesn&#8217;t address it.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/book-review-generative-ai-design?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/book-review-generative-ai-design?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/book-review-generative-ai-design?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><h2><strong>Compared to the Books That Set the Standard</strong></h2><p>Kleppmann&#8217;s <em>Designing Data-Intensive Applications</em> is the obvious comparison point. Not because the subject matter overlaps much, but because DDIA established what a technically serious book about system design looks like: it grounds every architectural decision in the specific failure modes that motivated it, it doesn&#8217;t pretend trade-offs don&#8217;t exist, and it treats the reader as someone who will need to make real decisions under real constraints. The chapter on replication in DDIA doesn&#8217;t just explain what replication is. It explains what breaks under each replication model, under what conditions, and what the operational consequences are when it does.</p><p><em>Generative AI Design Patterns</em> reaches for that standard and gets partway there. The pattern descriptions are more rigorous than anything you&#8217;d find in a typical AI blog post or conference talk, and the attempt at a consistent vocabulary is valuable. But the failure mode analysis is shallow. Patterns are presented with their intended use and their main benefit, with trade-offs that feel curated for tractability rather than completeness.</p><p>Richards and Ford&#8217;s <em>Fundamentals of Software Architecture</em> is another useful comparison. That book is explicitly about trade-off reasoning as a core engineering skill, the idea that architectural decisions can&#8217;t be evaluated without a clear understanding of what you&#8217;re optimizing for and what you&#8217;re giving up. The AI design patterns book nods to this perspective without fully committing to it. The result is a pattern catalog that tells you what to do more confidently than it tells you when not to.</p><p>This is a pattern in books about emerging technology. The urge to be comprehensive and usable leads to a certain softening of the failure cases. But in a field where the systems are expensive to run, difficult to evaluate, and often opaque in their failure modes, the failure cases are exactly what you need to understand.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><strong>Who Gets the Most From This</strong></h2><p>If you&#8217;re a software architect who hasn&#8217;t yet built a production LLM system and needs a structural map of the problem space before diving in, this book will save you significant time. The vocabulary alone is worth something. Knowing the difference between a naive RAG system and a corrective RAG system, or between an LLM router and a semantic cache, means you&#8217;re not inventing terminology when you discuss these patterns with your team.</p><p>Engineering managers overseeing teams working on AI products will find it useful for the same reason. You need enough conceptual vocabulary to evaluate technical proposals and ask the right questions without being able to evaluate every implementation decision yourself. This book gives you that vocabulary.</p><p>Senior engineers who have already shipped production RAG or agent systems will find it useful as a cross-reference rather than a primary resource. It&#8217;s a good way to check whether what you&#8217;ve built has a name, and whether the pattern you&#8217;ve evolved under operational pressure matches what the field is converging on. Sometimes it does. Sometimes you&#8217;ve invented something slightly different because your constraints were different, and that&#8217;s worth knowing too.</p><p>What the book is not is a substitute for operational experience. You will not finish this book and know how to run a production embedding pipeline reliably. You will not have clear intuitions about where your retrieval quality is degrading or how to detect it. Those things come from building systems, watching them fail, and understanding why. No catalog of patterns can compress that.</p><h2><strong>Final Thoughts</strong></h2><p>There is a version of generative AI engineering that looks, from a distance, like prompt engineering with a deployment step. You write some prompts, wire them to an API, point them at a vector store, and ship. That version of the discipline has a short ceiling. The systems stop being reliable at any meaningful scale, and the failure modes are hard to diagnose because the tooling assumes you care more about getting a response than about the consistency, correctness, and operational cost of that response across millions of requests.</p><p>The version of this discipline that will actually matter in three years looks a lot more like distributed systems engineering applied to a new class of component. The LLM is a stateless service with a nondeterministic response function and a token-denominated cost model. The vector store is a database with specific consistency trade-offs and a dependency on an external embedding model for write operations. The orchestration layer is a workflow engine that needs retry semantics, timeout budgets, and structured error handling. These are not novel problems. They are familiar problems wearing unfamiliar clothes.</p><p><em>Generative AI Design Patterns</em> understands this at some level, which is what makes it worth reading. The patterns it names are real patterns. The structural thinking it encourages is the right structural thinking. Where it falls short is in the operational depth that would make it genuinely useful to engineers building systems that have to stay up.</p><p>The book that will eventually exist for this field, the one that does for AI system design what DDIA did for data-intensive applications, will need to spend as much time on failure modes and operational realities as on architecture diagrams. It will need to treat retrieval quality degradation, embedding model lifecycle management, and agent reliability under composition as first-class topics rather than footnotes. It will need to be written by someone who has been paged because their RAG system started hallucinating confidently after a corpus update, and who can explain exactly what happened and why.</p><p>We&#8217;re not there yet. But this book is a step in the right direction, which is a more useful thing to say about it than either uncritical praise or dismissal. Read it with that framing and you&#8217;ll get real value from it. Just don&#8217;t mistake the map for the territory.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Why Software Architecture Matters More in the AI Era]]></title><description><![CDATA[LLMs changed how software is built, but scalability, reliability, orchestration, and architectural trade-offs still determine whether AI systems survive production.]]></description><link>https://nidly.substack.com/p/why-software-architecture-matters</link><guid isPermaLink="false">https://nidly.substack.com/p/why-software-architecture-matters</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 25 May 2026 00:51:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8fhN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8fhN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8fhN!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!8fhN!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!8fhN!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!8fhN!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8fhN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg" width="3447" height="2254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2254,&quot;width&quot;:3447,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1556998,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/195978083?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39e67691-8a19-491f-9afe-66dca88c30fa_3447x5170.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8fhN!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!8fhN!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!8fhN!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!8fhN!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8caa9d-ccb5-494c-8b1a-b964d66fea72_3447x2254.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@bormot?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Alexandr Bormotin</a> on <a href="https://unsplash.com/photos/tall-apartment-building-with-many-windows-and-balconies-vI3BV8AwHQc?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>...</figcaption></figure></div><p>There is a persistent belief moving through engineering teams right now that AI will reduce the need for software architecture. The logic sounds convincing at first. Models can generate code, reason over documents, summarize workflows, and increasingly handle behavior that once required carefully designed application logic. If intelligence is becoming an API call, then perhaps the surrounding system matters less. That assumption usually survives right until production. The difficult parts never disappeared. They moved.</p><p>The complexity that once lived inside deterministic application code now leaks outward into orchestration layers, retrieval pipelines, evaluation systems, fallback strategies, latency management, and operational governance. AI systems are still distributed systems. In some ways, they are harder to reason about than traditional ones. Their components are probabilistic, their behavior shifts over time, and many of their failures degrade silently before anyone notices.</p><p>This is why so many AI applications look impressive in a demo and become deeply unstable under real usage. Early prototypes hide the operational burden. A prompt works a few times. Retrieval appears accurate enough. Latency seems acceptable at small scale. Then production traffic arrives, data changes, embeddings age, upstream providers evolve their models, costs become unpredictable, and suddenly the system is no longer behaving like the clean diagram everyone approved three months earlier.</p><p>What makes this difficult is that the failures rarely appear all at once. They accumulate gradually. Retrieval precision slips after incremental indexing changes. A model upgrade subtly alters output structure that downstream services quietly depended on. Context windows grow, latency increases, and fallback logic becomes tangled enough that nobody fully understands which path a request actually took. Traditional systems usually fail loudly. AI systems often fail slowly.</p><p>The industry did not eliminate architectural complexity. It redistributed it into places many teams were not prepared to observe, govern, or design carefully. That shift is precisely why software architecture matters more now, not less.</p><div><hr></div><h2>The Model Is Not the System</h2><p>The model is not your system. Most teams realize this late, usually after something expensive has already broken.</p><p>Integrating a language model into a production application does not mean deploying a model. It means deploying a distributed system that happens to contain a probabilistic component with weak behavioral guarantees. That distinction matters more than most teams initially expect.</p><p>Traditional infrastructure components fail in relatively legible ways. A database connection drops. A downstream service returns a 503. A message queue backs up and lag increases. These problems are difficult, but they are observable. Their behavior is bounded enough that engineers can usually reason about them systematically. Language models behave differently.</p><p>They do not throw exceptions when reasoning degrades. They return text. That text may be correct, approximately correct, structurally inconsistent, subtly misleading, or confidently wrong. Worse, many of these failures do not appear consistently. The same prompt may succeed repeatedly and then fail under slightly different context, retrieval quality, or model behavior after a provider update.</p><p>This changes the nature of system design around the component itself. Failures become harder to reproduce, harder to monitor, and harder to govern using traditional operational patterns. A model can silently degrade on edge cases your evaluation suite never covered. Retrieval quality can decline gradually as embeddings age and incremental indexing introduces inconsistencies. A model upgrade can alter output structure that downstream services had quietly started depending on.</p><p>None of these failures announce themselves immediately. Most accumulate slowly enough to evade dashboards while user trust erodes underneath them.</p><p>This is not a criticism of language models. It is a description of the architectural contract they expose. Unlike most infrastructure components, their behavior is probabilistic, evolving, and only partially observable. The surrounding system has to absorb that uncertainty. That is an architectural problem. Not a prompting problem.</p><div><hr></div><h2 style="text-align: justify;">Retrieval Is an Architectural System</h2><p style="text-align: justify;">Retrieval pipelines are where many teams first discover that AI systems are fundamentally architectural systems, not just model integrations.</p><p><a href="/__u/nidly.substack.com/p/rag-isnt-about-embeddings?r=a3p8i">RAG</a> became the dominant pattern for understandable reasons. It keeps information current, avoids constant fine-tuning, and allows systems to work with knowledge the model never saw during training. But retrieval is not a feature you attach to an application. It is its own distributed subsystem, with operational characteristics, trade-offs, and failure modes that evolve over time. Embedding freshness is one of the least visible and most operationally important problems in production retrieval systems.</p><p>Embeddings capture semantic relationships at a particular moment, using a particular model, against a particular distribution of data and queries. As documents evolve, user behavior shifts, and embedding models change, retrieval quality begins drifting in ways that are difficult to detect directly. Nothing crashes. No alert fires. The system continues returning results that appear plausible while becoming progressively less useful.</p><p>This is what makes retrieval degradation dangerous. Failure rarely looks catastrophic. It looks almost correct. Chunks still match keywords. Similarity scores still look healthy. Offline evaluation metrics may remain stable enough to avoid concern. Meanwhile, users increasingly fail to retrieve the information they actually needed because semantic relevance has drifted away from operational relevance. The precision-recall trade-off compounds this further.</p><p>Increasing recall improves the probability of retrieving relevant context, but it also increases noise, latency, token consumption, and the chance that the model anchors on misleading information. Optimizing aggressively for precision reduces noise but increases the risk of silently excluding critical context altogether.</p><p>These are not isolated tuning decisions. Retrieval depth, chunking strategy, metadata filtering, query rewriting, hybrid search, and re-ranking layers interact continuously with one another. Changing one layer alters the behavior of the entire pipeline. Teams often discover this coupling accidentally, usually after a retrieval improvement in one metric causes degradation somewhere else in production.</p><blockquote><p>At that point, the problem is no longer prompting quality. It is system architecture.</p></blockquote><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Latency in AI Systems Is an Architectural Problem</h2><p style="text-align: justify;">Latency in AI systems is not just a performance concern. It changes the shape of the entire system around it.</p><p>Traditional backend latency problems were difficult but relatively predictable. Database queries could be optimized. Caches reduced repeated work. Slow downstream services could be isolated behind retries, queues, or asynchronous workflows. Most of the time, engineers were optimizing deterministic operations whose behavior stayed reasonably stable under known conditions. Language model inference behaves differently.</p><p>Inference is inherently slow, computationally expensive, and directly exposed to the user experience in many applications. Unlike batch-oriented distributed systems, users are often waiting interactively for the result. A single request may involve retrieval, re-ranking, prompt construction, multiple inference calls, tool execution, and post-processing before a response is returned. The latency accumulates surprisingly fast, especially once orchestration layers become more sophisticated.</p><p>RAG pipelines intensify this further. Retrieval adds network and search latency. Re-ranking introduces additional model execution. Query rewriting may trigger another inference step before retrieval even begins. Agentic workflows compound the problem because each reasoning step can generate additional downstream calls whose execution path is not fully predictable ahead of time.</p><p style="text-align: justify;">The traditional optimization techniques still matter, but they behave differently around probabilistic systems.</p><p>Caching remains essential, but exact-match caching becomes less useful when semantically similar requests produce equivalent outcomes. Teams often move toward semantic caching strategies, which immediately introduces new trade-offs around similarity thresholds, freshness guarantees, invalidation complexity, and retrieval consistency. The system is no longer deciding only whether data is fresh. It is deciding whether meaning is close enough.</p><p>Streaming responses help reduce perceived latency, but they introduce operational complexity of their own. Partial outputs may need moderation before completion. Structured responses become harder to validate incrementally. Errors occurring midway through generation are significantly more difficult to recover from cleanly than failures in traditional request-response systems.</p><p style="text-align: justify;">Model selection itself becomes an architectural negotiation between latency, quality, and cost.</p><p>Smaller models are faster and cheaper but may fail unpredictably on edge cases that larger models handle reliably. Larger models improve reasoning quality but increase inference time and operational cost substantially. Most teams eventually realize there is no globally correct choice. The optimal model depends on workload distribution, acceptable failure rates, user expectations, and the economic constraints of the product itself.</p><p>What makes this especially difficult is that the quality side of the trade-off is hard to measure rigorously. In traditional systems, performance optimization usually preserves deterministic correctness. A faster database query that returns the same result is unambiguously better. AI systems do not offer that clarity. Reducing latency often changes output quality in ways that are subtle, domain-specific, and difficult to evaluate comprehensively.</p><p>This pushes engineering teams into a problem space that looks much closer to experimentation infrastructure than traditional software optimization. Offline evaluation pipelines, regression detection, benchmark datasets, human review loops, and production telemetry become necessary parts of the architecture. Without them, teams are effectively optimizing blind.</p><p style="text-align: justify;">Operating blind around deterministic systems is dangerous. Operating blind around probabilistic systems is substantially worse.</p><div><hr></div><h2 style="text-align: justify;">The Observability Gap in AI Systems</h2><p style="text-align: justify;">Observability becomes significantly harder once probabilistic components enter the system.</p><p>Traditional observability tooling was designed around deterministic infrastructure. Traces, metrics, and logs can usually explain where a request traveled, how long each operation took, and which component failed. That model works well when correctness is relatively easy to define. AI systems break this assumption.</p><p>A trace may show that an embedding request completed successfully, retrieval returned results within latency targets, and inference finished without error. None of that tells you whether the retrieved context was actually relevant or whether the generated response was correct. From the perspective of conventional monitoring, the system appears healthy while users quietly receive degraded results.</p><blockquote><p>This is not a tooling gap alone. It is a structural property of probabilistic systems.</p></blockquote><p style="text-align: justify;">Correctness becomes harder to observe directly because failures are often semantic rather than operational. Retrieval quality drifts gradually. Model outputs become subtly inconsistent after provider updates. Edge cases emerge only under combinations of context that evaluation suites never captured. Many failures do not trigger alerts because technically nothing crashed.</p><p>Closing this gap requires building infrastructure that many teams underestimate because it does not resemble product development.</p><p>Evaluation harnesses, regression datasets, human review workflows, production feedback loops, prompt version tracking, retrieval diagnostics, and logging of intermediate artifacts become essential operational components. Without them, observability becomes mostly cosmetic: dashboards that accurately measure latency while telling you almost nothing about output quality or behavioral degradation. Fallback strategies matter for the same reason.</p><p>If retrieval confidence drops, what behavior does the system fall back to? If a provider experiences elevated latency, does the application degrade gracefully or simply stall? If a model update changes output structure and breaks downstream parsing, is there isolation between components or does the failure cascade through the workflow?</p><blockquote><p>These are not exceptional edge cases. They are normal operating conditions for production AI systems.</p></blockquote><p>Teams that treat them as future optimizations usually discover the cost later, when reliability problems have already spread across tightly coupled workflows. Retrofitting resilience into probabilistic systems is substantially more expensive than designing for uncertainty from the beginning.</p><div><hr></div><h2 style="text-align: justify;">AI Complexity Lives in the Orchestration Layer</h2><p style="text-align: justify;">Orchestration is where a significant amount of AI system complexity ultimately accumulates, and most teams underestimate it early.</p><blockquote><p>The prompt is not the system.</p></blockquote><p>Prompts matter, but they exist inside a much larger orchestration layer responsible for retrieval, context assembly, conversation state, retries, fallback handling, guardrails, tool coordination, response validation, and interaction with external systems. That layer is still software architecture. It can be modular or tightly coupled, observable or opaque, resilient or fragile. As capabilities expand, orchestration complexity grows disproportionately.</p><p>A single-turn interaction with one model is relatively straightforward. Add conversational memory and state management appears. Add retrieval and the system inherits indexing, ranking, and freshness problems. Add tool execution and the failure surface expands beyond incorrect responses into incorrect actions against external systems. Add multiple models or specialized agents and the architecture begins behaving like a distributed coordination system with asynchronous workflows, competing execution paths, and increasingly difficult debugging characteristics.</p><p style="text-align: justify;">The model does not absorb this complexity. The surrounding system does. This is why many AI systems feel deceptively simple in prototypes and operationally chaotic in production. Capability expands faster than coordination discipline around it.</p><p style="text-align: justify;">Model provider coupling introduces another architectural risk that often remains invisible until migration becomes necessary.</p><p>Models are not interchangeable infrastructure components. Their behavior, latency characteristics, pricing structures, context handling, output formatting, and failure patterns differ in ways that leak into application logic surprisingly quickly. Teams frequently discover that what appeared to be &#8220;just an API call&#8221; had quietly become embedded throughout prompts, parsers, workflows, evaluation assumptions, and operational tooling.</p><p style="text-align: justify;">At that point, changing providers or upgrading model versions stops being a configuration change and becomes a system migration.</p><p>Designing abstraction boundaries around model interactions is therefore not premature optimization or unnecessary indirection. It is evolutionary architecture applied to a component category that is guaranteed to evolve faster than the surrounding system built around it.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/why-software-architecture-matters?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/why-software-architecture-matters?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/why-software-architecture-matters?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p style="text-align: justify;"></p><div><hr></div><h2 style="text-align: justify;">Governance Is an Architectural Concern</h2><p style="text-align: justify;">Governance becomes unavoidable once AI systems begin interacting with real operational workflows.</p><p>Many engineering teams instinctively treat governance as a process problem rather than a system design problem. It sounds bureaucratic, adjacent to compliance, and disconnected from technical architecture. In practice, the opposite is true. Governance determines what the system must be capable of observing, restricting, auditing, and recovering from. That makes it architectural.</p><p>Questions around auditability, decision tracing, permission boundaries, output validation, and acceptable system behavior are not organizational overhead layered on top of the application afterward. They directly shape storage design, orchestration logic, observability requirements, access control models, and operational workflows from the beginning. The importance of this increases alongside system capability.</p><p>A model that summarizes documents has a relatively constrained failure surface. A system that executes actions, modifies records, writes code, triggers workflows, or communicates externally introduces consequences that extend beyond the interface itself. Failures stop being informational and become operational.</p><p>At that point, reversibility matters. Audit trails matter. Constraint enforcement matters. Human override paths matter. Designing highly capable systems without these controls is not moving quickly. It is accumulating operational liability that eventually surfaces under production pressure.</p><p>The architectural characteristics that have always mattered in distributed systems still matter here: reliability, scalability, observability, maintainability, security, and evolvability. What changed is not their importance. What changed is the difficulty of achieving them.</p><p>AI systems introduce components whose behavior is probabilistic, partially observable, and continuously evolving on timelines controlled by external providers. The surrounding architecture has to absorb that instability while still presenting predictable behavior to users and downstream systems. That work does not happen inside the model. It happens in the architecture around it.</p><div><hr></div><h2>Most AI Failures Are Software Failures</h2><p>Most production failures in AI systems are not model failures.</p><p>That is the conclusion many teams eventually reach after enough operational incidents and honest post-mortems. The model usually behaved within the range of behavior the system already allowed for. The actual failure happened elsewhere.</p><p>Retrieval quality degraded because embeddings were stale and nobody was measuring semantic drift. A tool invocation failed and the orchestration layer exposed raw failure behavior directly to users because fallback handling was incomplete. A model upgrade introduced subtle output changes that downstream parsers quietly depended on. Evaluation pipelines failed to detect regressions because benchmark datasets captured only ideal cases rather than real production variability.</p><blockquote><p>These are not fundamentally new categories of failure. They are software failures.</p></blockquote><p>They emerge from the same root causes that have existed in distributed systems for decades: weak observability, poor boundary design, insufficient testing, operational shortcuts, hidden coupling, and infrastructure assumptions that stopped being true under production conditions.</p><p>The model did not create most of these problems. The absence of architecture around the model did. This is why the belief that AI reduces the importance of software architecture consistently collapses in real systems.</p><p>AI applications are still distributed systems. They still inherit the properties that made distributed systems difficult long before language models existed: partial failure, unreliable networks, asynchronous workflows, eventual consistency, and limited visibility into global system state.</p><p>What AI systems add is another layer of instability on top of those existing constraints. Outputs become probabilistic. Evaluation becomes less deterministic. Behavior evolves across model versions controlled by external providers. Retrieval quality drifts gradually instead of failing explicitly. Latency, cost, orchestration complexity, and provider coupling become tightly interconnected.</p><p>None of this simplifies system design. It increases the amount of architectural discipline required to keep systems reliable over time.</p><p>The teams building resilient AI systems are rarely the ones obsessed with prompts alone. They are the teams treating retrieval pipelines as operational infrastructure, investing early in evaluation systems, designing graceful degradation paths, instrumenting semantic behavior instead of only technical performance, and applying the same engineering rigor to AI components that they would apply to any other unreliable distributed dependency.</p><p>That work is less visible than a polished demo. It is also the reason the demo still works six months later.</p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[I Replaced My Entire Code Review Process With Claude Code. Here's What Broke]]></title><description><![CDATA[A real-world experiment replacing human code review with AI and what actually failed.]]></description><link>https://nidly.substack.com/p/i-replaced-my-entire-code-review</link><guid isPermaLink="false">https://nidly.substack.com/p/i-replaced-my-entire-code-review</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 18 May 2026 06:44:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-xCV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-xCV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-xCV!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!-xCV!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!-xCV!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-xCV!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-xCV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1645873,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/195537465?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-xCV!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!-xCV!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!-xCV!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-xCV!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f8fe1c-a119-4feb-a2ec-300f52ce673c_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We ran an experiment most teams avoid for a reason. Every pull request went through <a href="https://claude.com/product/claude-code">Claude Code</a> before any human touched it. Not as a suggestion engine or a linting layer, but as the primary reviewer.</p><p>For the first few weeks, it looked like an upgrade. Feedback was instant, the review queue collapsed, and engineers stopped waiting two days to be told their naming was inconsistent or a null check was missing. On paper, everything improved. Review turnaround dropped from 47 hours to under 3, merge velocity increased, and Slack got quieter in all the right places. It felt like we had removed a bottleneck.</p><p>Then the failures started quietly. No outage, no obvious regression, just a growing set of changes that were technically correct and strategically wrong. Code that passed review while encoding assumptions we would never have approved explicitly. Code that handled the happy path perfectly, but ignored entire failure modes we had already seen in production. Code that looked clean, consistent, and completely misaligned with how the system actually behaves.</p><p>The worst part was that we didn&#8217;t notice. From the outside, the system looked healthy. Reviews were happening, feedback was detailed, and everything was moving faster right up until it wasn&#8217;t.</p><div><hr></div><h2>Why We Tried This</h2><p>Code review was the slowest part of our development cycle, and not for the reasons people like to pretend. Engineers were not lazy. The problem was attention. Good review requires sustained attention, and that is the scarcest resource on any team.</p><p>A typical pull request would sit for 24 to 48 hours before getting a first look. By then, the author had already context-switched. The reviewer had to reconstruct the intent from scratch. What should have been a one-hour review turned into a two-day back-and-forth over small corrections and missed assumptions. We were a team of twelve engineers serving around 2 million active users, and this bottleneck was quietly limiting how fast we could ship.</p><p>We tried the standard fixes. Smaller PRs, better descriptions, dedicated review windows. They improved things at the margins, but they did not change the core constraint. Human review is expensive, and we did not have enough of it.</p><p>Claude Code changed the equation. The promise was simple: a reviewer with no queue, no context switching, and no attention decay. It could read the entire diff in seconds and return structured feedback immediately.</p><p>We ran a four-week experiment. Every pull request went through Claude Code first, and human reviewers saw that feedback before adding their own. After two weeks, we pushed further and made human review optional for most changes. That was the mistake. Not using Claude Code, but making the human optional.</p><div><hr></div><h2>What Actually Improved</h2><p>The improvements were real, and it would be dishonest to pretend otherwise. Claude Code is extremely consistent. Human reviewers are not. A pattern that gets approved on Monday can get flagged on Thursday, depending on who is reviewing and how much attention they have left. Claude Code applies the same standard every time. For enforcing style, naming, and structural consistency, that reliability is valuable.</p><p>It also catches the obvious bugs humans miss when they skim. Null pointer risks, missing error handling in async paths, off-by-one errors. The kind of issues that slip through because everyone assumes someone else will catch them. Claude Code does not assume. It reads the entire diff.</p><p>Feedback speed mattered more than we expected. Especially for junior engineers. When feedback arrives in minutes, you are still inside the problem. The context is fresh, and the feedback is immediately actionable. When it arrives two days later, you have already moved on, and coming back feels like overhead.</p><p>We also saw a reduction in review politics. Every team has unspoken patterns where certain reviewers push back on certain styles, and authors learn to agree and move on. Claude Code has no history with anyone. It evaluates the code without context, and that removed a layer of friction we had normalized.</p><p>None of this was trivial. These were real improvements to a real workflow. But they all operated at the surface layer of code review, which turned out not to be the part that mattered most.</p><div><hr></div><h2></h2><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>What It Could Not See</h2><h4>Hidden Context</h4><p>Claude Code reviews the code in the diff. It does not review the assumptions behind the code. That sounds obvious until it fails in production.</p><p>Take a simple case: input validation before writing to a database. Claude Code checks that the logic is correct, covers common edge cases, and is implemented cleanly. What it cannot check is whether those validation rules match what the product team agreed on in a Slack thread six weeks ago, or whether they contradict a compliance requirement buried in a PDF no one links to.</p><p>We hit this with rate limiting. The implementation was solid. Correct token bucket, well-tested, clean code. Claude Code approved it with minor comments. What it missed was that we already had a contractual obligation with an enterprise client that required different behavior for their tier. That contract lived outside the codebase.</p><p>The code worked. The violation was invisible. A human reviewer with context would not need to know the answer. They would just ask the question: does this interact with our tiered rate limiting?</p><blockquote><p>Claude Code cannot ask questions it does not know exist.</p></blockquote><div><hr></div><h4>Architectural Judgment</h4><p>There is a difference between &#8220;does this work&#8221; and &#8220;should we be doing this at all.&#8221;</p><blockquote><p>Good code review occasionally catches the second.</p></blockquote><p>A reviewer might ask: why are we adding a new service instead of extending an existing one? Why are we reimplementing something we already have? This works now, but what happens in three months when the data model changes?</p><p>Claude Code evaluates the implementation of a decision. It does not challenge the decision itself.</p><p>We introduced a caching layer in a place where it worked technically but was wrong architecturally. It improved latency for a local metric, and Claude Code gave useful feedback on TTLs and invalidation. No one questioned whether caching belonged there at all. Everyone assumed it had been vetted.</p><p>Four months later, we had to rip it out. It created consistency issues we had already solved elsewhere. The cost of that rework was far higher than a single uncomfortable question during review.</p><div><hr></div><h4>Security Blind Spots</h4><p>Claude Code understands security patterns. It will flag injection risks, missing sanitization, and hardcoded secrets. These are real catches. What it cannot do is model threats in your system.</p><p>Security is not just pattern matching. It is context. Who can access this? What can they do? What happens if they abuse it?</p><p>We added an internal admin endpoint with basic auth. Claude Code flagged it and suggested token-based authentication. Fair point. The engineer explained it was behind a VPN. Claude Code accepted the resolution.</p><p>What neither caught was that the endpoint allowed bulk data export with no rate limiting and no audit logging. A compromised internal account could extract the entire dataset without detection.</p><p>The problem was not the authentication mechanism. It was the combination of capability, access, and lack of observability. That is a threat modeling problem. Not a syntax problem.</p><div><hr></div><h4>The Social Layer</h4><p>This was the part we underestimated. Code review is not just about correctness. It is how teams distribute knowledge. When a senior engineer reviews a PR and asks a question, that question carries context. It tells you what matters, what exists already, and what has failed before.</p><p>Claude Code gives answers. It does not ask those questions. Over four weeks, junior engineers got better at writing code that passes automated review. They did not get better at understanding the system. The feedback loop optimized for correctness at the line level, and quietly broke the feedback loop for architectural intuition.</p><p>Code review is one of the few ways teams move tacit knowledge around. When you automate around it, that knowledge stops moving.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/i-replaced-my-entire-code-review?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/i-replaced-my-entire-code-review?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/i-replaced-my-entire-code-review?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h3>The Failure Mode</h3><p>The real danger is not that Claude Code gives wrong answers. It is that it gives wrong answers with the same tone it uses for correct ones.</p><p>Human reviewers signal uncertainty. They say things like &#8220;I might be wrong here&#8221; or &#8220;worth checking with someone closer to this part of the system.&#8221; Those signals matter. They tell you where to trust the feedback and where to think.</p><p>Claude Code does not communicate uncertainty in a way engineers act on. It produces structured, detailed, authoritative feedback. To a junior engineer, it all looks equally valid.</p><p>We saw this in practice. Engineers would apply the suggested changes, resolve the comments, and move on. They were not evaluating the feedback. They were complying with it.</p><blockquote><p>The authority of the tool replaced their own judgment.</p></blockquote><p>This is not a single-bug problem. It is a system-level risk. You are training a team to stop questioning answers that look confident. That habit does not stay contained to code review.</p><p>Scale amplifies it. A wrong assumption in a human reviewer affects a handful of pull requests. A wrong assumption in Claude Code affects every pull request. You are distributing plausible mistakes at the throughput of your entire team.</p><div><hr></div><h3>The Slip</h3><p>In week three, a pull request changed how we handled session expiration for long-running background jobs.</p><p>The code was clean. The logic was correct. Claude Code approved it, suggested a minor logging improvement, the author applied it, and the PR merged.</p><p>Two weeks later, failures started showing up in a data export pipeline for enterprise clients. Jobs were expiring mid-run. Not immediately, but around four hours in, on workloads that used to complete in three.</p><blockquote><p>The logic was not broken. The assumption was.</p></blockquote><p>The implementation assumed that background jobs would complete within the same session window used for interactive users. That had been true when the code was originally written. It stopped being true when we onboarded a new client tier with larger datasets six weeks earlier.</p><p>No one connected those dots during review. Claude Code could not. A human reviewer who had context from the onboarding might have. Or might not have. But the question would at least have existed.</p><p>We spent three days diagnosing the issue, one day fixing it, and two more explaining it to the client. The fix was a one-line configuration change. The missing review comment was a single question: does this assume anything about job duration?</p><div><hr></div><h3>What We Kept</h3><p>We did not remove Claude Code. We changed how we trust it. It is now the mandatory first pass on every pull request. It runs automatically, and engineers see the feedback before requesting human review. This strips out surface-level issues before a human spends attention on them. That part worked.</p><p>Human review is mandatory. Making it optional was the actual mistake. AI review does not replace the human. It shifts what the human needs to focus on.</p><p>We defined explicit trust boundaries. Claude Code is trusted for style, obvious bugs, and test coverage gaps. It is not trusted for security decisions, core data model changes, or anything tied to external contracts or compliance. Those always require a human reviewer with context.</p><p>We also changed how we use it. For infrastructure-heavy changes, we explicitly prompt it to surface assumptions and flag anything it cannot verify from the codebase. It does not do this perfectly, but often enough to be useful. Prompting for uncertainty is one of the few levers you actually have.</p><p>The workflow is slower than full automation. It is faster than what we had before. That is the point. The goal was never maximum speed. It was sustainable velocity without hidden risk.</p><div><hr></div><h3>Conclusion</h3><p>Claude Code <em>did not fail</em>. It exposed how we were using code review. We treated review as validation: does this work, does it follow standards. Claude Code handles that well. What we undervalued was everything else review does. It distributes context, surfaces architectural questions, and forces someone to think about how a change fits into the system.</p><p>When you add Claude Code to a team with strong reviewers, it makes them faster. When you add it to a team without them, it gives you the appearance of good review without the substance. It produces the artifact of review, not the outcome.</p><p>The engineers who benefit most are the ones who already know what to look for. For everyone else, it accelerates their ability to ship code that passes review. Those are not the same thing. AI review does not replace judgment. It reveals whether your team had it.</p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Scaling to Millions (Part 1): The System Design Mistake That Hides in Plain Sight]]></title><description><![CDATA[Your system handles millions of requests. Latency is fine. No errors in sight. And it's slowly becoming wrong in ways no monitor will ever catch.]]></description><link>https://nidly.substack.com/p/scaling-to-millions-part-1-the-system</link><guid isPermaLink="false">https://nidly.substack.com/p/scaling-to-millions-part-1-the-system</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 11 May 2026 02:38:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!p1t1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!p1t1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!p1t1!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!p1t1!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!p1t1!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!p1t1!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!p1t1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg" width="1024" height="682" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:682,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:135322,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/194421411?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!p1t1!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!p1t1!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!p1t1!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!p1t1!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e771758-3a43-429b-b0f2-cc2685edf636_1024x682.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Your system handles millions of requests. Latency is fine. No errors in sight. And it&#8217;s quietly returning the wrong results.</p><p>This isn&#8217;t hypothetical. It&#8217;s the default failure mode of most systems once they reach a certain scale. Not a crash, not a 500, not even a spike in latency. Just a slow, invisible drift away from truth. The dashboard stays green, the on-call stays asleep, and users start noticing things your system doesn&#8217;t measure.</p><p>Most engineering conversations about scale focus on throughput, availability, and latency. These are real problems, and we&#8217;ve gotten very good at solving them. Horizontal scaling is well understood. Load balancing is a commodity. Auto-scaling is expected. Somewhere along the way, we started treating &#8220;the system is fast and available&#8221; as &#8220;the system is correct.&#8221; They are not the same thing.</p><div><hr></div><h2><strong>The Illusion of a Healthy System</strong></h2><p>Consider a search service running in production. It handles millions of requests per day, keeps P99 latency under 80ms, and maintains an error rate below 0.01%. The on-call engineer hasn&#8217;t been paged in weeks. By every metric on the dashboard, the system looks healthy.</p><p>Now consider two users searching for the same product at the same moment. They receive different results. Not because of personalization or an experiment, but because one request hits a replica that processed a catalog update fifteen seconds ago, and the other hits one that hasn&#8217;t. Both replicas are healthy. Both responses are returned successfully. Only one reflects reality.</p><p>Nothing in the monitoring system flags this as a problem. No errors are thrown, no thresholds are crossed, and no alerts fire. The system behaves exactly as designed, yet it is wrong for a meaningful fraction of the requests it serves.</p><p>This is what most production correctness failures look like. They don&#8217;t announce themselves; they accumulate. They surface through user behavior long before they appear in metrics, which is exactly why they persist for so long without being addressed.</p><p>Consider a profile service. A user updates their display name and receives a 200 response. The write succeeds, but for the next 45 seconds, parts of the product still show the old name because multiple downstream services cache the value with a TTL that hasn&#8217;t been revisited since the initial deployment. The user refreshes and still sees stale data. Eventually, they open a support ticket. Meanwhile, the system continues to report a near-zero error rate.</p><p>Or consider a payment system where a consumer restarts mid-processing and a message is redelivered. The same event is processed twice. No exception is thrown, no alert fires, and both executions complete successfully. Duplicate handling was never implemented because the team assumed exactly-once delivery, which the system never actually guaranteed.</p><p>None of these are edge cases. They are the expected behavior of distributed systems under normal conditions. The problem isn&#8217;t that they happen; the problem is that most systems are instrumented around the wrong definition of failure.</p><p>Green dashboards indicate availability. They say nothing about correctness. The system isn&#8217;t failing; it&#8217;s diverging from reality.</p><div><hr></div><h2><strong>What &#8220;Working&#8221; Actually Means at Scale</strong></h2><p>When an engineer says a service is working, they almost always mean: it&#8217;s responding, it&#8217;s within SLA, and it hasn&#8217;t thrown any errors recently. These are necessary conditions. At low scale, they&#8217;re usually sufficient.</p><p>You can&#8217;t serve a million requests per second and have a meaningful fraction of them return stale or inconsistent results without noticing. Everything happens fast enough that divergence doesn&#8217;t have time to accumulate in visible ways. Scale breaks that assumption.</p><p>The more requests you serve, the more paths through your system a request can take. The more paths, the more opportunities for those paths to return different answers to logically identical queries. And under higher load, replication lag, cache freshness, and event backlogs start to matter in ways they didn&#8217;t when the system was smaller.</p><p>A stricter definition of &#8220;working&#8221; requires three properties at the same time. The response should be correct, meaning it reflects the actual state of the data. It should be consistent, meaning the same query returns the same result regardless of which path handled it. And it should be fresh, meaning the data is recent enough to be useful.</p><p>Most production systems guarantee none of these explicitly. They guarantee availability and latency. The rest is left to chance.</p><p>Take replication lag as an example. In a typical leader-follower setup, writes go to the primary and propagate asynchronously to replicas. Under normal conditions, lag is small. Under high write throughput, it grows. During recovery or network issues, it can spike to seconds.</p><p>In that window, a user can write data and immediately read from a replica that hasn&#8217;t caught up. From their perspective, the write didn&#8217;t happen. They retry. Now you have duplicate state. The replica is healthy. Replication is working as designed. The lag is within configured thresholds. The system is returning incorrect data.</p><p>Monitoring shows nothing unusual because it measures whether the replica is running, not whether it is current. Throughput is a proxy metric. It says nothing about truth.</p><div><hr></div><h2></h2><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><strong>How Correctness Decays, Without a Single Crash</strong></h2><p>This class of failure persists because it accumulates gradually. Each component behaves correctly in isolation. The problem emerges from their interaction over time.</p><h4><strong>Stale cache:</strong></h4><p>Every cache is a bet that the data won&#8217;t change before the TTL expires. That bet always has a cost.</p><p>At low scale, stale reads are rare enough to ignore. At scale, they compound. A catalog updated nightly tolerates long TTLs. A pricing system reacting to inventory does not. A user account may require near-real-time correctness while other data in the same cache does not.</p><p>The issue isn&#8217;t caching itself. It&#8217;s that cache policies are rarely revisited. TTLs are set once and then forgotten. Caches get reused for contexts where staleness has very different costs. And most systems don&#8217;t track how stale cached data actually is relative to the source of truth.</p><h4><strong>Replica divergence:</strong></h4><p>Replication lag is tracked, but usually against operational thresholds, not correctness requirements. A system might treat <em>500ms</em> lag as acceptable because it keeps dashboards green, not because anyone evaluated what that inconsistency means in practice.</p><p>Lag also isn&#8217;t constant. It increases under load, spikes during recovery, and varies across replicas. Most monitoring shows averages or current values, not worst-case behavior. The result is predictable: the system appears stable, while correctness degrades in bursts.</p><h4><strong>Event ordering:</strong></h4><p>Asynchronous systems do not guarantee ordering unless explicitly designed to. Events that are logically sequential can arrive out of order. If consumers process them without enforcing dependencies, the resulting state becomes inconsistent.</p><p>For example, an update event arrives before a create event. The update is ignored or partially applied. No exception necessarily occurs. The system continues, but the state is wrong.</p><h4><strong>Schema drift:</strong></h4><p>Over time, producers and consumers evolve independently. New fields are added but never consumed. Old fields disappear quietly. Data types change. Most of these changes do not cause failures. They degrade meaning.</p><p>Individually, these changes seem minor. Over time, across many services, they become a major source of data inconsistency. Nothing breaks. Everything slowly diverges.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/scaling-to-millions-part-1-the-system?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/scaling-to-millions-part-1-the-system?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/scaling-to-millions-part-1-the-system?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2><strong>The Signals You&#8217;re Probably Ignoring</strong></h2><p>The signals exist in every system experiencing this kind of correctness decay. They&#8217;re just not instrumented, not routed to the right people, or not recognized as systemic rather than incidental.</p><p>User complaints are often the most reliable early warning system, and the one engineering teams are slowest to act on. When error rates are flat, and latency is normal, but support volume is rising around a specific feature, something is wrong in a way your instrumentation doesn&#8217;t capture.</p><p>This is usually misclassified as a UX issue, a communication problem, or an isolated incident. It is rarely treated as a systemic correctness failure worth measuring.</p><p>Part of the reason is that user complaints are noisy and qualitative. &#8220;My results look different&#8221; is difficult to translate into an alert. But patterns matter. When multiple users report the same symptom, stale data, inconsistent behavior, or changes that don&#8217;t persist,  you&#8217;re no longer dealing with noise. You&#8217;re looking at correctness failure at scale. The inability to alert on it doesn&#8217;t make it less real.</p><p>Retry patterns in otherwise healthy systems are another signal that goes unnoticed. If clients are retrying requests that returned 200, something is wrong. Either the response was incorrect, the operation didn&#8217;t complete despite being acknowledged, or the client has some way of detecting that the result isn&#8217;t valid.</p><p>Most retry monitoring focuses on failure cases: 5xx responses, timeouts, and network errors. Retries on successful responses are rarely tracked, and almost never analyzed for what they imply about correctness.</p><p>Response variance across replicas is one of the most direct signals of consistency degradation, and one of the least measured. If the same logical request produces different results depending on which replica handles it, that divergence is your consistency window made visible.</p><p>Measuring it requires intentional design. For example, issuing the same request to multiple replicas in parallel and comparing responses in shadow mode. Very few systems implement this. It requires treating consistency as something to observe, not something to assume.</p><p>Data freshness is another blind spot. Replication lag is tracked. Cache TTLs are configured. But the actual age of the data in a response,  measured from when the underlying record last changed,  is almost never surfaced.</p><p>That is the metric that most directly answers a simple question: is this response still true? - It is missing from most dashboards.</p><p>Monitoring tells you when the system is down. It does not tell you when it is wrong. The gap between those two widens over time. Monitoring investment tends to focus on known failure modes, while system complexity keeps growing. New services, new async flows, and new caches each introduce new paths for data to diverge from reality.</p><p>None of them comes with correctness instrumentation by default.</p><div><hr></div><h2>Designing for Truth, Not Just Throughput</h2><p>The shift required here is specific and consequential. Most architecture conversations about scale start from two questions: is the system available, and is it fast? Both are legitimate. Neither is sufficient. The right first question is: are the results still correct?</p><p>That reordering changes what gets designed, what gets measured, and what gets prioritized when there&#8217;s a conflict. It means defining correctness properties explicitly before choosing replication topology instead of after. It means having SLOs for data freshness and consistency, not just for latency and error rate. It means treating the gap between availability failures and correctness failures as an instrumentation problem that needs to be solved, not an acceptable ambiguity.</p><p>Correctness SLOs are uncommon but not complicated to define once you take them seriously. What is the maximum acceptable age of data in a given response? What is the upper bound on the window during which two users can get different answers to the same query? What is the tolerable frequency of an event being processed out of order or processed more than once? These aren&#8217;t abstract questions. They&#8217;re design constraints that belong in a design document, reviewed by the team, and monitored in production.</p><p>The reason they&#8217;re rare is that they require making trade-offs explicit that most teams prefer to leave implicit. If you define a staleness SLO, you are committing to measuring staleness, which means you have to build the instrumentation, which means you are accountable for it. Leaving the trade-off implicit is more comfortable until it produces a correctness failure significant enough to become visible.</p><p>Making the trade-off explicit also changes how you reason about where consistency matters and where it can be relaxed. Eventual consistency in a user profile service is a reasonable engineering choice. The name showing up stale for 30 seconds has limited impact. Eventual consistency in an inventory system where two users can both reserve the same last item is not a reasonable choice, and a correctness SLO makes that distinction visible. Without it, both decisions get made the same way by default, by whoever set up the infrastructure with no visibility into the difference in consequence.</p><p>The distinction between availability failures and correctness failures also matters for how you run incident response and postmortems. A correctness failure that serves wrong data to ten percent of users for two hours is a significant incident. It will almost never trigger an automated alert under most monitoring configurations, and it will rarely make it into a postmortem unless a user complaint or a manual audit surfaces it. An availability failure that takes the service down for five minutes will trigger everything. The asymmetry in how these two failure modes are treated is a direct product of what gets measured and that&#8217;s a choice.</p><p>Correctness has to be designed in. It can&#8217;t be monitored in after the fact. You can observe latency and availability on a system that wasn&#8217;t designed with them in mind, and the metrics will still be meaningful. You cannot observe correctness on a system that wasn&#8217;t designed with explicit consistency properties, because the instrumentation requires understanding what &#8220;correct&#8221; means for each operation &#8212; a design-time decision that can&#8217;t be retrofitted cleanly.</p><blockquote><p>Scale is an engineering problem. Correctness is a design problem.</p></blockquote><div><hr></div><h2>What&#8217;s Coming in Part 2</h2><p>Part 1 is the diagnosis. The system appears healthy by every metric on the dashboard. It is not. It&#8217;s returning incorrect, inconsistent, or stale data at a rate the monitoring can&#8217;t see, and the gap between what the system reports and what users experience grows as the system adds load, adds services, and adds complexity.</p><p>Part 2 is the architecture. Specifically: how do you build systems that remain correct under continuous change, at scale, without paying the throughput cost of strong consistency everywhere?</p><p>That means idempotency guarantees in practice, not the simplified version where you hash the request and check a table, but the version that holds when consumers restart mid-batch, when clocks skew, and when the same logical operation arrives from two different upstream paths simultaneously.</p><p>It means ordering guarantees in async systems: what they actually require from the infrastructure, where they&#8217;re achievable without sacrificing throughput, and where accepting unordered delivery is the right trade-off if you design the consumer to be commutative.</p><p>It means cache invalidation is done at the correctness boundary rather than the expiry boundary, which requires knowing what the correctness boundary is, a design decision that most teams skip.</p><p>And it means making deliberate, documented decisions about where inconsistency is acceptable and structuring the system so those boundaries are explicit, enforced, and observable rather than accidental side effects of infrastructure defaults.</p><p>The goal isn&#8217;t consistency everywhere. That&#8217;s neither achievable at scale nor worth pursuing. The goal is knowing exactly where you&#8217;ve accepted inconsistency, being able to demonstrate that the decision was intentional, and having the instrumentation to know when the system is violating its own stated trade-offs. Scaling is easy to demonstrate. Staying correct is harder to prove.</p>]]></content:encoded></item><item><title><![CDATA[BDD(Behavior-Driven Development) Breaks in Distributed Systems]]></title><description><![CDATA[Given/When/Then assumes a deterministic world. Distributed systems don't have one.]]></description><link>https://nidly.substack.com/p/bddbehavior-driven-development-breaks</link><guid isPermaLink="false">https://nidly.substack.com/p/bddbehavior-driven-development-breaks</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 04 May 2026 07:54:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!VKFR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!VKFR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!VKFR!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!VKFR!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!VKFR!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!VKFR!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!VKFR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg" width="1024" height="706" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:706,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:127414,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/195274792?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!VKFR!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!VKFR!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!VKFR!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!VKFR!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6398603-11bf-4489-bc8d-30eee0634eec_1024x706.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Here is a scenario that looks perfectly reasonable:</p><pre><code><code>Given a user has $100 in their account  
When they transfer $60 to another user  
Then their balance should be $40  
</code></code></pre><p>Clean. Readable. Everyone agrees. Stakeholders understand it, QA signs off, and it passes locally every time. <strong>Then you run it in a real system.</strong> There&#8217;s a message broker. Two payment services. A database replicating across availability zones. Now it starts failing. Not consistently, maybe one in twenty runs.</p><p>Sometimes the balance is still $100 because the read hit a replica that hasn&#8217;t caught up yet. Sometimes the transfer runs twice because a consumer retried after a timeout. Sometimes the API returns <code>200 OK</code>, but the debit hasn&#8217;t happened yet; it&#8217;s still in a queue. Nothing about the business rule changed. The logic is still correct.</p><blockquote><p>What broke is the meaning of <strong>&#8220;Then&#8221;</strong>.</p></blockquote><h2>What BDD Was Actually Selling</h2><p>BDD came from a real problem. Unit tests could verify code, but they couldn&#8217;t tell if the system matched business intent. Dan North&#8217;s idea was to close that gap by describing behavior in a way everyone could read and agree on.</p><p>The goal wasn&#8217;t test coverage. It was a shared understanding. And in the systems BDD grew up in, that worked. A request comes in, something happens, and a response goes out. The system is mostly synchronous, and side effects are immediate.</p><p>When you say, <em>&#8220;Then the balance should be $40,&#8221;</em> you mean something very concrete: the request finishes, you check the database, and the value is there. Simple. Direct. Mostly true. If your system still behaves like that, BDD holds up fine.</p><blockquote><p><strong>The problems start when you carry that assumption into systems that don&#8217;t.</strong></p></blockquote><p>In practice, BDD scenarios drift. They grow. Someone maintains them. They pass in CI and fail in staging. Eventually, someone adds a <code>sleep(500)</code> to &#8220;stabilize&#8221; things. At that point, the scenarios are no longer describing the system. They&#8217;re describing a simplified version of it, one that behaves nicely, responds immediately, and never surprises you.</p><blockquote><p><strong>That version doesn&#8217;t exist in production.</strong></p></blockquote><p>Now you&#8217;re paying the cost of maintaining human-readable specs and getting false confidence in return.</p><div><hr></div><h2>The Determinism That BDD Requires</h2><p>Every Given/When/Then scenario makes an implicit contract: given this exact state, when this exact thing happens, then this exact outcome will follow.</p><p>That contract depends on determinism. The system has to be in a known state before the action, the action has to produce a predictable sequence of effects, and the result has to be observable at the moment you assert it.</p><p>This is not an implementation detail. It is structural. The format itself has no way to express statements like &#8220;Then the balance will eventually be $40&#8221; or &#8220;Then at least one of these outcomes will happen.&#8221; It assumes a single snapshot in time and assumes that the snapshot is meaningful.</p><p>In a monolithic system with synchronous execution and strong transactional guarantees, that assumption holds. In most systems being built today, it does not.</p><div><hr></div><h2>How Distributed Systems Actually Work</h2><p>A distributed system does not process a request. It processes a chain of events, partial states, concurrent operations, and compensating actions that together produce an outcome over time. That outcome may not be visible immediately, and in some cases may not become visible at all.</p><p>Consider the same payment scenario. A user submits a transfer. The API validates the request and publishes a <code>transfer_requested</code> event. A service consumes that event and debits the sender, then emits a <code>sender_debited</code> event. Another service credits the recipient. A third service updates a read model that the balance API queries.</p><p>Each step can fail independently. Each can be retried. Each can execute out of order. The read model can lag behind the write path by milliseconds or seconds, depending on load. At what point is the &#8220;Then&#8221; evaluated?</p><ul><li><p>Immediately after the HTTP response, when the transfer is still in-flight</p></li><li><p>After a fixed delay, which is effectively a guess</p></li><li><p>By polling until the condition becomes true, which turns the scenario into a different kind of test</p></li></ul><p>Partial failure makes this harder. In a real system, success is not binary. The debit may succeed while the credit fails. A saga may roll the debit back. A dead-letter queue may hold the credit event for later recovery. From the outside, the request returns 200. Inside the system, the money may be in an intermediate state that no single BDD scenario can describe.</p><p>Concurrency introduces another layer of complexity. Two transfers initiated at the same time against the same account may both read the same balance, both pass validation, and both attempt to debit, resulting in an overdraft that no isolated scenario would predict.</p><blockquote><p>BDD evaluates scenarios in isolation. Distributed systems fail through interaction.</p></blockquote><p>Eventual consistency is not a bug. It is a deliberate tradeoff that allows systems to scale. BDD, however, assumes strong consistency at the moment of assertion. These two assumptions are fundamentally incompatible.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2><strong>The Gap Between Scenario and System</strong></h2><p>When BDD scenarios describe a distributed system, they tend to fall into one of three patterns. All of them are problematic. The <strong>first</strong> is abstraction. The scenario describes a system that does not exist.</p><p>It talks to a facade that <em>synchronously</em> resolves all the asynchronous work before returning. The test passes, but it is not testing the real system. It is testing a <em>simulation</em> that removes exactly the properties you need to be confident about in production. The <strong>second</strong> is brittleness.</p><p>The scenarios exercise the real system, but rely on timing assumptions, polling, or retry logic hidden inside step definitions. They start to flap. Failures show up with no clear cause. Someone investigates, finds no real bug, and marks it as <em>flaky</em>. Over time, the suite stays green, but nobody actually trusts it. The <strong>third</strong> is incompleteness.</p><p>The scenarios cover the happy path and a few obvious errors, but the <em>interesting failure modes</em> never make it in. Nobody writes something like:</p><blockquote><p>Given a transfer is in-flight when the credit service goes down<br>Then the saga coordinator retries the credit within 30 seconds<br>And the read model is consistent within 60 seconds</p></blockquote><p>That scenario is both <em>true</em> and <em>important</em>. It is also awkward to write, hard to automate, and not readable for non-technical stakeholders. So it gets skipped. The format pushes you toward scenarios that are <em>easy to express in BDD</em>, not scenarios that actually capture the <strong>risk surface</strong> of the system.</p><div><hr></div><h2><strong>AI Systems Make This Untenable</strong></h2><p>If distributed systems stretch BDD to its limits, AI systems break it completely. A large language model does not produce deterministic output. Given the same input, it can return different responses across runs. A recommendation system shifts as models are retrained. An anomaly detection system changes behavior as the definition of &#8220;normal&#8221; evolves. So what does a BDD scenario even mean here?</p><pre><code><code>Given a user asks "What is the capital of France?"
When the model processes the request
Then the response should be "Paris"</code></code></pre><p>This might pass today. It might fail after the next fine-tuning run. Even if it passes, it tells you almost nothing about the system&#8217;s quality. The real questions are not about exact answers. They are about <em>distributions</em>:</p><ul><li><p>How often is the model factually wrong?</p></li><li><p>How often does latency exceed acceptable thresholds?</p></li><li><p>How does quality degrade at the edges of the input space?</p></li></ul><p>None of these fit into a Given/When/Then structure. The same issue appears in any system driven by embeddings, feedback loops, or learned features. The output is no longer just a function of the input. It depends on:</p><ul><li><p>model weights</p></li><li><p>training data</p></li><li><p>feature pipelines</p></li><li><p>accumulated distribution shift</p></li></ul><p>At that point, an assertion like <em>&#8220;Then the recommendation should be X&#8221;</em> is no longer meaningful.</p><div><hr></div><h2><strong>What Actually Works</strong></h2><p>None of this means you stop specifying system behavior. It means you choose tools that match the <em>actual properties</em> of the system you are building.</p><p><strong>Invariants over exact outcomes.</strong><br>Instead of asserting that a specific balance is reached, assert what must <em>always</em> hold. The sum of all balances remains constant. No account goes below zero without an overdraft agreement. Every transfer eventually results in either a completed credit or a compensating action. These statements survive timing, retries, and partial failure. They describe the system as it really behaves, not as we wish it behaved.</p><p><strong>Contract testing at boundaries.</strong><br>In distributed systems, the useful question is not &#8220;does the system do X end-to-end,&#8221; but &#8220;does service A honor what service B depends on.&#8221; Tools like Pact let you verify these contracts independently. This gets you closer to BDD&#8217;s original goal, <em>shared understanding</em>, without requiring unrealistic end-to-end determinism.</p><p><strong>Tolerance windows instead of exact assertions.</strong><br>Eventual consistency forces you to define <em>how eventual is acceptable</em>. Saying &#8220;the read model reflects the write within 500ms under normal load&#8221; is a real requirement. Saying &#8220;the balance is immediately updated after the request&#8221; is not, at least not in a system built on asynchronous replication.</p><p><strong>Chaos and property-based testing.</strong><br>The interesting scenarios are not happy paths. They are failure modes. What happens when a consumer is slow? When is a service unavailable? When messages are duplicated? Property-based testing lets you say &#8220;this invariant must always hold&#8221; and then actively search for the inputs that break it.</p><p><strong>Observability as a requirement.</strong><br>In production, systems are understood through metrics, traces, and logs, not pass/fail assertions. Being able to detect &#8220;this saga failed to complete within its SLA&#8221; is often more valuable than a suite of BDD scenarios that only pass when everything behaves nicely.</p><p><strong>For AI systems, evaluation replaces assertion.</strong><br>You define metrics. You build datasets that represent real inputs. You track how the system performs over time. A regression is not &#8220;this scenario failed.&#8221; A regression is &#8220;quality dropped 4% this week, concentrated in this category.&#8221;</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/bddbehavior-driven-development-breaks?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/bddbehavior-driven-development-breaks?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/bddbehavior-driven-development-breaks?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2><strong>The Right Tool for the Complexity You Have</strong></h2><p>BDD is not a bad idea. For systems with clear business rules, synchronous flows, and a need for shared understanding between product and engineering, it still works. It forces useful conversations. It creates readable documentation.</p><blockquote><p>The problem is not BDD itself. It is where we apply it.</p></blockquote><p>A tool that works for a monolith does not automatically work for a system of multiple services communicating over a message bus. A testing model built on deterministic, synchronous assertions does not translate cleanly into a system where timing varies, outcomes are probabilistic, and failure is expected.</p><p>Engineers who have lived through this recognize the pattern. The BDD suite grows. Flakiness grows with it. Trust in the tests declines. Eventually, the tests remain in CI, but nobody really believes them.</p><p>And when production incidents happen, they happen <em>between</em> scenarios. In the failure modes nobody wrote down. The issue is not that we lack better tools. We already have them: invariant testing, contract testing, chaos engineering, property-based testing, SLOs, and observability. The issue is inertia.</p><p>BDD worked in simpler systems, so it keeps getting applied even when the system has outgrown it. It feels familiar. The tooling is mature. It is easy to keep doing the same thing. But every hour spent maintaining scenarios that assume a synchronized, deterministic system is an hour not spent understanding how the system actually fails. That is what breaks in distributed systems.</p><p>Not BDD as a format, but the assumption behind it: that behavior can be fully captured as a set of discrete, deterministic, human-readable scenarios. In systems where behavior is <em>emergent</em>, <em>time-dependent</em>, and often <em>probabilistic</em>, that assumption does not hold. And building confidence on something that does not hold is how you end up with a green build and a production incident.</p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Why Clean Architecture Breaks Down at Scale, Part 2: Architecture for Change, Not Clarity]]></title><description><![CDATA[How to design systems that stay adaptable when change is the only constant.]]></description><link>https://nidly.substack.com/p/why-clean-architecture-breaks-down-57c</link><guid isPermaLink="false">https://nidly.substack.com/p/why-clean-architecture-breaks-down-57c</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Mon, 27 Apr 2026 06:12:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Im_A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Im_A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Im_A!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png 424w, /__u/substackcdn.com/image/fetch/$s_!Im_A!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png 848w, /__u/substackcdn.com/image/fetch/$s_!Im_A!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Im_A!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Im_A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png" width="1456" height="867" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:867,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2473882,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/194799487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Im_A!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png 424w, /__u/substackcdn.com/image/fetch/$s_!Im_A!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png 848w, /__u/substackcdn.com/image/fetch/$s_!Im_A!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Im_A!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8de57a7-377d-4ff8-abfe-d39b790c84c4_1625x968.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><a href="/__u/nidly.substack.com/p/why-clean-architecture-breaks-down?r=a3p8i">Part 1</a> ended with a diagnosis. <a href="https://blog.cleancoder.com/uncle-bob/2012/08/13/the-clean-architecture.html">Clean Architecture</a> doesn&#8217;t fail loudly. It slows down; Changes take longer than they should, and small requests start touching more of the system than expected. Boundaries multiply, abstractions age, and indirection turns simple decisions into long paths.</p><p>Nothing is technically broken. The system still works. It just becomes harder to change. That&#8217;s the important part. The problem isn&#8217;t that the architecture is wrong. It&#8217;s that it was designed for a different environment, one where structure is relatively stable, and change is incremental. Most real systems don&#8217;t stay in that state for long.</p><p>So the question isn&#8217;t what breaks. That part is already clear. The question is what actually holds when change becomes constant. Not a new pattern, and not a cleaner version of the same idea. Something more practical: a way to decide where boundaries should exist, how much structure is enough, and when it starts getting in the way.</p><p>The argument here is simple. Architecture should optimize for rate of change, not structural clarity. That doesn&#8217;t mean structure doesn&#8217;t matter. It means the structure has to follow how the system evolves, not how it looks on a diagram.</p><div><hr></div><h3>Stable vs. Volatile Systems</h3><p>The most useful observation about any large system is that it doesn&#8217;t change uniformly. Some parts barely move. The core domain model of an e-commerce system, order total calculations, inventory reservation flows. These tend to evolve slowly. The business has strong opinions about them. They&#8217;ve been debated, refined, and stabilized over time. They&#8217;re load-bearing.</p><p>Other parts don&#8217;t sit still at all. Feature experiments run for a few weeks and disappear. Third-party integrations shift every time a vendor changes an API. Notification flows get rewritten with every product decision. The frontend surface changes faster than anything beneath it.</p><blockquote><p>The mistake Clean Architecture invites is treating these very different parts with the same level of structural seriousness.</p></blockquote><p>If you draw equally rigid boundaries around your checkout core and an experimental loyalty feature, you&#8217;re applying the same cost model to two completely different kinds of change. The checkout core benefits from careful layering, explicit contracts, and conservative evolution. The loyalty feature needs the opposite. It needs room to move fast, to be wrong, and to be deleted without ceremony.</p><p>When a system is small, this distinction doesn&#8217;t hurt much. But as systems grow and teams multiply, it starts to matter in a very real way. Over-engineering volatile parts creates drag. Under-engineering the stable core creates risk. Treating them the same creates both at once.</p><p>The design question isn&#8217;t where to draw all the boundaries. It&#8217;s which ones deserve to be treated as long-term decisions, and which ones should stay loose enough to change without a migration plan.</p><div><hr></div><h3>Boundaries Based on Change, Not Layers</h3><p>Layer-based thinking is intuitive. Controllers receive requests, use cases execute logic, entities represent the domain, and repositories handle persistence. Separate them cleanly and you get a system that&#8217;s easy to explain. Every engineer knows where to go for a given type of change. The problem is that layers don&#8217;t match how systems actually evolve.</p><p>A small product change, like adding a field to a user profile, ends up crossing every layer. Not because each layer has something meaningful to say about that field, but because the architecture requires it. The change is forced to move through multiple checkpoints, even when most of them are just passing data along. That&#8217;s where the cost shows up. Not in complexity, but in traversal.</p><blockquote><p>An alternative is to draw boundaries where change actually concentrates.</p></blockquote><p>Features, flows, and capabilities tend to evolve together. A checkout flow changes when payment requirements change. A notification system shifts with the communication strategy. A reporting surface evolves with analytics needs. These are not layer concerns. They are cohesion boundaries. Code that changes together should live together. This leads to a different kind of structure.</p><p>Instead of grouping controllers, use cases, and repositories by their technical role, you group everything around a feature. The checkout flow owns its input handling, its business logic, its data access, and its API contract. A change to checkout stays within that boundary. It doesn&#8217;t require navigating a global structure.</p><p>This isn&#8217;t a new idea. It shows up as vertical slices, feature modules, or flow-based design. What matters is not the name, but the outcome. Fewer handoffs. Shorter paths. Less coordination for small changes.</p><p>That said, not everything fits this model. Some concerns are genuinely cross-cutting. Authentication, observability, and core domain concepts often need shared abstractions. But a lot of what gets labeled as cross-cutting is just incidental similarity, code that looked reusable at the time.</p><p>Feature-based organization forces a harder question. Is this shared because it must be, or because it was convenient?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Shortening the Path of Change</h3><p>Every layer a change must travel through is a tax. Sometimes that tax is worth paying. A layer might enforce a real constraint, a consistency check, a security boundary, or a validation rule. But often, the tax exists for a different reason. The architecture requires the layer, even when the change doesn&#8217;t.</p><blockquote><p>In practice, the most expensive layers are the ones that only translate.</p></blockquote><p><a href="https://martinfowler.com/eaaCatalog/dataTransferObject.html">DTO</a>s mapping domain objects to presentation models. Mappers converting persistence records into entities. Response builders reshaping data into slightly different forms. These layers rarely add logic. They add distance.</p><p>The usual argument is insulation. If each layer has its own types, changes in one don&#8217;t ripple into others. That&#8217;s true in theory. In practice, most changes still propagate. A domain concept evolves, and now every representation of it needs to be updated, along with every mapper in between. You end up paying an ongoing cost to protect against a kind of change that rarely happens independently.</p><p>The real distinction isn&#8217;t between layers and no layers. It&#8217;s between meaningful translation and mechanical translation.</p><p>If a field carries the same meaning from database to domain to API, renaming and remapping it at every step doesn&#8217;t make the system cleaner. It just makes the path longer. Letting it move through as-is isn&#8217;t careless design. It&#8217;s recognizing that structure has a cost, and only paying it when it buys something real.</p><p>There is a trade-off here, and it&#8217;s worth stating directly. Fewer translation layers mean more visible coupling. If domain objects flow into API responses, changes in the domain can affect the API. That risk is real. But in many systems, that coupling already exists, just hidden behind layers that still require coordinated updates.</p><p>The question isn&#8217;t whether coupling exists. It&#8217;s whether you want it to be explicit and manageable, or indirect and disguised.</p><div><hr></div><h3>Duplication vs. Abstraction</h3><p>Avoiding duplication is treated like a rule of good engineering. Breaking it feels wrong. The problem is that most bad abstractions start as an attempt to remove duplication too early.</p><p>Abstraction is a bet on the future. It assumes that two pieces of code are the same in all the ways that will matter later. In fast-moving systems, that assumption is usually wrong.</p><p>Two email flows might look identical today. A few months later, one needs retries, another needs templating, and a third needs compliance logging. The abstraction that unified them becomes the place where all differences have to fit. What started as simplification turns into constraint. The alternative is delayed abstraction.</p><p>Write it twice. Let it evolve. See how it changes under real use. When a stable pattern actually emerges, extract it then. The duplication costs you a bit of readability in the short term. The abstraction you get later is grounded in reality, not prediction.</p><p>This shows up in practice more often than people admit. Two services need to send transactional emails. The natural instinct is to build a shared notification service. It looks clean. It feels right. A few months later, requirements diverge. One flow needs delayed delivery, another needs strict ordering, another needs tracking and reporting.</p><p>Now you have two options. Expand the shared service with conditionals until it barely resembles its original shape, or bypass it and fork the logic locally. Either way, the abstraction stops helping.</p><p>Starting with separate implementations would have been simpler. When the behavior diverged, the code would have diverged with it. No shared contract to renegotiate. No coordination overhead. Just two pieces of code evolving independently. The system ends up larger, but easier to reason about.</p><p>That doesn&#8217;t mean duplication is always the right choice. Some things do need to be centralized. Security logic, billing rules, and data consistency boundaries are not places to experiment. The point is more subtle.</p><p>Duplication is not always a failure. Sometimes it&#8217;s the safer bet. The real mistake is treating abstraction as a default instead of a decision.</p><div><hr></div><h3></h3><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/why-clean-architecture-breaks-down-57c?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/why-clean-architecture-breaks-down-57c?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/why-clean-architecture-breaks-down-57c?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><h3>Rethinking Replaceability</h3><p>Part 1 described a notification system that looked replaceable in code but wasn&#8217;t in practice. The interface hid the vendor, but the rest of the system didn&#8217;t. Analytics depended on specific event formats. Operations relied on the provider&#8217;s UI. Finance workflows were built around its billing model. Swapping the implementation turned into months of work that had little to do with code.</p><p>That&#8217;s the gap. Replaceability at the code level is a much weaker guarantee than it appears. When you design something to be &#8220;replaceable,&#8221; you&#8217;re designing its interface. What you&#8217;re not designing is everything that grows around it over time. Dashboards assume certain metrics. Alerts depend on specific behaviors. Teams build intuition about how the system fails and recovers. None of that lives in the interface.</p><p>So when the implementation changes, the code may compile, but the system still shifts under your feet. Real replaceability looks different. It&#8217;s less about clean interfaces and more about being able to observe, compare, and migrate behavior safely.</p><p>That means designing for change at the operational level. Running dual writes during migration. Using shadow reads to compare results. Defining metrics that are stable enough to detect behavioral drift across implementations.</p><p>These are not abstractions. They are capabilities. They acknowledge that replacing a component is not just a code change, but a system-level transition. It also means being honest about what replaceability actually buys you.</p><p>If you&#8217;ve been running on a payment provider for years, you&#8217;ve accumulated knowledge that doesn&#8217;t live in code. Edge cases, incident patterns, performance quirks. Switching vendors means losing that familiarity and rebuilding it under pressure. The new system will behave differently in ways you don&#8217;t expect. No interface protects you from that.</p><p>The real question isn&#8217;t whether a component is replaceable. It&#8217;s whether your system is observable enough to detect the differences, and whether your team can respond before those differences turn into incidents.</p><div><hr></div><h3>Managing Indirection</h3><blockquote><p>Indirection is useful when it reflects real variability.</p></blockquote><p>An abstraction over database access makes sense if you actually support multiple storage backends, or need to isolate tests from a real database. An abstraction over email providers makes sense if you genuinely switch between them. In these cases, the indirection earns its cost because it manages something real.</p><p>The problem is that indirection rarely stays tied to that reality.</p><p>Interfaces outlive the variability they were introduced for. Factories continue to exist long after there is only one implementation left. The code still works, but the indirection stops doing useful work. It remains because removing it feels like removing safety.</p><p>The cost is not obvious at first. It shows up as distance. You&#8217;re debugging unexpected behavior. You find the interface, then the implementation, then the factory that selects it, then the configuration that drives the factory. Several steps in, you still haven&#8217;t reached the logic that actually matters.</p><p>The answer is there. It&#8217;s just far away. That distance accumulates. Engineers stop tracing behavior from first principles and start relying on familiarity with the structure. Understanding becomes slower, more dependent on context, and harder to transfer. When the people who &#8220;know the system&#8221; leave, that cost becomes visible.</p><blockquote><p>Local reasoning is an underrated design goal.</p></blockquote><p>If someone can open a file and understand how a feature behaves without jumping across the system, that has real value. Often more than the theoretical benefit of being able to swap an implementation that hasn&#8217;t changed in years.</p><p>Indirection should follow variability. When the variability disappears, the indirection should go with it. Keeping it around for the sake of consistency turns structure into overhead.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Ownership and Team Alignment</h3><p>Conway&#8217;s Law is often framed as a statement about management. It&#8217;s really about information.</p><p>Teams build systems that mirror how they communicate because people make better decisions about what they understand, and they understand best what they own end to end. Layer-based team splits work against this.</p><p>If one team owns controllers and another owns use cases, no team fully owns a feature. A change to something like user onboarding has to move across multiple teams because the logic is split across layers. The coordination cost isn&#8217;t accidental. It&#8217;s the direct result of aligning ownership with architecture instead of with the product.</p><p>End-to-end ownership looks different. A team owns a slice of the system: its data, its logic, its API, and its behavior in production. Most changes stay within that boundary. Decisions are faster because they don&#8217;t require constant coordination. Understanding is deeper because the team can trace behavior without crossing ownership lines.</p><p>That doesn&#8217;t come for free. Teams can drift. Local decisions can diverge. Shared infrastructure still requires coordination. But the coordination that remains is real. It&#8217;s about genuinely shared concerns, not routing a small change through multiple teams.</p><blockquote><p>Architecture shapes how teams work.</p></blockquote><p>If you want independent teams, your system has to allow independence. If small changes regularly cross ownership boundaries, the system will be slow no matter how strong the engineers are.</p><div><hr></div><h3>Practical Heuristics</h3><p>A few signals that are reliable enough to act on:</p><ul><li><p>If a simple change crosses more than two ownership boundaries, the boundaries are in the wrong place.</p></li><li><p>If an abstraction needs more than a sentence to explain, it has likely outlived its usefulness.</p></li><li><p>If two pieces of code look similar but serve different capabilities, don&#8217;t rush to unify them. Let them evolve. Abstract later if they converge.</p></li><li><p>If you can&#8217;t understand a component&#8217;s behavior by reading its file, the indirection has gone too far.</p></li><li><p>If engineers are shaping their changes to avoid crossing boundaries instead of expressing intent clearly, the architecture is in the way.</p></li></ul><p>These are not strict rules. They&#8217;re early warnings. Ignore them long enough and they stop being warnings.</p><div><hr></div><h3>Conclusion</h3><p>Architecture isn&#8217;t something you design once and then live inside. It&#8217;s a continuous set of decisions about where to add friction and where to remove it.</p><p>Friction belongs where stability matters. It protects the slow-moving parts of the system. It does not belong in the path of changes that happen every day.</p><p>The shift from clarity to change isn&#8217;t about new patterns. It&#8217;s about being honest about how your system behaves over time.</p><p>Where does it change most?<br>How often do those changes happen?<br>Who pays the cost when change is slow?</p><p>Those questions matter more than any diagram. Systems that survive are not the ones that predicted change correctly. They&#8217;re the ones that assumed they wouldn&#8217;t, and designed for that reality.</p><p>They keep stable parts protected, volatile parts cheap to modify, and boundaries aligned with how the system actually evolves.</p><blockquote><p>The goal was never a clean diagram.</p></blockquote><p>It was a system you can still change, confidently, years after you built it.</p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Don’t Waste 2026 on the Wrong Career: ML vs AI Engineer]]></title><description><![CDATA[Most engineers are learning both and mastering neither. Here&#8217;s the real difference and how to pick one path that actually leads somewhere.]]></description><link>https://nidly.substack.com/p/dont-waste-2026-on-the-wrong-career</link><guid isPermaLink="false">https://nidly.substack.com/p/dont-waste-2026-on-the-wrong-career</guid><dc:creator><![CDATA[Alireza Rahmani Khalili]]></dc:creator><pubDate>Wed, 22 Apr 2026 06:03:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hLDL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hLDL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hLDL!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!hLDL!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!hLDL!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hLDL!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_webp, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hLDL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1216081,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://nidly.substack.com/i/194725168?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!hLDL!, /__u/nidly.substack.com/w_424, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!hLDL!, /__u/nidly.substack.com/w_848, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!hLDL!, /__u/nidly.substack.com/w_1272, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hLDL!, /__u/nidly.substack.com/w_1456, /__u/nidly.substack.com/c_limit, /__u/nidly.substack.com/f_auto, /__u/nidly.substack.com/q_auto:good, /__u/nidly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc3a23f-f549-4dda-9407-177f65fbaf94_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There&#8217;s a pattern showing up in engineering circles right now. A backend engineer with a few years of real experience decides to move into AI and starts doing what seems reasonable. A course on neural networks, a few transformer papers, some PyTorch tutorials, and at the same time, small experiments with LLM APIs and maybe a basic <a href="/__u/nidly.substack.com/p/rag-isnt-about-embeddings?r=a3p8i">RAG</a> system. Individually, none of this is wrong. But six months later, the outcome is usually the same.</p><p>They know a little about training, a little about inference, a little about embeddings, and a little about everything else. But they are not competitive as an ML engineer, and they haven&#8217;t built anything solid as an AI engineer. There is no depth, no real leverage, and nothing they can confidently stand behind in an interview or in production. They end up stuck between two different disciplines, mistaking exposure for progress.</p><p>This is no longer a niche mistake. It is quickly becoming the default path for engineers who decide to &#8220;get into AI&#8221; without first deciding what that actually means. And the cost is not just time. It is six months of building the wrong instincts, the wrong mental models, and a completely distorted view of what the work actually looks like.</p><p>ML Engineering and AI Engineering are not two points on the same spectrum. They are different jobs with different skill sets, different constraints, and different definitions of success. Treating them as variations of the same thing is how you quietly lose a year without even realizing it</p><div><hr></div><h2><strong>The Real Split</strong></h2><p>Strip away the terminology, and the distinction is architectural. ML Engineering is about producing models. The work happens upstream of deployment. It involves collecting and cleaning training data, designing features, running experiments, and improving model performance over time. You are responsible for the quality of the model itself, which means you need to understand what is happening inside it, not just what comes out.</p><p>AI Engineering is about using models. The work happens downstream. It involves integrating models into systems, handling the data flowing through them, managing failure modes in production, and making sure the whole system behaves reliably under real-world constraints. The model is just one component, often a black box you don&#8217;t control.</p><p>The word &#8220;model&#8221; shows up in both roles, but the relationship to it is completely different. One side is shaping the model. The other is building everything around it.</p><div><hr></div><h2><strong>What ML Engineers Actually Do</strong></h2><p>The honest version of ML engineering looks less like programming and more like applied statistics mixed with data janitorial work.</p><p>The core artifact is the training pipeline. Building one means deciding how data is ingested, how it&#8217;s cleaned, how it&#8217;s transformed into features, how batches are constructed, and how experiments are tracked so results are reproducible. A large part of the job is debugging why a model that looked good in a notebook falls apart on real data. Most of the time, the issue isn&#8217;t the model. It&#8217;s that the real-world data distribution doesn&#8217;t match what you trained on.</p><p>Feature engineering hasn&#8217;t disappeared. For tabular problems, it&#8217;s still central. For language and vision, it shows up in different forms: tokenization choices, input formatting, data augmentation, and how training examples are constructed. These decisions look small, but they shape the outcome in ways you only see after running expensive experiments.</p><p>Evaluation is where things get difficult. A single metric on a test set is not enough. ML engineers spend time designing evaluation frameworks that reflect real-world behavior, looking at slices of performance, failure cases, and how the model degrades as data shifts over time. This requires statistical judgment, domain context, and a tolerance for being wrong in ways that are hard to detect.</p><p>The academic nature of the work is real. Not because you need to publish papers, but because you need to be comfortable with math, with reading research, and with understanding what&#8217;s happening under the surface. The bar is shaped by a candidate pool that often includes people with deep theoretical backgrounds, and that shows up in how problems are discussed and evaluated.</p><p>None of this makes the path impossible. People do transition into ML engineering without formal degrees. But it takes longer than most expect, the depth is often underestimated, and the competition at serious ML-driven companies is real.</p><div><hr></div><h2><strong>What AI Engineers Actually Do</strong></h2><p>AI engineering is software engineering that happens to involve models. The right mental model is systems thinking, not statistics.</p><p>The core work is integration. You are dealing with a language model, an embedding model, a vector database, a retrieval layer, orchestration logic, and upstream data sources. Your job is to make this system behave reliably. That means getting the right context into the model, handling cases where retrieval fails, parsing structured output from something that speaks natural language, and doing all of it within real latency and cost constraints.</p><blockquote><p>Calling an LLM API is the visible part of the job. It&#8217;s also the easy part. Most of the effort sits around it.</p></blockquote><p>Retrieval systems make this obvious. Naive vector search gives you semantic similarity, but similarity is not the same as relevance. Hybrid approaches improve recall, but introduce new tuning problems. Add reranking, and now you&#8217;re operating a multi-stage retrieval pipeline where each step can fail quietly and errors compound across the system.</p><p>Production constraints dominate. Latency is user-facing. Cost is tracked. Edge cases don&#8217;t stay in staging; they show up in production. AI engineers spend time on prompt strategies, caching layers, fallback paths for malformed outputs, token budget management, and building observability that makes inference behavior understandable.</p><p>Most of this maps directly to backend engineering. Distributed systems intuition applies. API design patterns apply. Experience with queues, caching, and failure handling transfers cleanly. The model is just another dependency, with its own latency profile, failure modes, and cost curve.</p><div><hr></div><h2><strong>Why the Market Has Shifted</strong></h2><p>A few years ago, the dominant question was how to build better models. That question still matters, but it is now concentrated in a small number of organizations with the compute, data, and research talent to push it forward. For everyone else, the question has changed. It&#8217;s no longer about building models. It&#8217;s about using the ones that already exist.</p><p>This shift happened when model quality crossed a practical threshold. Not perfect, but good enough to power real products. GPT-4&#8211;class systems can solve useful problems. Open-weight models are competitive in many domains. That was enough to move the bottleneck.</p><blockquote><p>The constraint is no longer model capability. It&#8217;s integration.</p></blockquote><p>Companies don&#8217;t need another model. They need systems that can use models reliably, at scale, under real constraints. That means handling messy data, building retrieval layers, managing latency and cost, and making outputs usable in actual workflows.</p><p>That demand maps directly to AI engineering. For engineers with a system design background, this creates a structural advantage. The hard parts of the job are already familiar: service boundaries, data flow, observability, and failure handling. The model itself is just another dependency with its own quirks.</p><p>What&#8217;s new is learnable. Working with LLM APIs, building retrieval pipelines, and managing prompt logic takes effort, but it&#8217;s measured in months, not years. That difference matters more than people think.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2><strong>The Hard Truth About the ML Path</strong></h2><p>It is possible to become an ML engineer without a research background. People do it. But the path is narrower than it looks, and the effort required is higher than most expect.</p><p>The mathematical foundations are not optional. Linear algebra, probability, optimization, and information theory these are not decorative. They are part of the job. You don&#8217;t need to re-derive everything, but you need enough fluency to understand what a paper is actually saying, to interpret a loss curve beyond &#8220;it went down,&#8221; and to reason about why a model behaves the way it does.</p><p>Without that, you&#8217;re not really doing ML engineering. You&#8217;re just using tools without understanding their limits.</p><p>The competition reflects this. At companies doing serious ML work, you&#8217;re compared to people who have spent years on these problems. They&#8217;ve read the literature, built intuition, and developed a way of thinking that doesn&#8217;t come from a few months of courses. A backend engineer transitioning into ML is not on equal footing in that context, and ignoring that reality leads to a job search strategy that doesn&#8217;t match the market.</p><p>There are more accessible entry points. ML platform engineering, MLOps, evaluation and testing infrastructure, these are real roles with real impact. But they are different from model development. Treating them as stepping stones is fine. Treating them as the same job is not.</p><div><hr></div><h2><strong>The Practical Advantage of System Builders</strong></h2><p>The engineers best positioned for AI engineering right now are not the ones who studied language models the most. They are the ones who have built real systems in production.</p><p>A backend engineer who has designed ingestion pipelines, dealt with distributed consistency issues, and built monitoring for services that could not go down already has the right instincts. They understand latency in a way that is tied to users, not benchmarks. They have seen systems fail in production in ways that never showed up in staging. They know that when something breaks, the real problem is not the failure itself, but whether you can understand why it happened. This is not just helpful background. It is the core of the job.</p><p>The engineers who build reliable AI systems are not the ones who can explain attention mechanisms in detail. They are the ones who treat the model as a dependency and apply the same discipline they would apply to a database or an external service. They design around it, monitor it, and assume it will fail in ways that are not obvious.</p><p>The transition from backend engineering to AI engineering is shorter than it looks. Not because the work is easy, but because most of the required skills are already there. What needs to be learned sits at the application layer: how prompting behaves, how retrieval affects output quality, and how to evaluate model outputs in a structured way.</p><p>That learning curve is real, but it is different. It is about applying existing engineering instincts to a new type of system, not rebuilding your foundation from scratch.</p><div><hr></div><h2><strong>Daily Work, Concretely</strong></h2><p>The day-to-day difference between these roles is not subtle. An ML engineer&#8217;s work is structured around experiment cycles. You form a hypothesis about why performance is below target. Maybe the data underrepresents a class. Maybe the feature encoding is losing signal. Maybe regularization is off. You design an experiment, run it, wait, interpret the results, and refine your thinking. Progress is real, but slow. Iteration cycles are measured in hours or days, and the feedback loop is not tight.</p><p>An AI engineer operates much closer to production. Something breaks or behaves unexpectedly. Retrieval quality drops for a certain class of queries. Outputs degrade in edge cases. You inspect logs, trace how data flows through the system, analyze where the failure is happening, and decide whether the fix belongs in retrieval, prompting, or system logic. You ship a change and watch what happens in production. The loop is tighter, the feedback is immediate, and the work is more reactive.</p><p>Neither is better. They fit different ways of thinking. If you enjoy slow iteration, isolating variables, and converging on explanations, ML work will feel natural. If you care more about shipping, debugging real systems, and seeing immediate impact in production, AI engineering tends to be a better fit.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/dont-waste-2026-on-the-wrong-career?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading AI &amp; Software Engineering at Production Scale &#8212; A Newsletter! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://nidly.substack.com/p/dont-waste-2026-on-the-wrong-career?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/nidly.substack.com/p/dont-waste-2026-on-the-wrong-career?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2><strong>What a Real AI Engineering System Looks Like</strong></h2><p>Take a concrete example: a document Q&amp;A system for an internal knowledge base. The input is a corpus of documents such as product specs, runbooks, and design docs. The output is an answer grounded in those documents, ideally with references.</p><p>The engineering work starts before any model is involved. Documents need to be ingested, cleaned, and chunked. Chunking is not trivial. If chunks are too small, you lose context. If they are too large, relevance drops. Metadata needs to be preserved because it directly affects filtering and ranking. This is data engineering work.</p><p>Then comes retrieval. Each chunk is embedded and stored in a vector database. At query time, the user&#8217;s input is embedded and used to retrieve similar chunks. This works, but not perfectly. Semantic similarity is not the same as relevance. Hybrid search can help by combining dense and keyword-based retrieval, but it adds complexity. Reranking can improve results further, but now you have a multi-stage pipeline where each step can fail in ways that are not obvious.</p><p>The generation layer takes the retrieved context and the user&#8217;s query, builds a prompt, and calls the model. The output is not automatically trustworthy. It needs to be checked. Does it answer the question? Is it grounded in the provided context? Does it fabricate citations? Post-processing handles formatting, extraction of references, and fallback cases when no useful answer is found.</p><p>Evaluation is continuous. Not whether the system works in a demo, but whether it works across real queries. That requires a test set, a way to measure answer quality, and a feedback loop from user behavior back into the system.</p><p>Every part of this is software engineering. The model is just one step in a pipeline that only works because everything around it is designed carefully.</p><div><hr></div><h2><strong>What Automation Doesn&#8217;t Simplify</strong></h2><p>There&#8217;s a reasonable concern that AI engineering might be automated away as models get better. The answer is more structural than optimistic.</p><p>Real systems operate under constraints that no general-purpose model understands on its own. Latency budgets, cost limits, reliability requirements, and product expectations do not come from the model. They come from the system you are building. Someone has to translate those constraints into architecture decisions that actually hold in production. Integration complexity does not go away. It increases.</p><p>A system with multiple model calls, a retrieval layer, and a reranker has more failure paths than a simple pipeline. Failures do not show up cleanly. They propagate. A weak retrieval result leads to a bad prompt, which leads to a plausible but incorrect answer. The system does not crash. It just becomes wrong.</p><p>Handling that is engineering work. You need to understand where things break, how to isolate failures, and how to instrument the system so that problems are visible before users notice them.</p><p>Reliability and correctness get harder as more of the system becomes probabilistic. Managing that requires evaluation pipelines, output validation, fallback strategies, and continuous testing against real usage patterns. Better models do not remove this work. They make it more important.</p><div><hr></div><h2><strong>How to Choose</strong></h2><p>The decision comes down to a few honest questions. Choose ML engineering if you are genuinely interested in the modeling work itself. If reading papers on optimization is engaging, not a chore. If improving a metric through careful experimentation feels rewarding. And if you are willing to invest serious time in mathematical foundations, not just tooling.</p><p>Choose AI engineering if you care about building systems that actually work in production. If your instincts are around system design, data flow, and reliability. If you want to ship things users interact with, not spend most of your time tuning models in isolation. With a backend background, you are already most of the way there. The rest is learnable.</p><p>If you are unsure, look at your past work. The signal is usually there. Did you enjoy isolating variables and converging on precise explanations over time, or did you enjoy assembling systems, shipping them, and seeing them handle real usage? Both are valid. They just lead to different careers.</p><p>A useful way to frame the difference is this. ML engineering pushes the frontier forward. AI engineering makes that frontier usable.</p><p>Both matter. What does not work is trying to do both at once and ending up shallow in each. The engineers who will be in the strongest position going into 2026 are not the ones who tried to learn everything. They are the ones who chose a direction and went deep.</p>]]></content:encoded></item></channel></rss>