<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Andrea’s Substack]]></title><description><![CDATA[My personal Substack]]></description><link>https://a4al6a.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png</url><title>Andrea’s Substack</title><link>https://a4al6a.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 17:13:07 GMT</lastBuildDate><atom:link href="/__u/a4al6a.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Andrea Laforgia]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[a4al6a@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[a4al6a@substack.com]]></itunes:email><itunes:name><![CDATA[Andrea Laforgia]]></itunes:name></itunes:owner><itunes:author><![CDATA[Andrea Laforgia]]></itunes:author><googleplay:owner><![CDATA[a4al6a@substack.com]]></googleplay:owner><googleplay:email><![CDATA[a4al6a@substack.com]]></googleplay:email><googleplay:author><![CDATA[Andrea Laforgia]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA["What is so bad with Waterfall?"]]></title><description><![CDATA[A Brief History of a Methodology That Never Existed]]></description><link>https://a4al6a.substack.com/p/what-is-so-bad-with-waterfall</link><guid isPermaLink="false">https://a4al6a.substack.com/p/what-is-so-bad-with-waterfall</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Fri, 07 Aug 2026 17:51:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!MDHd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This is what Lars wrote on LinkedIn:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!MDHd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!MDHd!, /__u/a4al6a.substack.com/w_424, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp 424w, /__u/substackcdn.com/image/fetch/$s_!MDHd!, /__u/a4al6a.substack.com/w_848, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp 848w, /__u/substackcdn.com/image/fetch/$s_!MDHd!, /__u/a4al6a.substack.com/w_1272, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!MDHd!, /__u/a4al6a.substack.com/w_1456, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!MDHd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp" width="1082" height="1124" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1124,&quot;width&quot;:1082,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:150160,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://a4al6a.substack.com/i/210254742?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!MDHd!, /__u/a4al6a.substack.com/w_424, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp 424w, /__u/substackcdn.com/image/fetch/$s_!MDHd!, /__u/a4al6a.substack.com/w_848, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp 848w, /__u/substackcdn.com/image/fetch/$s_!MDHd!, /__u/a4al6a.substack.com/w_1272, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp 1272w, /__u/substackcdn.com/image/fetch/$s_!MDHd!, /__u/a4al6a.substack.com/w_1456, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c825885-259d-43d4-8230-457b5bbd348c_1082x1124.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I&#8217;ll tell you, Lars, what&#8217;s wrong:<br><br><strong>&#8220;Waterfall was coined in the 70&#8217;s as a type of project structure for software development.&#8221;</strong></p><p>Wrong twice over.</p><p>Nobody coined it as a methodology. There is no founding paper, no author, no manifesto, no standards body that sat down and said &#8220;let us call this Waterfall.&#8221; The word was applied retrospectively, by other people, mostly as criticism. Its first known appearance in print is Bell and Thayer&#8217;s 1976 ICSE paper <em>Software Requirements: Are They Really A Problem?</em>, and they were describing someone else&#8217;s diagram.</p><p>That someone else was Winston Royce, 1970, <em>Managing the Development of Large Software Systems</em>. He never used the word. More importantly, he drew the sequential model in Figure 2 and then wrote, in the very next breath, that the implementation as described &#8220;is risky and invites failure.&#8221; The remaining two thirds of the paper is five recommended fixes, one of which is literally &#8220;do it twice&#8221;, a prototype iteration, and another of which is &#8220;involve the customer.&#8221; Royce is cited as the father of waterfall by people who read one diagram and stopped.</p><p>And the sequential phase model predates Royce by fourteen years anyway. Herbert Benington described it in 1956, presenting how SAGE was built. In his 1983 retrospective he added the detail everyone omits: they built a prototype first.</p><p>Waterfall as an actual mandated practice comes from procurement, not engineering. DOD-STD-2167 in 1985 effectively baked it into US defence contracting. MIL-STD-498 unbaked it in 1994 by explicitly permitting evolutionary and incremental lifecycles. So the one institution that genuinely did codify waterfall spent nine years doing it and then stopped.</p><p><strong>&#8220;Non-software engineering fields use traditional sequential frameworks that mirror the waterfall model.&#8221;</strong></p><p>This is the load-bearing claim and it is not true.</p><p>Naval architects have used the <em>design spiral</em> since at least Evans in 1959. It is called a spiral because you go round it repeatedly. Aerospace conceptual design runs iterative sizing loops. Toyota&#8217;s product development uses set-based concurrent engineering, deliberately carrying multiple design alternatives forward rather than committing early, documented by Ward, Liker and Sobek in <em>The Second Toyota Paradox</em>. Construction has the Last Planner System and Integrated Project Delivery, and lean construction has been an organised movement since 1997. SpaceX&#8217;s entire hardware method is build, fly, explode, learn, repeat.</p><p>Systems engineering does not use waterfall either. It uses the V-model, which is not a sequence but a decomposition paired with verification and validation looping back to each corresponding level. ISO/IEC/IEEE 15288 and the INCOSE handbook both explicitly enumerate incremental and evolutionary lifecycle models as legitimate.</p><p>What is true is that construction <em>execution</em> is sequential, because you cannot roof a building before you wall it. That is physics, not project management philosophy. The design phase preceding it is iterative as hell.</p><p><strong>&#8220;Only software development engineering frown upon it.&#8221;</strong></p><p>See above. Lean construction, lean product development, set-based design and the Toyota Production System are all pushbacks on sequential single-pass design, and none of them came out of software.</p><p><strong>The unexamined question: why do other disciplines front-load design?</strong></p><p>Because their cost of change curve is brutal. Moving a load-bearing wall after the concrete cures costs a fortune. Recalling a shipped airframe costs more. When change is expensive, you buy insurance by thinking harder up front, and that is rational.</p><p>Software&#8217;s marginal cost of change is close to zero, and modern engineering practice, continuous delivery, automated testing, trunk based development, is a deliberate campaign to keep it there. Copying the process of a discipline whose economics are the opposite of yours is cargo culting. The post treats the sequence as the principle. The sequence is a consequence.</p><p><strong>&#8220;Fixed price, fixed date. I asked an agile team when they&#8217;d be finished and they said we don&#8217;t know.&#8221;</strong></p><p>They were a bad team. You have described one anecdote and generalised it to a discipline.</p><p>Empirical forecasting is not exotic. Throughput, cycle time distributions and Monte Carlo simulation over historical data will give you a date with a confidence interval, which is more than a Gantt chart has ever given anyone. Vacanti&#8217;s <em>Actionable Agile Metrics for Predictability</em> is the standard reference. A plan produces a confident date. Empirical forecasting produces an accurate one. These are different things and the industry has spent fifty years confusing them.</p><p>Also, fixed price and fixed date do not make an estimate correct. They transfer risk to the supplier, who prices the risk in and then manages scope quietly. Everyone flexes scope. Agile flexes it visibly and early. Waterfall flexes it in month eleven of a twelve month contract.</p><p>And fixed price agile contracts exist and have for two decades: the Norwegian PS2000 model, Sutherland&#8217;s &#8220;money for nothing, change for free&#8221;, capped time and materials with a fixed scope reserve.</p><p><strong>&#8220;Traditional engineering is better at hitting dates.&#8221;</strong></p><p>Flyvbjerg&#8217;s dataset of over 16,000 projects says otherwise. Roughly 47.9% come in on budget. About 8.5% hit budget and schedule. Around 0.5% hit budget, schedule and benefits. The exemplars of the sequential framework include the F-35, the 787, Berlin Brandenburg Airport, Crossrail and HS2. If you want to hold up traditional engineering as the predictability benchmark, you should look at the data first.</p><p>Note also what Flyvbjerg actually recommends: modularity, repeatable units, and slow deliberate planning that is itself iterative, the Pixar model of planning by rewriting. That is not waterfall.</p><p><strong>&#8220;Agile is a software development framework, it was intended for software development.&#8221;</strong></p><p>Historically backwards. Scrum takes its name and its central idea from Takeuchi and Nonaka&#8217;s 1986 HBR article <em>The New New Product Development Game</em>, a study of overlapping-phase product development at Fuji-Xerox, Canon, Honda, NEC and Epson. Copiers, cameras, cars and printers. Scrum was imported <em>into</em> software from manufacturing, not exported out of it.</p><p>Kanban is Toyota. Lean is Toyota. PDCA is Shewhart and Deming, from statistical process control in manufacturing. Software borrowed nearly all of it. Arguing that agile cannot leave software is arguing that it cannot go home.</p><p>This is also just the genetic fallacy. Where an idea originated tells you nothing about where it applies.</p><p><strong>&#8220;Worldwide router upgrade, package upgrade, cloud migration. NO.&#8221;</strong></p><p>Here Lars is half right, for entirely the wrong reason.</p><p>The variable that matters is not software versus not-software. It is how much uncertainty the work contains. Snowden&#8217;s Cynefin framework is the useful lens: work that is <em>clear</em> or <em>complicated</em> has knowable answers and expert practice, so plan it and execute the plan. Work that is <em>complex</em> has answers that only emerge through probing, so run experiments and adapt.</p><p>A worldwide router upgrade is complicated. Plan it. Fine. Nobody serious is proposing a daily standup to discuss whether the router should exist.</p><p>But cloud migrations fail precisely because people classify them as complicated when they are complex. The undocumented dependency, the licence that does not port, the service nobody knew was still in the critical path. Which is why competent migrations run in waves, with a pilot, learning from each wave. That is incremental delivery with feedback. You can refuse to call it agile if the word upsets you, but that is what it is.</p><p>And note the third category you have skipped: your example of &#8220;upgrading a software package to a new version&#8221; being unsuitable for agile is odd, given that shipping small, frequent, reversible changes is exactly what DORA&#8217;s decade of research associates with <em>both</em> higher throughput and higher stability. Big-bang version upgrades are the failure mode, not the safe option.</p><p><strong>Where this actually leaves us</strong></p><p>The post is arguing about two labels when the real variable is the cost of being wrong divided by the speed at which you find out. Front-load your thinking when discovering you were wrong is expensive and slow. Shorten your feedback loops when it is cheap and fast. Most real work sits between the two, and the sensible answer is a plan you fully intend to revise, which is neither of the two tribes the post is asking you to pick.</p><p><strong>Waterfall is not a framework anyone designed. It is a shape people noticed in a diagram drawn by a man who was warning them about it.</strong></p>]]></content:encoded></item><item><title><![CDATA[The Relay Method]]></title><description><![CDATA[How a chain of specialized AI agents &#8212; each allowed to talk only to its neighbours &#8212; turns a vague human wish into working, audited software]]></description><link>https://a4al6a.substack.com/p/the-relay-method</link><guid isPermaLink="false">https://a4al6a.substack.com/p/the-relay-method</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Mon, 22 Jun 2026 12:07:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Summary</h2><p>Most &#8220;AI coding&#8221; setups are a single model doing everything: it talks to you, designs the solution, writes the code, and tells you it&#8217;s done. That works until it doesn&#8217;t &#8212; and when it doesn&#8217;t, you can&#8217;t see <em>where</em> the reasoning went wrong, because it all happened inside one opaque conversation.</p><p><strong>The Relay Method</strong> takes the opposite stance. It splits the work across a small chain of specialized agents, each with one job and one rule: <strong>you may only talk to the agent on your left and the agent on your right.</strong> A human problem flows down the chain, getting sharper at every step; working software and status flow back up. Every message between agents is a file on disk, appended to a single audit log you can replay line by line.</p><p>The result is a system where each handoff is a <em>translation</em> between levels of abstraction, where implementation detail physically cannot leak up to the human and the human&#8217;s framing cannot leak down into the code, and where &#8212; months later &#8212; you can reconstruct exactly how &#8220;I want a 3D Tetris game&#8221; became a specific commit.</p><p>This article describes the method, each persona in the chain, how they communicate, and why the constraints are the point.</p><div><hr></div><h2>The core idea: a bucket brigade for abstraction</h2><p>Picture five stations in a line:</p><pre><code><code>Owner (human)  &#8644;  Interpreter  &#8644;  Analyst  &#8644;  Examiner  &#8644;  Builder
&#9492;&#9472;&#9472; live chat &#9472;&#9472;&#9496;  &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472; relay messages &#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
</code></code></pre><p>Each arrow is an <strong>edge</strong>, and every edge is a translation to a different level of abstraction. Going down the chain, the same intent is progressively sharpened:</p><ul><li><p>The <strong>Owner</strong> has a problem, in plain human words.</p></li><li><p>The <strong>Interpreter</strong> restates it as a <em>need to implement</em> &#8212; still in the Owner&#8217;s language, never in solution terms.</p></li><li><p>The <strong>Analyst</strong> reframes that need as an <em>observable behaviour</em> &#8212; what must be true, with no hint of <em>how</em>.</p></li><li><p>The <strong>Examiner</strong> decomposes the behaviour into <em>expectations</em> &#8212; precise, checkable statements.</p></li><li><p>The <strong>Builder</strong> writes code and produces <em>evidence</em> that each expectation holds.</p></li></ul><p>The neighbours-only rule is not bureaucracy; it is the mechanism. Because the Builder can only speak to the Examiner, it can never leak file names and data structures up to the human. Because the Interpreter can only speak to the Analyst (downward) and the Owner (upward), it can never smuggle business framing into the code. Each role is forced to speak the vocabulary of its own edge &#8212; and that is exactly what keeps the abstraction levels clean.</p><p>There&#8217;s a sixth participant who stands <em>outside</em> the line &#8212; the <strong>Sentinel</strong> &#8212; whose job is to make sure everyone honours their edge. More on them later.</p><div><hr></div><h2>The personas</h2><h3>1. The Owner &#8212; the human with a problem</h3><p>The Owner is you. You don&#8217;t write code, you don&#8217;t design the architecture, and you don&#8217;t manage the agents. You <strong>state a problem</strong>, answer clarifying questions, approve a plan, and then gate each increment: continue, stop, or change course.</p><p>This is the quiet revolution of the method: your role shifts from <em>author</em> to <em>editor</em>. You are no longer producing the solution; you are evaluating proposals and steering. You speak to exactly one agent &#8212; the Interpreter &#8212; and only ever in the language of needs and outcomes.</p><h3>2. The Interpreter &#8212; the Owner&#8217;s voice inside the machine</h3><p>The Interpreter is the bilingual diplomat. Upward, it speaks human: it asks you clarifying questions <em>before</em> planning (so ambiguity is resolved rather than assumed), proposes a <strong>roadmap</strong> of potentially-shippable iterations, presents finished <strong>increments</strong>, and asks whether to keep going. Downward, it hands the Analyst one <strong>behaviour to implement</strong> at a time &#8212; still expressed as a need (&#8221;the player should see the piece fall and lock at the bottom&#8221;), never as a design.</p><p>Crucially, the Interpreter is the only edge that is a <em>live conversation</em>: you type in its window and it talks back. Everything below it happens through files, asynchronously, while you watch.</p><h3>3. The Analyst &#8212; from need to observable behaviour</h3><p>The Analyst is the translator that strips solutions out of needs. It receives a behaviour-to-implement and reframes it as a pure <strong>behaviour</strong>: an actor, an observable outcome, and explicit boundaries &#8212; with every trace of <em>how</em> removed.</p><p>&#8220;Build a falling piece&#8221; becomes: <em>&#8220;Without any user interaction, the piece descends by exactly one cell at a steady, observable cadence; it stays aligned to the grid at each step; this covers descent through empty space only &#8212; what happens at the floor is a separate behaviour.&#8221;</em></p><p>Notice what&#8217;s absent: no mention of timers, frames, data structures, or languages. The Analyst answers <strong>what must be observably true</strong>, and hands that to the Examiner. When status comes back up, the Analyst translates it the other way &#8212; reporting to the Interpreter only <em>which problem was solved</em>, never expectations or tests.</p><h3>4. The Examiner &#8212; expectations, not assertions</h3><p>The Examiner is where the method gets its rigor, and it runs on <strong>Expectation-Driven Development</strong> (EDD) &#8212; a practice for human&#8211;AI collaboration in which correctness is established by stating expectations in plain language and then demanding <em>evidence</em> that they hold, rather than by anyone declaring &#8220;done.&#8221; (See the original write-up: <em><a href="/__u/a4al6a.substack.com/p/expectation-driven-development-a">Expectation-Driven Development</a></em>.)</p><p>The Examiner takes a behaviour and decomposes it into a <strong>set of expectations</strong> &#8212; <code>E1..En</code> &#8212; each a precise, checkable statement of what must be true, plus an end-to-end <em>integration</em> expectation that ties them together. For the falling piece: <em>&#8220;E1: the piece&#8217;s vertical position decreases by exactly one cell per tick. E2: between ticks it is stationary. E3: it never leaves the grid&#8230;&#8221;</em> Still no &#8220;how&#8221; &#8212; expectations describe outcomes, not mechanisms.</p><p>It sends these to the Builder and waits for <strong>evidence</strong>. Then it judges. In EDD terms, the Examiner insists on <em>executed</em> evidence &#8212; real runs, real output, real measured values &#8212; over <em>generative</em> evidence, where the AI merely narrates what it believes would happen. If an expectation is unmet or the evidence is unconvincing, the Examiner returns a <strong>verdict</strong> saying what still fails, and the Builder iterates. This <code>expectation &#8594; evidence &#8594; verdict</code> loop repeats until every expectation is satisfied.</p><p>The striking consequence, straight from EDD: <strong>the code may carry no unit tests at all.</strong> The expectations and their evidence &#8212; recorded permanently in the audit log &#8212; <em>are</em> the proof of correctness. The Examiner is the test oracle, and the ledger is the test report.</p><h3>5. The Builder &#8212; the only one who touches code</h3><p>The Builder is the implementer, and the only agent that writes and runs real code. It receives expectations from the Examiner and produces <strong>evidence</strong> that they now hold: it builds, it runs the program, it captures screenshots and measured outputs, and it reports back <strong>which expectations are fulfilled</strong> &#8212; and <em>only</em> that. It is forbidden from leaking implementation detail upward; &#8220;I used a hash map keyed by cell coordinates&#8221; is a contract violation. The correct report is &#8220;E2 now holds; here is the run that shows it.&#8221;</p><p>The Builder talks to no one but the Examiner. It doesn&#8217;t know who the Owner is or what the roadmap looks like. It knows expectations, and it knows how to produce evidence. That narrowness is its strength.</p><h3>6. The Sentinel &#8212; the auditor outside the chain</h3><p>The five-station line is a closed pipeline, but pipelines drift. Over a long run, the Builder starts slipping code internals into its evidence; the Interpreter starts smuggling implementation hints into a behaviour. Who watches the contracts?</p><p>The <strong>Sentinel</strong> does. It stands outside the chain &#8212; not a link in it &#8212; and is the one party allowed to read the <em>entire</em> conversation. It periodically audits every message against its edge&#8217;s contract (does this builder&#8594;examiner message leak implementation detail? does this analyst&#8594;examiner message prescribe a solution?), writes findings to an audit log, and &#8212; by design &#8212; may message any agent directly with an <code>advisory</code>, a <code>warning</code>, or a <code>directive</code> to pull it back on-contract. In practice it earns its keep: it will catch the Builder naming functions in its evidence and the Interpreter drifting into solutioning &#8212; corrections a single mega-prompt would never surface.</p><p>The Sentinel never blocks or rewrites messages. It observes, reports, and nudges. It is the method&#8217;s conscience.</p><div><hr></div><h2>How they communicate: files you can replay</h2><p>There is no server and no message broker. Every message is a small JSON file written to a mailbox directory, and simultaneously appended as one line to a single, gap-free, append-only <strong>ledger</strong>. That ledger is the system of record. You can replay the entire conversation message by message, filter by agent, and trace any decision&#8217;s full lineage &#8212; from the Owner&#8217;s first sentence to a specific commit &#8212; because every message carries the id of the message it replies to.</p><p>The rules live in one place: a topology file that fixes <strong>which edges exist</strong> and <strong>which message types are legal on each edge</strong>. A wrong neighbour or a wrong message type is rejected before anything is written. This is &#8220;who-talks-to-whom&#8221; enforced as data, not as etiquette.</p><p>The message vocabulary mirrors the abstraction ladder:</p><ul><li><p><strong>Owner &#8596; Interpreter:</strong> <code>problem</code>, <code>clarification</code>, <code>roadmap</code>, <code>roadmap-verdict</code>, <code>increment</code>, <code>continue-query</code>, <code>feedback</code>, <code>result</code>, <code>question</code>.</p></li><li><p><strong>Interpreter &#8594; Analyst:</strong> <code>behaviour-to-implement</code>.</p></li><li><p><strong>Analyst &#8594; Examiner:</strong> <code>behaviour</code>.</p></li><li><p><strong>Examiner &#8594; Builder:</strong> <code>expectation</code>, <code>verdict</code>.</p></li><li><p><strong>Builder &#8594; Examiner:</strong> <code>evidence</code>.</p></li><li><p><strong>Examiner &#8594; Analyst &#8594; Interpreter:</strong> <code>behaviour-status</code> (climbing back up, re-translated at each step).</p></li><li><p><strong>Sentinel &#8594; any agent:</strong> <code>advisory</code>, <code>warning</code>, <code>directive</code>.</p></li></ul><h3>The rhythm of one behaviour</h3><pre><code><code>Interpreter --behaviour-to-implement--&gt; Analyst
Analyst     --behaviour--------------&gt;   Examiner
Examiner    --expectation------------&gt;   Builder    (E1..En + an integration expectation)
Builder     --evidence---------------&gt;   Examiner
Examiner    --verdict (if unmet)-----&gt;   Builder    &#10226; loop until every expectation holds
Examiner    --behaviour-status-------&gt;   Analyst
Analyst     --behaviour-status-------&gt;   Interpreter
Interpreter --increment + continue?--&gt;   Owner
</code></code></pre><p>Each agent runs in its own session and stays idle &#8212; costing nothing &#8212; until a tiny dispatcher notices a message waiting in its inbox and wakes it. Work flows by itself; the human only re-enters at the gates.</p><div><hr></div><h2>Why the constraints are the point</h2><p>It&#8217;s tempting to see the neighbours-only rule and the per-edge vocabularies as friction. They are the opposite &#8212; they are what make the system legible:</p><ul><li><p><strong>Clean abstraction boundaries by construction.</strong> Detail cannot leak up; framing cannot leak down. The structure of the conversation enforces separation of concerns better than any code review.</p></li><li><p><strong>A complete, tamper-evident audit trail.</strong> The ledger is append-only with a gap-free sequence, so a missing or out-of-order message is itself a signal. Commit it alongside the code and you have a permanent record of <em>how</em> the software was produced, not just <em>what</em> was produced.</p></li><li><p><strong>Resilience.</strong> Because all state is inspectable files, a stuck chain is just a message sitting in an inbox. A crash or restart loses only the running processes; relaunch the agents and they pick up exactly where they left off.</p></li><li><p><strong>Specialization over a bigger brain.</strong> The win isn&#8217;t a smarter agent &#8212; it&#8217;s a <em>team</em> of narrow agents communicating under strict, auditable contracts.</p></li></ul><p>We put it to the test by building <strong>ThreeDeeBlocks</strong>, a 3D Tetris clone compiled to WebAssembly, end to end. The swarm worked through dozens of behaviours largely autonomously, leaving a complete trail of every expectation, every piece of executed evidence, and every course-correction &#8212; including the Sentinel flagging contract drift in real time, and the Interpreter surfacing an honest engineering substitution (a different toolchain) to the Owner for sign-off rather than burying it.</p><div><hr></div><h2>Conclusions</h2><p>The Relay Method is a bet that the path to trustworthy AI-built software is not a single, ever-more-capable agent, but a <strong>structured conversation</strong> between modest ones &#8212; constrained so tightly that the structure itself does much of the quality work.</p><p>Three ideas do the heavy lifting:</p><ol><li><p><strong>A linear chain of single-purpose agents</strong>, each speaking only to its neighbours, so every handoff is a clean translation across abstraction levels.</p></li><li><p><strong>Expectation-Driven Development at the core</strong>, so correctness is proven by plain-language expectations and executed evidence rather than asserted &#8212; sometimes with no unit tests in the code at all.</p></li><li><p><strong>A file-based, append-only audit trail plus an external Sentinel</strong>, so the whole thing is inspectable, replayable, and self-policing.</p></li></ol><p>The human comes out changed too: not an author hunched over a keyboard, but an editor who states problems, reviews increments, and decides. If you&#8217;ve ever wished you could <em>see</em> how an AI arrived at a piece of code &#8212; and trust that it actually does what it claims &#8212; the Relay Method is one concrete way to get there.</p><div><hr></div><h2>References &amp; further reading</h2><ul><li><p><strong>Expectation-Driven Development</strong> &#8212; the practice the Examiner is built on (expectations, executed vs. generative evidence, stabilizing into tests): <a href="/__u/a4al6a.substack.com/p/expectation-driven-development-a">https://a4al6a.substack.com/p/expectation-driven-development-a</a></p></li><li><p><strong>Conway&#8217;s Law</strong> (Melvin Conway, 1968) &#8212; that systems mirror the communication structure of the organizations that build them; the Relay Method designs that communication structure deliberately.</p></li><li><p><strong>Behaviour-Driven Development</strong> (Dan North) &#8212; a precursor in spirit: describing behaviour in structured natural language. EDD relaxes its framework constraints for human&#8211;AI collaboration.</p></li><li><p><strong>The test oracle</strong> &#8212; the classic testing notion of an authority that decides whether an output is correct; in this method, the Examiner plays that role and the ledger records its verdicts.</p></li><li><p><strong>Continuous Delivery / Modern Software Engineering</strong> (Dave Farley) &#8212; on working in small, verifiable increments behind explicit gates, echoed here in the roadmap-and-increment rhythm.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Four Eyes Without Pull Requests]]></title><description><![CDATA[How to design a review control that survives an audit and produces real safety, not just artefacts]]></description><link>https://a4al6a.substack.com/p/four-eyes-without-pull-requests</link><guid isPermaLink="false">https://a4al6a.substack.com/p/four-eyes-without-pull-requests</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Mon, 04 May 2026 14:29:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>Abstract</strong></h2><p>A reader of this newsletter, who runs a team that practises trunk-based development with TDD and pair programming, recently received feedback from an ISO auditor. Nothing in the team&#8217;s system of record explicitly recorded that a second pair of eyes had verified each change, nor whose eyes those were. He saw two options: log the second pair of eyes in commit messages, or fall back to short-lived pull requests with a single approval.</p><p>Both framings miss the point. The auditor&#8217;s complaint is correct on a narrow procedural matter, and wrong on the underlying mechanism. The answer is concrete, and made of five pieces: a written policy, structured commit trailers, automated enforcement with a tight exception path, supplementary telemetry, and a sampling audit with a defined protocol. Together they form a control that is testable, deliberate, and honest about its limits.</p><h2><strong>A reader&#8217;s question</strong></h2><p>The reader wrote a careful letter. His team had moved to trunk-based development with pair programming. They were happier and produced better software. Then the auditor noticed that the team&#8217;s working practice was not visible in any system of record. He suggested co-authored commits or a return to pull requests. He wanted advice.</p><p>The answer matters beyond one team. Many engineering leaders who would otherwise drop pull requests hesitate at the audit stage, and many keep pull requests precisely because they assume audit compliance demands them. <strong>That assumption is wrong</strong>, and it is worth dismantling carefully so the alternative becomes something a serious auditor will sign off, not a hopeful sketch.</p><h2><strong>The category error</strong></h2><p>ISO does not require pull requests. Neither does SOC 2. Neither does the great majority of compliance regimes that ordinary software teams find themselves under.</p><p>ISO/IEC 27001:2022 has Annex A controls covering change management and secure development. Their text, in plain language, requires that changes are authorised, that they are tested, that environments are separated, and that the process is documented. SOC 2 has CC8.1, which requires that the organisation manages changes through a process that includes authorisation and testing. The phrase &#8220;pull request&#8221; appears in neither. The shape of the review is left to the team.</p><p>A pull request is one possible implementation of a change-review control. A pairing session followed by a co-authored commit is another. A signed-off patch series sent through email, in the style the Linux kernel has used for two decades, is a third. The standard cares about the control objective. It is the auditor&#8217;s job to assess whether the team&#8217;s chosen mechanism meets that objective, regardless of its shape.</p><h2><strong>What a control actually needs</strong></h2><p>A review control that an auditor will accept has four properties. It is <em>defined</em> in writing, so the auditor knows what to test against. It is <em>implemented</em> in practice, so what the writing says is what actually happens. It produces <em>evidence</em>, so the implementation can be inspected after the fact. And it is <em>enforced</em>, so bypassing it is hard, or at least visible.</p><p>Mature controls have all four. Weak controls have three of the four, and it is usually the implementation that is missing. This is the awkward fact at the heart of the typical pull request workflow.</p><h2><strong>The implementation gap in pull request review</strong></h2><p>A pull request creates an artefact. It does not, on its own, create a review. The artefact and the review are independent variables. A diff is opened. Someone clicks approve. The change merges. Did the approver read the diff? Did they understand it? Did they push back on anything? The artefact does not tell us. The auditor sees a green tick and is satisfied. The control objective, in substance, is unmet.</p><p>I am not arguing that all pull request review is theatre. Some teams review carefully and reject changes that need rework. The argument is narrower. The artefact does not distinguish between deep review and rubber stamping, so the artefact alone is not evidence that review happened. When the artefact is the evidence, the control is weak even where the practice is strong.</p><p>Pair programming is the opposite trade. Two engineers engage with the code in real time, arguing about names, test design, edge cases, and whether the change should exist at all. The review is continuous, and it shapes the code as it is written rather than inspecting it afterwards. The findings in <em>Accelerate</em> and the annual State of DevOps reports indicate that high-performing teams favour smaller batches, more frequent integration, and trunk-based development. The kind of continuous review that pairing makes possible fits the pattern those findings describe better than long-lived branches do.</p><p>The catch is that pairing leaves no native artefact. The substance is present, the artefact is missing. The auditor&#8217;s complaint is correct on this narrow point. The fix is not to throw away the substance. The fix is to add the artefact.</p><h2><strong>The composite that closes the gap</strong></h2><p>Five pieces, working together, close the implementation and evidence gap that pure pairing leaves open. Compared with a typical pull request workflow, the composite is stronger on the question that matters most for safety (whether real review happened) and weaker on the question of independent-action provability that pull requests answer through platform authentication. That trade is worth making for most teams under ISO 27001 and SOC 2, but the article will be honest about both sides.</p><h3><strong>The policy</strong></h3><p>The team&#8217;s change management policy must say, in plain language, that pair programming or mob programming is the team&#8217;s primary control for the four-eyes requirement on code changes, that every change reaching the trunk is produced by at least two engineers working together, and that the second engineer is recorded on the commit. Without this document, an auditor has no benchmark against which to test the team&#8217;s behaviour. With it, every other artefact has meaning. Half a page, signed by the engineering lead and reviewed annually, is sufficient.</p><h3><strong>The commit trailers</strong></h3><p>Git supports structured trailers natively, and the conventions are not new. The Linux kernel has used <code>Signed-off-by</code>, <code>Reviewed-by</code>, <code>Tested-by</code>, and <code>Acked-by</code> for two decades to attribute who did what to a patch. The same machinery is available to any team. A pairing team adopts <code>Co-authored-by</code> for the active partner and, where appropriate, <code>Reviewed-by</code> for an additional asynchronous reviewer. A typical commit looks like this.</p><pre><code><code>feat(billing): apply VAT for cross-border invoices

Co-authored-by: Jane Doe &lt;jane@example.com&gt;
Reviewed-by: Marco Rossi &lt;marco@example.com&gt; (commit b3f7a2c)
</code></code></pre><p>These trailers are machine-readable. A script can count them, sample them, and verify their format. They are also human-readable: an auditor opening the commit sees, at a glance, who participated.</p><h3><strong>The enforcement</strong></h3><p>A pre-commit hook on the developer&#8217;s machine refuses commits that do not meet the rules. A CI check on the integration branch is the backstop, in case the local hook is bypassed. The rules need to be specific, otherwise enforcement is theatre too.</p><p>A defensible minimum has five rules. Every commit on the integration branch must contain at least one <code>Co-authored-by</code> or <code>Reviewed-by</code> trailer. The trailer&#8217;s email must resolve against the company directory, so anonymous attestation is impossible. The author and the named co-author or reviewer must be different identities. A commit without a co-author must contain an explicit <code>Solo-work:</code> line naming an asynchronous reviewer by directory identity and the commit hash that reviewer inspected. A commit lacking any of the above is rejected by CI, and the rejection is logged.</p><p>The exception path is deliberately narrow. There is no zero-review path. Either the work was paired (<code>Co-authored-by</code>), or it was reviewed asynchronously (<code>Reviewed-by</code>), or it was solo with an explicit asynchronous reviewer attached to a specific commit hash (<code>Solo-work:</code>). All three produce structured evidence; none allows a change to reach the trunk unattested.</p><h3><strong>The telemetry</strong></h3><p>Where pairing tools such as Tuple, Pop, or mob.sh are in use, their session records corroborate the trailers. Calendar invites for recurring mob sessions add a second source. None of this is the primary evidence; it is corroboration, and it makes systematic falsification harder. The composite does not lean on telemetry, because individual sources are easy to dismiss in isolation. The trailers and the audit do the work.</p><h3><strong>The sampling audit</strong></h3><p>This is the linchpin of the composite, and it deserves a defined protocol rather than a vague intent.</p><p>A defensible minimum specification: each quarter, draw a random sample of commits from the integration branch using a seeded selection script. The sample size is the greater of ten percent of commits in the period or twenty commits, capped at fifty. The audit is run by a person who is not the engineering lead of the audited team, such as a peer team lead, a quality engineer, or an internal audit function, in order to avoid self-review. For each sampled commit, both named engineers are interviewed separately, with three standard questions: what was the purpose of this change, what alternatives were considered, and what would you do differently now. The interviewer records answers in a shared log alongside the commit hash and the date.</p><p>Failure criteria are explicit. If two or more sampled engineers cannot recall a change they are recorded on, the result is a process finding that triggers a documented remediation: a refresh of the policy, a refresh of the team&#8217;s pairing protocol, or, in serious cases, a return to short-lived pull requests until the team rebuilds the practice. The criterion does not catch deliberate fraud reliably; it does catch process drift, which is the failure mode most relevant under ISO 27001 and SOC 2.</p><p>This is the step that converts a culture of trust into a control an external auditor can test, because it gives the auditor a mature internal control to inspect rather than a claim to take on faith. Auditors recognise this structure. It is the same shape they apply to financial controls.</p><h2><strong>Three honest objections</strong></h2><p>A serious reader will already have noticed three problems. Let me address them rather than pretend they do not exist.</p><h3><strong>Self-attestation</strong></h3><p>A <code>Co-authored-by</code> trailer is, in the end, a string the committer types. A pull request approval is stronger on this narrow point: it is tied to an authenticated user action with its own audit trail. The composite does not pretend to match that property.</p><p>What the composite does instead is shift the weight of the control from a <em>preventative</em> mechanism (a click that allegedly happens before merge) to a <em>detective</em> mechanism (a sampling audit that verifies, after the fact, that review happened in substance). This is a recognised pattern in audit design: detective controls are weaker on the moment of action and stronger on the cumulative behaviour of the system. Two engineers conspiring to falsify a trailer must also conspire to give consistent, detailed answers to separate interviews about a commit they did not work on, and they must do so every time a sampling audit lands on a falsified commit. Casual misuse becomes uneconomic; deliberate fraud becomes detectable through pattern analysis.</p><p>The composite does not eliminate the self-attestation problem. It raises the cost of circumvention to the point where the dominant residual risk is process drift, not collusion. For ISO 27001 and SOC 2, that is sufficient. For regimes that demand non-repudiation as a primary property of the control, it is not, and the regulatory carveouts below apply.</p><h3><strong>Distributed teams</strong></h3><p>Pairing across timezones is genuinely hard. A team that cannot pair synchronously cannot rely on pairing as its primary control without modification. The honest answer is a hybrid: synchronous pairing within a timezone, asynchronous review with <code>Reviewed-by</code> trailers across timezones, and a policy that distinguishes the two cases. The asynchronous case shares structure with a pull request workflow, but the review is recorded in the commit history rather than in a separate workflow, the integration is not delayed by it, and the same enforcement and sampling apply. The article does not claim distributed teams can drop the second pair of eyes. It claims they can record it without inheriting the bottleneck of long-lived branches.</p><h3><strong>Solo work</strong></h3><p>Some changes are genuinely solo: a quick fix, a documentation typo, a configuration tweak that does not warrant a pairing session. The exception path in the hook handles this, but only with a name and a commit hash attached, and only with the same sampling treatment as paired commits. The team&#8217;s policy should set a threshold above which pairing is required regardless. This is no different from the way mature pull request workflows distinguish trivial changes from substantive ones; the difference is that the distinction is recorded in the commit, not in the branch lifecycle.</p><h2><strong>The regulatory carveouts</strong></h2><p>There are regimes where the literal separation of author and approver is required, and where genuine pairing cannot satisfy the requirement on its own. SOX controls over financially material systems, FDA rules for software in medical devices, parts of the GxP family in pharmaceutical manufacturing, and certain defence-sector standards come to mind. In those contexts, the team can still pair as its way of working, but a separate, documented approval step by a person who was not part of the pairing is also required. The composite described above is not a universal substitute. It is a substitute for the great majority of teams under ISO 27001 and SOC 2, which is the relevant audience here. A team operating under one of the stricter regimes should read the actual control text and design accordingly, and should be wary of any blog post, including this one, that promises a single mechanism for all cases.</p><h2><strong>Talking to the auditor</strong></h2><p>The conversation with the auditor is a conversation about controls, not artefacts. Open by asking which control objective the auditor is testing, in their own words. Almost always, the answer reduces to &#8220;evidence that a person other than the author has reviewed the change before it reached production&#8221;. Once that objective is on the table, the team presents the policy, the trailers, the enforcement rules, the telemetry, and the sampling audit log as the answer. Bring the policy. Bring an example of a sampled commit with the corresponding interview record. Treat the auditor as a serious reader who needs to be shown a serious control.</p><p>Most auditors respond well to a team that has thought clearly about what the standard is asking for and has produced a deliberate answer. Some do not. If the auditor insists on pull requests as a fixed shape, regardless of the control objective they implement, the team has three options. First, escalate within the audit firm: some firms have specialists who understand modern engineering practice, and some auditors can be replaced. Second, accept the audit firm&#8217;s preference and run a pull request workflow alongside the team&#8217;s real practice, knowing that the pull request is a record-keeping ritual rather than the actual review. Third, move the certification to a different firm.</p><p>None of these is comfortable, and the article will not pretend they are. Escalation is not always available, especially for smaller customers of larger audit firms; switching firms is disruptive and rarely fast; running a parallel ritual erodes the cultural commitment to pairing over time. The honest position is that a team should attempt the first option first, prepare for the second as a fallback during the certification cycle in question, and consider the third only if the audit relationship cannot be repaired.</p><h2><strong>The choice</strong></h2><p>The reader&#8217;s question looked technical. It is, in the end, a choice between two kinds of compliance.</p><p>One kind is built from the appearance of review. It is cheap, recognisable to a tired auditor, and satisfies the letter of the standard at the cost of producing very little of the substance the standard intends. The other kind is built from real review, made auditable. It costs discipline, a small amount of tooling, and a defined sampling protocol. It produces both real safety and a paper trail, and it makes the team better at building software along the way.</p><p>I know which one I think is worth buying. So does the reader, which is why he wrote in the first place. The good news is that he does not have to abandon what is working in order to satisfy an auditor. He has to write the policy, install the hooks, define the enforcement rules, set up the sampling protocol, run the audits, and have the conversation with care. The pairing stays where it belongs, in front of the code, with two pairs of eyes that actually see it.</p><h2><strong>References</strong></h2><ol><li><p>A. Laforgia, <em>Stop Using Pull Requests</em>, <a href="/__u/a4al6a.substack.com/p/stop-using-pull-requests">https://a4al6a.substack.com/p/stop-using-pull-requests</a></p></li><li><p>A. Laforgia, <em>The Illusion of Compliance</em>, <a href="/__u/a4al6a.substack.com/p/the-illusion-of-compliance">https://a4al6a.substack.com/p/the-illusion-of-compliance</a></p></li><li><p>ISO/IEC 27001:2022, <em>Information security, cybersecurity and privacy protection. Information security management systems. Requirements</em>, Annex A controls on change management and secure development.</p></li><li><p>AICPA, <em>Trust Services Criteria for Security, Availability, Processing Integrity, Confidentiality, and Privacy</em>, criterion CC8.1 on change management.</p></li><li><p>Linux kernel project, <em>Submitting patches: the essential guide to getting your code into the kernel</em>, <code>Documentation/process/submitting-patches.rst</code>, on the use of <code>Signed-off-by</code>, <code>Reviewed-by</code>, <code>Tested-by</code>, and <code>Acked-by</code> trailers.</p></li><li><p>Git documentation on commit trailers and the <code>Co-authored-by</code> convention as supported by GitHub and GitLab.</p></li><li><p>P. Hammant and contributors, <em>Trunk Based Development</em>, https://trunkbaseddevelopment.com</p></li><li><p>N. Forsgren, J. Humble, G. Kim, <em>Accelerate: The Science of Lean Software and DevOps</em>, IT Revolution Press, 2018.</p></li><li><p>K. Beck, <em>Test-Driven Development: By Example</em>, Addison-Wesley, 2002.</p></li><li><p>M. Feathers, <em>Working Effectively with Legacy Code</em>, Prentice Hall, 2004.</p></li></ol>]]></content:encoded></item><item><title><![CDATA[No, AI Did Not Kill Agile]]></title><description><![CDATA[Why execution speed doesn't solve a knowledge problem]]></description><link>https://a4al6a.substack.com/p/no-ai-did-not-kill-agile</link><guid isPermaLink="false">https://a4al6a.substack.com/p/no-ai-did-not-kill-agile</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Thu, 09 Apr 2026 09:56:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>TL;DR</strong></h2><p>The &#8220;AI killed Agile&#8221; argument confuses two different problems. There is a speed problem &#8212; how fast can we turn an idea into working software? And there is a knowledge problem &#8212; how do we figure out which idea is the right one? AI largely solves the first. It does not touch the second. Agile was created for the second. The push for &#8220;write a detailed spec, then have AI build it&#8221; is not innovation &#8212; it is waterfall with faster computers. And the person who (accidentally) invented waterfall warned us it would fail.</p><p>The principle that built many Agile practices &#8212; learn what to build by building, observing, and adapting &#8212; is more relevant than ever. When building is nearly free, the only reason not to iterate is if you never understood why iteration mattered in the first place.</p><div><hr></div><h2><strong>1. The Argument (and Why It Feels Right)</strong></h2><p>A new narrative is gaining traction: AI has made Agile obsolete.</p><p>Steve Jones, Executive VP at Capgemini, puts it bluntly: &#8220;Agentic SDLCs are too fast for Agile.&#8221; Jennifer Jones-Mitchell of HumanDrivenAI argues that &#8220;Agile&#8217;s iterative cycles still rely on humans to ideate, test, and refine over weeks or months. Generative AI tools can produce viable prototypes, campaigns, or solutions in minutes&#8221; [6, 15].</p><p>The argument is seductive, and it would be dishonest not to acknowledge why. Anyone who has sat through a performative sprint planning meeting, a daily standup that was just a status report, or a retrospective that changed nothing feels the appeal. If AI can ship in a day what used to take a sprint, why keep the ceremony?</p><p><strong>That frustration is legitimate.</strong> Martin Fowler, one of the Agile Manifesto&#8217;s signatories, coined &#8220;semantic diffusion&#8221; to describe how Agile lost its original meaning. Many organisations practice what he calls &#8220;faux-agile&#8221;: implementing Scrum ceremonies while ignoring the adaptive principles that made those ceremonies useful [4].</p><h2><strong>2. The Confusion: Speed vs. Knowledge</strong></h2><p>The &#8220;AI killed Agile&#8221; argument treats the speed of coding as the constraint Agile was designed to address. It wasn&#8217;t.</p><p>In February 2001, seventeen practitioners met at Snowbird, Utah &#8212; representing Extreme Programming, Scrum, Crystal, Feature-Driven Development &#8212; united against heavyweight, documentation-driven processes that were failing to deliver working software [3]. The result was the Agile Manifesto:</p><ul><li><p><strong>Responding to change</strong> over following a plan</p></li><li><p><strong>Working software</strong> over comprehensive documentation</p></li><li><p><strong>Customer collaboration</strong> over contract negotiation</p></li><li><p><strong>Individuals and interactions</strong> over processes and tools [1]</p></li></ul><p>None of these say &#8220;write code faster.&#8221; They address a different problem: <strong>we do not know what to build until we start building it</strong>. Not because we are bad at planning, but because the information needed to plan correctly does not exist until real users interact with real software.</p><p>This is an epistemological claim, not a methodological preference. The Manifesto&#8217;s Principle 2 states: &#8220;Welcome changing requirements, even late in development.&#8221; Principle 11: &#8220;The best architectures, requirements, and designs <em>emerge</em> from self-organising teams&#8221; [2]. The word <em>emerge</em> is doing heavy lifting. Requirements are not specified and then executed. They are discovered through building, delivering, and observing.</p><p>Research bears this out: 40-60% of software project failures originate in requirements, not slow implementation [17]. Only about 20% of features deliver the positive impact initially intended [14]. The problem was never &#8220;we code too slowly.&#8221; It was always &#8220;<em>we don&#8217;t know which features are the right ones until we put them in front of real users.</em>&#8221;</p><p>AI makes coding faster. It does not make stakeholders know what they want. It does not reduce market uncertainty. It does not reveal whether users will actually use what you built.</p><h2><strong>3. The Inversion: Faster Means More, Not Fewer</strong></h2><p>Here is the part the &#8220;AI kills Agile&#8221; crowd gets exactly backwards.</p><p>If AI makes it cheaper to produce working software, the logical conclusion is not &#8220;we don&#8217;t need iterations.&#8221; It is &#8220;we can afford <em>more</em> iterations.&#8221;</p><p>Eric Ries&#8217;s Lean Startup methodology is built on the Build-Measure-Learn loop: build the smallest thing that tests your assumption, measure what happens when real users interact with it, learn whether to continue or pivot. The antidote to uncertainty is not better planning &#8212; it is &#8220;validated learning: a rigorous method for demonstrating progress when one is embedded in the soil of extreme uncertainty&#8221; [8].</p><p><strong>If AI collapses the &#8220;Build&#8221; phase from weeks to hours, a team can run more Build-Measure-Learn cycles per month. That is not the death of Agile. It is Agile unleashed.</strong></p><p>But here is the nuance: AI makes Build so cheap it rounds to zero. It does not make Measure or Learn cheap. Observing how real users behave, understanding <em>why</em> they behave that way, synthesising that understanding into the next decision &#8212; these remain expensive, slow, and irreducibly human. So the bottleneck shifts. In a world of instant builds, the constraint moves from &#8220;how fast can we code?&#8221; to &#8220;how fast can we learn?&#8221; That does not kill iteration. It recenters it on what always mattered: the learning.</p><p>Kent Beck, co-author of the Agile Manifesto, sees it this way. AI has shifted the cost landscape: &#8220;The whole landscape of what&#8217;s &#8216;cheap&#8217; and what&#8217;s &#8216;expensive&#8217; has all just shifted.&#8221; His recommendation: test approaches previously deemed too expensive. Not &#8220;abandon discipline.&#8221; Not &#8220;skip feedback.&#8221; Experiment more [5].</p><h2><strong>4. &#8220;But I Can Iterate on the Spec&#8221;</strong></h2><p>There is a smarter version of the &#8220;AI kills Agile&#8221; argument that deserves engagement. It goes: &#8220;I write a spec, AI builds a prototype in minutes, I show it to users, I revise the spec, AI rebuilds. The loop is tight and fast. That IS iteration.&#8221;</p><p>This is reasonable, and the workflow it describes can genuinely work &#8212; when the spec is treated as a <strong>disposable hypothesis</strong>. Write a one-page prompt, generate a prototype, show it to users, throw the prompt away, start fresh. That is rapid experimentation with a generative tool. Call it agile if you want. It might well be.</p><p><strong>But that is not what most people mean by &#8220;spec-driven development,&#8221; and it is not what the &#8220;write a solid spec upfront&#8221; advocates are proposing.</strong> Their version is different: invest in a comprehensive, detailed specification &#8212; a document that captures requirements, edge cases, business rules &#8212; then hand it to AI for execution. The spec is the artifact that is reviewed, signed off, refined, and version-controlled. It becomes the primary record of truth.</p><p>And the moment that happens, the spec takes on a gravity of its own. Each revision carries forward assumptions from previous iterations that nobody re-examines. The spec becomes a sedimentary record of decisions, many of which were wrong but are now buried under layers of refinement. You are no longer exploring. You are predicting &#8212; encoding guesses about what users need into a document and then executing those guesses with increasing fidelity.</p><p><strong>That is waterfall.</strong> <strong>It has always been waterfall.</strong> It is what killed the FBI&#8217;s $170 million Virtual Case File &#8212; not slow coding, but a specification that could not keep up with what the organisation was learning about its own needs [9].</p><p>Winston Royce &#8212; the person credited with inventing waterfall &#8212; understood this in 1970. In the same paper, he warned: <em>&#8220;I believe in this concept, but the implementation described above is risky and invites failure.&#8221;</em> He advocated passing through the process &#8220;at least twice&#8221; [16]. The waterfall model as practiced was a <em>misreading</em> of his paper. He presented it as a cautionary example.</p><p>Fifty-five years later, &#8220;write a detailed spec and have AI build it&#8221; recreates the same pattern. The spec is still wrong. It will always be wrong. Because requirements are not merely vague &#8212; they are <em>unknown and emergent</em>. The Manifesto says it plainly: &#8220;Welcome changing requirements, even late in development&#8221; [2]. Not &#8220;write better requirements.&#8221; <em>Welcome</em> the change, because the change is the learning.</p><p>The distinction is precise: <strong>specs are predictive</strong> &#8212; you guess what users need and encode the guess. <strong>Working software in users&#8217; hands is explorative</strong> &#8212; you discover what users actually do. A spec, no matter how rapidly iterated, is a <em>model</em> of what users want. Working software <em>is</em> what users interact with. The gap between model and reality is where projects fail.</p><h2><strong>5. The Real Threat</strong></h2><p>The &#8220;AI killed Agile&#8221; narrative has a political dimension that makes it genuinely dangerous.</p><p>One reason Agile became bureaucratic was that management wanted predictability. Sprints, story points, velocity charts gave managers the illusion of control over an inherently uncertain process. Agile was adopted by organisations that wanted its speed without accepting its core bargain: you cannot predict what you will deliver until you start delivering it.</p><p>AI makes this worse. The pitch is irresistible to a certain kind of stakeholder: &#8220;Just write the spec. The AI will handle it.&#8221; This is not a technology argument. It is a power argument &#8212; the desire to return to a world where management defines requirements, engineering executes them, and the humbling process of learning from users can be skipped.</p><p>Healthcare.gov is the canonical example. Over $2.1 billion spent. Six users on launch day. The team spent months building against frozen requirements that were wrong. AI would have generated the wrong code faster &#8212; and the team still would have discovered the integration failures at the same time, because they deferred testing until the end. The post-mortem recommended &#8220;incremental approaches such as betas, early testing and regular delivery&#8221; [10, 11]. The problem was not development speed. It was the organisational refusal to learn incrementally. <strong>Faster tools do not fix a process designed to avoid learning.</strong></p><p>The real threat is not that engineers will stop iterating. Most good engineers know better. The threat is that AI gives non-technical decision-makers a fresh justification for the approach they always preferred: define everything upfront, hand it off, and skip the uncomfortable discovery that the original vision was wrong.</p><p>Alistair Cockburn, another Manifesto co-author, distilled Agile to four imperatives: <strong>Collaborate, Deliver, Reflect, Improve</strong> [7]. Each one is a check on the natural organisational tendency to plan in isolation, build in darkness, and declare victory. AI does not make any of these obsolete. It makes the temptation to skip them more seductive and the consequences of skipping them more expensive.</p><h2><strong>Conclusions</strong></h2><p>No technology has ever solved the problem of building something nobody wanted. Faster horses, faster compilers, faster AI &#8212; the tool changes, the failure mode does not.</p><p>AI makes building cheaper. That is genuinely transformative. But when building is nearly free, learning what to build becomes the only thing that separates teams that ship products people use from teams that ship products nobody asked for.</p><p>Many Agile practices should evolve (&#8220;<em>We are uncovering better ways of developing software by doing it and helping others do it&#8221;)</em> but the core discipline survives: build something small, put it in front of real people, learn from what happens, adapt. The loop does not get shorter with AI. It gets <em>cheaper</em>. And when the loop is cheaper, you should run it more often, not less.</p><p>AI did not kill Agile. AI killed the <em>excuse</em> for not iterating.</p><div><hr></div><h2><strong>References</strong></h2><p>[1] Beck, K. et al. &#8220;Manifesto for Agile Software Development.&#8221; AgileManifesto.org, 2001. </p><p>https://agilemanifesto.org/</p><p>[2] Beck, K. et al. &#8220;Principles behind the Agile Manifesto.&#8221; AgileManifesto.org, 2001. <a href="https://agilemanifesto.org/principles.html">https://agilemanifesto.org/principles.html</a></p><p>[3] Agile Manifesto Authors. &#8220;History: The Agile Manifesto.&#8221; AgileManifesto.org, 2001. <a href="https://agilemanifesto.org/history.html">https://agilemanifesto.org/history.html</a></p><p>[4] Fowler, M. &#8220;Agile Software Guide.&#8221; MartinFowler.com. <a href="https://martinfowler.com/agile.html">https://martinfowler.com/agile.html</a></p><p>[5] Orosz, G. &#8220;TDD, AI agents and coding with Kent Beck.&#8221; The Pragmatic Engineer. </p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:165580001,&quot;url&quot;:&quot;https://newsletter.pragmaticengineer.com/p/tdd-ai-agents-and-coding-with-kent&quot;,&quot;publication_id&quot;:458709,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;The Pragmatic Engineer&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!6TJt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F5ecbf7ac-260b-423b-8493-26783bf01f06_600x600.png&quot;,&quot;title&quot;:&quot;TDD, AI agents and coding with Kent Beck&quot;,&quot;truncated_body_text&quot;:&quot;Stream the Latest Episode&quot;,&quot;date&quot;:&quot;2025-06-11T16:10:50.764Z&quot;,&quot;like_count&quot;:215,&quot;comment_count&quot;:1,&quot;bylines&quot;:[{&quot;id&quot;:30107029,&quot;name&quot;:&quot;Gergely Orosz&quot;,&quot;handle&quot;:&quot;pragmaticengineer&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58fed27c-f331-4ff3-ba47-135c5a0be0ba_400x400.png&quot;,&quot;bio&quot;:&quot;Big Tech and startups from the inside. Especially relevant for software engineers / AI engineers, useful for anyone working in tech.&quot;,&quot;profile_set_up_at&quot;:&quot;2021-09-06T16:08:47.417Z&quot;,&quot;reader_installed_at&quot;:&quot;2022-03-04T20:04:29.381Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:385140,&quot;user_id&quot;:30107029,&quot;publication_id&quot;:458709,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:458709,&quot;name&quot;:&quot;The Pragmatic Engineer&quot;,&quot;subdomain&quot;:&quot;pragmaticengineer&quot;,&quot;custom_domain&quot;:&quot;newsletter.pragmaticengineer.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Big Tech and startups, from the inside. Highly relevant for software engineers, AI engineers and engineering leaders, useful for those working in tech.&quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/5ecbf7ac-260b-423b-8493-26783bf01f06_600x600.png&quot;,&quot;author_id&quot;:30107029,&quot;primary_user_id&quot;:30107029,&quot;theme_var_background_pop&quot;:&quot;#FF6B00&quot;,&quot;created_at&quot;:&quot;2021-08-25T13:08:12.798Z&quot;,&quot;email_from_name&quot;:&quot;The Pragmatic Engineer&quot;,&quot;copyright&quot;:&quot;Gergely Orosz&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:null,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}}],&quot;twitter_screen_name&quot;:&quot;GergelyOrosz&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:10000,&quot;status&quot;:{&quot;bestsellerTier&quot;:10000,&quot;subscriberTier&quot;:1,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;bestseller&quot;,&quot;tier&quot;:10000},&quot;paidPublicationIds&quot;:[10845,256838,3525780],&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;podcast&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://newsletter.pragmaticengineer.com/p/tdd-ai-agents-and-coding-with-kent?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!6TJt!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F5ecbf7ac-260b-423b-8493-26783bf01f06_600x600.png" loading="lazy"><span class="embedded-post-publication-name">The Pragmatic Engineer</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title-icon"><svg width="19" height="19" viewBox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg">
  <path d="M3 18V12C3 9.61305 3.94821 7.32387 5.63604 5.63604C7.32387 3.94821 9.61305 3 12 3C14.3869 3 16.6761 3.94821 18.364 5.63604C20.0518 7.32387 21 9.61305 21 12V18" stroke-linecap="round" stroke-linejoin="round"></path>
  <path d="M21 19C21 19.5304 20.7893 20.0391 20.4142 20.4142C20.0391 20.7893 19.5304 21 19 21H18C17.4696 21 16.9609 20.7893 16.5858 20.4142C16.2107 20.0391 16 19.5304 16 19V16C16 15.4696 16.2107 14.9609 16.5858 14.5858C16.9609 14.2107 17.4696 14 18 14H21V19ZM3 19C3 19.5304 3.21071 20.0391 3.58579 20.4142C3.96086 20.7893 4.46957 21 5 21H6C6.53043 21 7.03914 20.7893 7.41421 20.4142C7.78929 20.0391 8 19.5304 8 19V16C8 15.4696 7.78929 14.9609 7.41421 14.5858C7.03914 14.2107 6.53043 14 6 14H3V19Z" stroke-linecap="round" stroke-linejoin="round"></path>
</svg></div><div class="embedded-post-title">TDD, AI agents and coding with Kent Beck</div></div><div class="embedded-post-body">Stream the Latest Episode&#8230;</div><div class="embedded-post-cta-wrapper"><div class="embedded-post-cta-icon"><svg width="32" height="32" viewBox="0 0 24 24" xmlns="http://www.w3.org/2000/svg">
  <path classname="inner-triangle" d="M10 8L16 12L10 16V8Z" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"></path>
</svg></div><span class="embedded-post-cta">Listen now</span></div><div class="embedded-post-meta">a year ago &#183; 215 likes &#183; 1 comment &#183; Gergely Orosz</div></a></div><p>[6] InfoQ. &#8220;Does AI Make the Agile Manifesto Obsolete?&#8221; InfoQ, February 2026. <a href="https://www.infoq.com/news/2026/02/ai-agile-manifesto-debate/">https://www.infoq.com/news/2026/02/ai-agile-manifesto-debate/</a></p><p>[7] Cockburn, A. &#8220;Heart of Agile.&#8221; HeartOfAgile.com. </p><p>https://heartofagile.com/</p><p>[8] Ries, E. &#8220;The Lean Startup - Principles.&#8221; TheLeanStartup.com. <a href="https://theleanstartup.com/principles">https://theleanstartup.com/principles</a></p><p>[9] Goldstein, H. &#8220;Who Killed the Virtual Case File?&#8221; IEEE Spectrum, September 2005. <a href="https://spectrum.ieee.org/who-killed-the-virtual-case-file">https://spectrum.ieee.org/who-killed-the-virtual-case-file</a></p><p>[10] CIO. &#8220;6 Software Development Lessons From Healthcare.gov&#8217;s Failed Launch.&#8221; CIO.com, 2013. <a href="https://www.cio.com/article/288541/developer-6-software-development-lessons-from-healthcare-gov-s-failed-launch.html">https://www.cio.com/article/288541/developer-6-software-development-lessons-from-healthcare-gov-s-failed-launch.html</a></p><p>[11] NPR. &#8220;This Slide Shows Why HealthCare.gov Wouldn&#8217;t Work At Launch.&#8221; NPR, November 2013. <a href="https://www.npr.org/sections/alltechconsidered/2013/11/19/246132770/this-slide-shows-why-healthcare-gov-wouldnt-work-at-launch">https://www.npr.org/sections/alltechconsidered/2013/11/19/246132770/this-slide-shows-why-healthcare-gov-wouldnt-work-at-launch</a></p><p>[12] Springer. &#8220;Continuous clarification and emergent requirements flows in open-commercial software ecosystems.&#8221; Requirements Engineering, 2016. <a href="https://link.springer.com/article/10.1007/s00766-016-0259-1">https://link.springer.com/article/10.1007/s00766-016-0259-1</a></p><p>[13] ScienceDirect. &#8220;Tackling Requirements Uncertainty in Software Projects: A Cognitive Approach.&#8221; 2021. <a href="https://www.sciencedirect.com/science/article/pii/S2666307421000218">https://www.sciencedirect.com/science/article/pii/S2666307421000218</a></p><p>[14] DevIQ. &#8220;Big Design Up Front (BDUF): A Software Development Antipattern.&#8221; DevIQ.com. <a href="https://deviq.com/antipatterns/big-design-up-front/">https://deviq.com/antipatterns/big-design-up-front/</a></p><p>[15] Jones-Mitchell, J. &#8220;How AI Killed the Agile Process (And Why That&#8217;s a Good Thing).&#8221; HumanDrivenAI, December 2024. <a href="https://humandrivenai.com/2024/12/04/how-ai-killed-the-agile-process-and-why-thats-a-good-thing/">https://humandrivenai.com/2024/12/04/how-ai-killed-the-agile-process-and-why-thats-a-good-thing/</a></p><p>[16] Royce, W.W. &#8220;Managing the Development of Large Software Systems.&#8221; Proceedings of IEEE WESCON, 1970.</p><p>[17] Requiment. &#8220;Why Do Software Development Projects Fail?&#8221; Requiment.com. <a href="https://www.requiment.com/why-do-software-development-projects-fail/">https://www.requiment.com/why-do-software-development-projects-fail/</a></p>]]></content:encoded></item><item><title><![CDATA[AI is solving the easy half of software quality]]></title><description><![CDATA[Most of our engineering practices only cover half the picture. AI is widening the gap.]]></description><link>https://a4al6a.substack.com/p/ai-is-solving-the-easy-half-of-software</link><guid isPermaLink="false">https://a4al6a.substack.com/p/ai-is-solving-the-easy-half-of-software</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Sun, 29 Mar 2026 12:10:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8mxA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every system we build does things. Some of those things it should do. Some it shouldn&#8217;t. This sounds obvious, but when we sit with it for a moment, we realise most of our engineering practices only address half the picture.</p><p>And AI-assisted development is about to make that imbalance much, much worse.</p><p>I&#8217;ve been thinking about this as a simple two-by-two grid. On one axis: does the system do the thing, or not? On the other: should it?</p><p>That gives us four quadrants:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8mxA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8mxA!, /__u/a4al6a.substack.com/w_424, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png 424w, /__u/substackcdn.com/image/fetch/$s_!8mxA!, /__u/a4al6a.substack.com/w_848, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png 848w, /__u/substackcdn.com/image/fetch/$s_!8mxA!, /__u/a4al6a.substack.com/w_1272, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8mxA!, /__u/a4al6a.substack.com/w_1456, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8mxA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png" width="1440" height="848" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:848,&quot;width&quot;:1440,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:74503,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://a4al6a.substack.com/i/192498919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8mxA!, /__u/a4al6a.substack.com/w_424, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png 424w, /__u/substackcdn.com/image/fetch/$s_!8mxA!, /__u/a4al6a.substack.com/w_848, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png 848w, /__u/substackcdn.com/image/fetch/$s_!8mxA!, /__u/a4al6a.substack.com/w_1272, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8mxA!, /__u/a4al6a.substack.com/w_1456, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4bdba3f-227a-460b-bed2-81501e9a2cbd_1440x848.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Functional</strong>: the system should do it, and it does. This is the happy place. Our login form authenticates users. Our checkout calculates the right total. We wrote a test, the test passes, everyone goes home.</p><p><strong>Missing</strong>: the system should do it, but it doesn&#8217;t. This is where backlogs live. Features not yet built. Edge cases not yet handled. It&#8217;s a known gap, and known gaps are manageable. We can prioritise them, plan for them, write a user story.</p><p><strong>Faulty</strong>: the system should not do it, but it does anyway. This is where things get uncomfortable.</p><p><strong>Safe</strong>: the system should not do it, and it doesn&#8217;t. This is where things get invisible.</p><p>Most of our engineering effort goes into the top half of this grid. We write requirements that describe what the system should do. We write tests that verify it does those things. We track missing functionality in backlogs and roadmaps. We have entire events dedicated to deciding what to build next.</p><p>But the bottom half? That&#8217;s where the real trouble lives. And it&#8217;s the half that AI coding tools don&#8217;t touch.</p><h2>AI is brilliant at the top half</h2><p>Let&#8217;s be fair. AI-assisted development is genuinely transforming the functional and missing quadrants.</p><p>Need to build a login form? An AI coding agent can produce one in seconds, complete with validation, error handling, and tests. Got a backlog of missing features? AI can chew through them faster than any team could before. The top half of the grid is where AI shines, because it&#8217;s a space defined by clear, expressible intent. We can describe what we want. The model can generate it. We can verify it works.</p><p>This is real progress. Teams are shipping faster. Backlogs are shrinking. Features that would have taken days are landing in hours. If our definition of productivity is &#8220;how quickly can we build the things we&#8217;ve decided to build,&#8221; then AI is a genuine leap forward.</p><p>But productivity isn&#8217;t the same as quality. And quality lives in all four quadrants, not just the top two.</p><h2>Why faulty is so hard to find</h2><p>The fundamental problem with faulty behaviour is that we can&#8217;t write a list of everything a system shouldn&#8217;t do. The space is, for all practical purposes, infinite.</p><p>We can write a requirement that says &#8220;the system should let users reset their password.&#8221; We can test that. But can we write a requirement that says &#8220;the system should not let users reset someone else&#8217;s password&#8221;? We could, but that&#8217;s just one of a thousand things the password reset flow shouldn&#8217;t do. It also shouldn&#8217;t expose the reset token in the URL. It shouldn&#8217;t send the new password in plain text. It shouldn&#8217;t allow a reset request to be replayed. It shouldn&#8217;t let someone enumerate valid email addresses by observing response times.</p><p>Each of those is a specific &#8220;should not&#8221; that someone had to think of. And someone had to think of it because, at some point, a system somewhere did exactly that.</p><p>This is why security vulnerabilities are so hard to eliminate. A vulnerability is almost always a case of the system doing something it was never intended to do. Not a missing feature. Not a broken feature. An extra, unwanted behaviour that emerged from the interaction between components that individually work fine.</p><p>And that&#8217;s the second reason faulty behaviour is hard to catch: it&#8217;s emergent. Our authentication module works. Our session management works. Our API gateway works. But the specific way they interact under a particular sequence of requests creates a behaviour that none of them were designed to produce. We can&#8217;t find it by testing each component in isolation. We can only find it by understanding the system as a whole, and the system as a whole is more complex than any single person can hold in their head.</p><p>There&#8217;s also a subtler problem. Faulty behaviour often looks like correct behaviour from the inside. The system is doing exactly what the code tells it to do. It&#8217;s following its instructions perfectly. The fault isn&#8217;t in the execution; it&#8217;s in the gap between what we told the system to do and what we meant for it to do. Every line of code is an instruction the system will follow faithfully. We just didn&#8217;t always think through the consequences of every instruction we gave.</p><h2>Why AI makes the bottom half worse</h2><p>Here&#8217;s the thing that nobody in the &#8220;AI will 10x your productivity&#8221; conversation wants to talk about: every line of code an AI generates is also a line of code that might do something we didn&#8217;t ask for.</p><p>When a human writes code slowly, there&#8217;s a natural friction that acts as a filter. We think about what we&#8217;re writing. We consider edge cases. We&#8217;ve been burned before by a particular pattern, so we avoid it. We have institutional memory, battle scars, and instincts shaped by years of production incidents.</p><p>An AI coding agent has none of that. It generates code that satisfies the stated requirement, and it does so fluently and confidently. But it has no concept of &#8220;should not.&#8221; It wasn&#8217;t there when our system went down at 3am because of a race condition. It doesn&#8217;t know that our payment provider&#8217;s API behaves differently in sandbox mode. It doesn&#8217;t understand that the specific way it structured a database query makes it vulnerable to timing attacks.</p><p>An AI can produce a perfectly functional password reset flow that also happens to leak information through response times. Every test passes. The feature works. The faulty behaviour is invisible to the test suite because nobody thought to test for it, and the AI certainly didn&#8217;t flag it.</p><p>And the volume problem makes this worse. If AI lets us ship ten times more code, we&#8217;ve also shipped ten times more surface area for faulty behaviour. The top half of the grid fills up faster. The bottom half gets more dangerous at exactly the same rate.</p><h2>Why safe is so hard to confirm</h2><p>If faulty behaviour is hard to find, safe behaviour is hard to even think about. How do we verify the absence of something? How do we test that something doesn&#8217;t happen?</p><p>We can test that a non-admin user cannot access the admin panel. That&#8217;s a specific assertion about a specific &#8220;should not.&#8221; But the number of things a non-admin user should not be able to do is essentially unbounded. They shouldn&#8217;t be able to modify other users&#8217; data. They shouldn&#8217;t be able to trigger a database export. They shouldn&#8217;t be able to craft a request that causes the server to leak its environment variables. Each of these is a test we could write, but only if we&#8217;ve already imagined the specific threat.</p><p>This is the core difficulty: the safe quadrant is defined by absence, and absence is invisible. Nobody notices the things a system doesn&#8217;t do until it starts doing them. The safe quadrant only becomes visible when something moves out of it and into the faulty quadrant. By then, we&#8217;ve got a production incident.</p><p>There&#8217;s also the problem of regression. Something that was safely excluded yesterday can quietly creep in today. A new dependency introduces a behaviour we didn&#8217;t ask for. A refactor changes the boundary conditions. A configuration change opens a path that was previously closed. The safe quadrant isn&#8217;t stable. It needs constant maintenance, but because we can&#8217;t see it, we don&#8217;t know when it&#8217;s eroding until something breaks through.</p><p>AI-generated code accelerates this erosion. When an AI agent refactors a module or introduces a new dependency, it&#8217;s optimising for the stated goal. It&#8217;s not thinking about what the previous version of the code was quietly preventing. The old code might have been clunky, but its clumsiness might have been the very thing keeping a dangerous path closed. The AI tidies it up, the tests still pass, and something that was safe is now faulty. Nobody notices until it matters.</p><h2>What this means for how we work</h2><p>Most of our testing practices are built for the top half of the grid. Unit tests verify that functions produce the right output for given inputs. Acceptance tests verify that the system exhibits the behaviours described in user stories. Both are firmly in the &#8220;should do&#8221; world.</p><p>To address the bottom half, we need different tools and, more importantly, a different mindset.</p><p>Property-based testing is one approach. Instead of testing specific inputs and outputs, we define invariants that should hold across all possible inputs. &#8220;No matter what string we pass to this function, the output should never contain unescaped HTML.&#8221; That&#8217;s a &#8220;should not&#8221; expressed as a universal constraint. It doesn&#8217;t cover everything, but it covers more than a handful of example-based tests.</p><p>Observability helps too. If we can&#8217;t enumerate everything the system shouldn&#8217;t do in advance, we can at least watch what it actually does in production and look for surprises. Anomaly detection, audit logging, runtime assertions. These are tools for spotting faulty behaviour after the fact, which isn&#8217;t as good as preventing it, but it&#8217;s far better than not spotting it at all.</p><p>Threat modelling is explicitly about the bottom half. We sit down and ask: what could go wrong? What could this system do that we don&#8217;t want it to? It&#8217;s a structured way of populating the &#8220;should not&#8221; column, and it&#8217;s one of the few practices that directly targets the faulty quadrant.</p><p>But perhaps the most important thing is simply to recognise that the bottom half exists. Most teams spend all their energy making the system do the right things and almost none making sure it doesn&#8217;t do the wrong things. The backlog is full of features to build. It&#8217;s never full of behaviours to prevent.</p><p>And that asymmetry is where the risk lives.</p><h2>The uncomfortable truth about AI and quality</h2><p>We&#8217;re quite good at building systems that do what they should. We&#8217;re not nearly as good at building systems that don&#8217;t do what they shouldn&#8217;t. And the reason is structural, not just a matter of effort. The &#8220;should do&#8221; space is finite and enumerable. The &#8220;should not&#8221; space is infinite and mostly invisible.</p><p>AI is an amplifier. It amplifies whatever we&#8217;re already doing. If we&#8217;re focused on building features, AI will help us build features faster than ever. If we&#8217;re focused on preventing unwanted behaviour, AI can help with that too. But almost nobody is focused on preventing unwanted behaviour, because our entire industry is structured around the top half of the grid.</p><p>So what AI actually amplifies, in practice, is the imbalance. More code, faster. More features, sooner. More surface area for things to go wrong, with no corresponding increase in our ability to spot them.</p><p>This doesn&#8217;t mean we should stop using AI. It means we should stop measuring AI&#8217;s impact purely by how fast it fills the top half of the grid. The real question isn&#8217;t &#8220;how quickly can we build?&#8221; It&#8217;s &#8220;how confident are we that what we built only does what it should?&#8221;</p><p>A green test suite means the system does the things we thought to check. It says nothing about the things we didn&#8217;t think to check. And when AI is writing the code, the gap between &#8220;what we checked&#8221; and &#8220;what the code actually does&#8221; grows wider with every prompt.</p><p>The four quadrants are a simple model, but they make one thing very clear: half of system quality is about what the system does, and the other half is about what it doesn&#8217;t. AI is making the first half easier and the second half harder. And until we take the second half as seriously as the first, faster will just mean faster into trouble.</p><h2>Conclusions</h2><p>The next time someone tells us AI will revolutionise software development, we should ask: which half? Because the half it&#8217;s good at was never the half we were struggling with.</p><p>We don&#8217;t need to build faster. We need to think wider. The quadrant that will hurt us most is the one we never thought to look at, and AI won&#8217;t look at it for us.</p><p>The most dangerous line of code isn&#8217;t the one that fails a test. It&#8217;s the one that passes every test and still does something nobody asked for. AI writes a lot of those lines now. We should probably start paying attention to that.<br><br>&#8221;How?&#8221; is a good question.</p>]]></content:encoded></item><item><title><![CDATA[The Change Advisory Board Must Die]]></title><description><![CDATA[How a 1980s governance relic is destroying your software delivery performance &#8212; and the psychology that keeps it alive]]></description><link>https://a4al6a.substack.com/p/the-change-advisory-board-must-die</link><guid isPermaLink="false">https://a4al6a.substack.com/p/the-change-advisory-board-must-die</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Sat, 21 Mar 2026 00:39:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3><strong>Abstract</strong></h3><p>Change Advisory Boards (CABs) were designed in the late 1980s to bring order to chaotic IT environments. Nearly four decades later, the largest empirical study of software delivery performance ever conducted &#8212; spanning over 23,000 responses from more than 2,000 organizations &#8212; has delivered a damning verdict: <strong>external change approval is worse than having no change approval process at all.</strong> CABs do not reduce failures. They slow delivery, encourage riskier deployments, erode team ownership, and create a false sense of security. Yet they persist &#8212; not because they work, but because of identity threat, loss aversion, and institutional self-preservation. This article examines the evidence, explores the psychology, and argues that the deployment pipeline &#8212; powered by TDD, automated testing, and progressive delivery &#8212; is the only &#8220;change advisory board&#8221; that actually works.</p><div><hr></div><h3><strong>TL;DR</strong></h3><p>The DORA research program (23,000+ responses, 2,000+ organizations, 2014-2024) found that external change approval by a CAB is negatively correlated with lead time, deployment frequency, and restore time, and has <strong>no correlation with change fail rate</strong> &#8212; the very metric it exists to improve. Organizations with heavyweight approval processes are <strong>2.6x more likely to be low performers</strong>. The UK Financial Conduct Authority found that CABs in financial services <strong>approved over 90% of changes</strong> while change-related incidents remained a top cause of disruption. Even ITIL 4 &#8212; the framework that invented CABs &#8212; has made them optional. The alternative is not anarchy: it&#8217;s peer review, TDD, automated pipelines, feature flags, and team ownership. CABs survive not because they work, but because removing them triggers identity threat in the people whose roles depend on them.</p><div><hr></div><h2><strong>1. A Solution from Another Era</strong></h2><p>In the late 1980s, the British Government had a problem. IT services were unreliable, changes were chaotic, and a single bad deployment could bring down an entire mainframe environment. The Central Computer and Telecommunications Agency (CCTA) responded by creating the IT Infrastructure Library (ITIL) &#8212; a framework that included, among many processes, the Change Advisory Board.</p><p>The CAB made sense in that world. Changes were infrequent and large &#8212; quarterly or annual releases. Infrastructure was physical and changes were difficult to reverse. There was no automated testing, no CI/CD, no easy rollback. <strong>The cost of failure was extremely high and recovery was slow.</strong> A committee of cross-functional stakeholders evaluating changes before they went live was a reasonable response to these constraints.</p><p>But that world no longer exists.</p><p>Today, elite-performing engineering organizations deploy code multiple times per day. They use automated testing suites that run thousands of checks in minutes. They deploy behind feature flags, roll out changes to 1% of users first, and roll back automatically if error rates spike. They have observability platforms that make production behavior visible in real time.</p><p><strong>And yet, in thousands of organizations, a committee still meets once a week to review change request forms and rubber-stamp deployments.</strong></p><div><hr></div><h2><strong>2. What the Data Actually Says</strong></h2><p>The DevOps Research and Assessment (DORA) team &#8212; founded by Nicole Forsgren, Jez Humble, and Gene Kim &#8212; has conducted the largest and longest-running research program on software delivery performance. Their research spans multiple years (2014-2024), has collected over 23,000 survey responses from more than 2,000 organizations across all sizes and industries, and uses peer-reviewed statistical analysis.</p><p>Their findings on change approval are unambiguous:</p><blockquote><p><strong>&#8220;We found that external approvals were negatively correlated with lead time, deployment frequency, and restore time, and had no correlation with change fail rate. In short, approval by an external body (such as a manager or CAB) simply doesn&#8217;t work to increase the stability of production systems, measured by the time to restore service and change fail rate. However, it certainly slows things down. It is, in fact, worse than having no change approval process at all.&#8221;</strong> &#8212; Forsgren, Humble, Kim, <em>Accelerate</em> (2018)</p></blockquote><p>Read that again. <strong>External approval by a CAB is worse than having no change approval process at all.</strong> This is not an opinion from a DevOps evangelist. It is the conclusion of the largest empirical study of software delivery ever conducted.</p><p>The 2019 State of DevOps Report added further specificity:</p><ul><li><p><strong>Peer review during development, supplemented by automation, is the most effective change approval approach.</strong></p></li><li><p><strong>No evidence was found supporting the hypothesis that formal, external review processes reduce change failure rates.</strong></p></li><li><p><strong>Organizations using heavyweight approval processes are 2.6 times more likely to be low performers.</strong></p></li><li><p>When team members have a clear understanding of the change approval process, this drives higher performance &#8212; <strong>the process clarity matters more than the process weight.</strong></p></li></ul><div><hr></div><h2><strong>3. The Regulator&#8217;s Own Data</strong></h2><p>If the DORA findings seem too theoretical, consider what happens when a financial regulator looks at the actual numbers.</p><p>The UK Financial Conduct Authority published &#8220;Implementing Technology Change&#8221; in February 2021, analyzing over 1 million production changes across UK financial institutions. What they found should alarm anyone who relies on a CAB for safety:</p><ul><li><p><strong>CABs approved over 90% of the major changes they reviewed.</strong></p></li><li><p><strong>Some firms had not rejected a single change during all of 2019.</strong></p></li><li><p>Despite this, <strong>change-related incidents were consistently one of the top causes of failure and operational disruption.</strong></p></li><li><p>Of high-severity customer-facing incidents, <strong>24% had change-related root causes.</strong></p></li><li><p>The sampled firms deployed nearly 68,000 major changes over 2019, resulting in 2,600 incidents. Major changes had a failure rate of <strong>3.8%</strong> &#8212; more than double the standard rate.</p></li><li><p>Firms using <strong>smaller, more frequent releases</strong> had higher change success rates.</p></li><li><p>Firms demonstrating <strong>greater adoption of agile methodologies</strong> were less likely to suffer change-related incidents.</p></li></ul><p>This is a regulator&#8217;s own data confirming that <strong>CABs are rubber stamps that do not prevent the failures they exist to prevent.</strong> A 90%+ approval rate is not a quality gate &#8212; it is a formality. The FCA found the same thing DORA found: it is not committees that reduce change risk, but better engineering practices &#8212; smaller batches, more automation, more testing, more frequent deployment.</p><div><hr></div><h2><strong>4. Eight Ways CABs Hurt You</strong></h2><h3><strong>They create bottlenecks</strong></h3><p>CABs typically meet on a fixed schedule &#8212; often weekly. A change that takes one hour to develop may wait days for approval. Stuart Rance, an ITIL practitioner, quantifies it directly: <strong>&#8220;If you hold a CAB meeting once a week, and change requests have to be submitted three days before the CAB meeting, then you delay every change by up to 10 days.&#8221;</strong></p><p>If a CAB meets weekly and reviews 20-30 changes in an hour, each change receives approximately 2-3 minutes of scrutiny. This is neither thorough review nor fast flow &#8212; <strong>it is the worst of both worlds.</strong></p><h3><strong>They reduce deployment frequency</strong></h3><p>Elite-performing teams deploy on demand, multiple times per day. A weekly CAB meeting caps deployment frequency at once per week at most &#8212; <strong>placing a hard ceiling on team performance.</strong></p><h3><strong>They do NOT improve stability</strong></h3><p>This is the most damning finding. The entire justification for CABs is that they reduce the risk of production failures. <strong>The data says they don&#8217;t.</strong> External approvals had no correlation with change fail rate (DORA, 2014-2019). None.</p><h3><strong>They shift accountability away from teams</strong></h3><p>When a CAB approves a change, accountability becomes diffused. The team is no longer solely responsible &#8212; the CAB &#8220;approved&#8221; it. <strong>The feedback loop between writing code and owning its impact in production is broken.</strong> The incentive to invest in quality shifts from &#8220;I own this&#8221; to &#8220;the CAB will catch problems.&#8221;</p><h3><strong>They encourage larger, riskier deployments</strong></h3><p>This is the critical paradox. CABs are designed to reduce risk, but by reducing deployment frequency, they force teams to batch changes. <strong>Larger batches are inherently more complex, harder to test, harder to understand, and harder to roll back.</strong> CABs often increase the risk of painful deployment failures as a direct result of the constraints they impose.</p><blockquote><p><strong>&#8220;Ask a programmer to review ten lines of code, he&#8217;ll find ten issues. Ask him to do five hundred lines, and he&#8217;ll say it looks good.&#8221;</strong> &#8212; Giray Ozil</p></blockquote><p>This is precisely the dynamic CABs create: large batches reviewed superficially versus small batches reviewed thoroughly.</p><h3><strong>They create learned helplessness</strong></h3><p>Engineers who must get deployments approved by committees <strong>start self-censoring innovative ideas.</strong> They tell other engineers that their features &#8220;won&#8217;t pass the committees&#8217; checks,&#8221; creating a culture of pre-emptive surrender. This creates a vicious cycle: <strong>the CAB reduces team autonomy, which reduces team ownership, which reduces quality, which &#8220;justifies&#8221; more CAB oversight.</strong></p><h3><strong>They inflate lead time</strong></h3><p>Elite performers have lead times of less than one hour; low performers have lead times of one to six months. <strong>Lead time is negatively correlated with external approval processes.</strong> CABs push organizations toward the latter.</p><h3><strong>They damage developer morale</strong></h3><p>Waiting for CAB approval is demoralizing for engineers who have already completed their work. The process signals distrust: <strong>&#8220;We don&#8217;t trust you to deploy your own code.&#8221;</strong> Filling out change request forms is widely perceived as bureaucratic busywork. The DORA research found that heavyweight change management processes contribute to burnout.</p><p>Nicole Forsgren noted: <strong>&#8220;The report has found year after year that a slower change management process can paradoxically result in more instability, not less.&#8221;</strong></p><div><hr></div><h2><strong>5. What the Industry&#8217;s Best Minds Say</strong></h2><p>The critique of CABs is not a fringe position. It comes from the most respected voices in software engineering:</p><p><strong>Jez Humble</strong> (DORA co-founder, co-author of <em>Continuous Delivery</em>):</p><blockquote><p><strong>&#8220;Most change management is Change Management Theatre. It is not about making things better, it is about covering your ass when things go wrong.&#8221;</strong></p><p>External change approval is <strong>&#8220;risk management theater&#8221;</strong> &#8212; like TSA&#8217;s enhanced airport security, <strong>&#8220;it accomplishes nothing at enormous cost, giving the impression that you&#8217;re managing the risk of making changes to the production environment, while actually making the situation worse.&#8221;</strong></p><p><strong>&#8220;How are outside reviewers expected to review thousands of lines of code by hundreds of programmers and predict its impact on the production environment? It can&#8217;t possibly work, which is why it doesn&#8217;t work.&#8221;</strong></p></blockquote><p><strong>Dave Farley</strong> (co-author of <em>Continuous Delivery</em>):</p><blockquote><p><strong>&#8220;Complex approaches to gatekeeping, like &#8216;Change Approval Boards&#8217; are negatively correlated with software quality, with the more complex compliance mechanisms around change associated with lower quality software.&#8221;</strong></p></blockquote><p><strong>Gene Kim</strong> (co-author of <em>The Phoenix Project</em>, <em>The DevOps Handbook</em>):</p><blockquote><p><strong>&#8220;If every time we want to make a code change we have to send our engineers to scores of committee meetings in order to get permission to make our changes&#8221;</strong> &#8212; the countermeasure is to <strong>&#8220;create more loosely-coupled architecture so that changes can be made safely and with more autonomy.&#8221;</strong></p></blockquote><p><strong>Charity Majors</strong> (co-founder of Honeycomb):</p><blockquote><p><strong>&#8220;The path to high performance is building better sociotechnical systems (fast deploys, observability, guard rails) that enable normal people to do great work.&#8221;</strong></p></blockquote><p><strong>Nicole Forsgren</strong> (DORA co-founder):</p><blockquote><p><strong>&#8220;Knowledge is power, and you should give power to those who have the knowledge.&#8221;</strong></p></blockquote><div><hr></div><h2><strong>6. The Proof: Organizations That Moved Beyond CABs</strong></h2><p>These are not startups. These are large, often regulated organizations that abandoned or dramatically reformed their CABs &#8212; and saw their performance soar.</p><p><strong>ING Bank</strong> went from three separate CABs (planning, technical, deployment) to one lightweight deployment advisory board. The risk value per change decreased, and the ratio of incidents per change dropped significantly. Mark Heistek, IT specialist at ING, described the old CABs as being <strong>&#8220;just like an insurance company, always rejecting your first claim.&#8221;</strong></p><p><strong>Barclays Bank</strong> &#8212; a 328-year-old institution &#8212; under Jon Smart&#8217;s leadership, delivered <strong>3x as much in 1/3 of the time</strong> with <strong>23x fewer production incidents</strong> and the <strong>highest ever employee engagement scores.</strong> These results came from empowering teams, not from adding more approval gates.</p><p><strong>Target</strong> went from deploying a few times per month to <strong>hundreds of times per day</strong> through their DevOps transformation.</p><p><strong>Amazon</strong> deploys code <strong>every 11-12 seconds on average.</strong> Every engineer has full ownership of their services. No centralized CAB for routine deployments.</p><p><strong>Netflix</strong> gives every engineer <strong>full access to the production environment from day one.</strong> Culture of accountability and continuous learning replaces approval gates.</p><p><strong>Facebook/Meta</strong> moved to quasi-continuous &#8220;push from master&#8221; in 2017. An IEEE-published study found that continuous deployment did not inhibit productivity or quality even as the engineering team scaled by 20x and the codebase by 50x.</p><div><hr></div><h2><strong>7. Even ITIL Agrees</strong></h2><p>Here is perhaps the most uncomfortable fact for CAB defenders: <strong>even ITIL itself &#8212; the framework that created the CAB &#8212; has moved on.</strong></p><p>ITIL 4, released in 2019, makes sweeping changes:</p><ol><li><p><strong>Renamed from &#8220;Change Management&#8221; to &#8220;Change Enablement.&#8221;</strong> The name itself signals a philosophical shift from controlling to enabling.</p></li><li><p><strong>CAB is no longer mandatory.</strong> ITIL 4 introduces the concept of a &#8220;Change Authority&#8221; that can be different for different types of changes and does not need to be a committee.</p></li><li><p><strong>Aligned with Agile, DevOps, and CI/CD.</strong> ITIL 4 explicitly recognizes that change enablement should enable rather than restrict.</p></li><li><p><strong>Automation encouraged.</strong> Unlike ITIL v3, which relied on manual approvals, ITIL 4 encourages automated change pipelines.</p></li><li><p><strong>Standard changes are pre-approved.</strong> Low-risk, repeatable changes do not require individual approval.</p></li></ol><p><strong>Organizations that defend their CABs by citing ITIL compliance are, paradoxically, not compliant with the current version of ITIL.</strong> ITIL 4 supports exactly the approach that DORA research recommends: risk-based differentiation, automation, peer review, and team empowerment.</p><div><hr></div><h2><strong>8. The Real Change Advisory Board: TDD and the Deployment Pipeline</strong></h2><p>If CABs don&#8217;t work, what does? The answer is not anarchy &#8212; it is engineering discipline.</p><p><strong>Test-Driven Development</strong> creates a comprehensive safety net that replaces the need for external approval. A team practicing TDD does not need to <em>ask permission</em> to deploy. They can <em>demonstrate</em> that their change is safe. The test suite is the evidence. The pipeline is the judge.</p><p>The evidence is strong. A landmark study by Nagappan et al. at IBM and Microsoft found that TDD reduced pre-release defect density by <strong>40-90%</strong> compared to similar projects not using TDD. DORA&#8217;s research found that test automation drives improved software stability, reduced team burnout, and lower deployment pain. DORA explicitly recommends that developers should <strong>&#8220;write unit tests before writing production code for all changes to the codebase&#8221;</strong> &#8212; a direct endorsement of TDD.</p><p><strong>Robert C. Martin</strong> offers a powerful analogy: <strong>&#8220;I want you to think of TDD the way accountants think of dual entry bookkeeping.&#8221;</strong> Both are disciplines where every intention is entered in two places and must ultimately reconcile. CABs are <em>bureaucratic</em> discipline &#8212; a committee imposing external control. TDD is <em>professional</em> discipline &#8212; teams imposing rigorous standards on themselves. The latter is both more effective and more empowering.</p><p><strong>Kent Beck</strong>, the creator of TDD: <strong>&#8220;Write tests until fear is transformed into boredom.&#8221;</strong> And: <strong>&#8220;Rather than apply minutes of suspect reasoning, we can just ask the computer by making the change and running the tests.&#8221;</strong></p><p>The deployment pipeline &#8212; as described by Humble and Farley in <em>Continuous Delivery</em> &#8212; is the automated mechanism that enforces this discipline. Dave Farley argues it is <strong>&#8220;the perfect vehicle for compliance&#8221;</strong> and that he has seen organizations move <strong>&#8220;from taking weeks, sometimes months, to ensure that releases were &#8216;compliant,&#8217; to generating genuinely compliant release candidates multiple times per day.&#8221;</strong></p><p>Consider the properties:</p><p>PropertyDeployment PipelineCAB<strong>Consistency</strong>Runs the same checks every timeAttention varies by fatigue, time pressure, who&#8217;s in the room<strong>Speed</strong>Feedback in minutesFeedback in days or weeks<strong>Objectivity</strong>Pass or fail &#8212; no politicsSusceptible to rubber-stamping, politics, bikeshedding<strong>Comprehensiveness</strong>Runs thousands of testsReviews a change description document<strong>Audit trail</strong>Automatically records what was tested, when, and the resultMeeting minutes at best</p><h3><strong>What about business risk?</strong></h3><p>There is a legitimate distinction between technical validation (&#8221;are the tests passing?&#8221;) and business-risk approval (&#8221;is this safe to release now?&#8221;). A technically correct feature can still be a bad business decision to release on Black Friday.</p><p>Modern practices address this without a committee:</p><ul><li><p><strong>Feature flags</strong> separate deployment from release. You deploy code to production without any user seeing it until a deliberate release decision is made.</p></li><li><p><strong>Progressive delivery</strong> (canary releases, percentage rollouts) manages blast radius mechanically, based on real production data rather than pre-deployment guesswork.</p></li><li><p><strong>Observability</strong> makes production behavior visible in real time, enabling detection and response in minutes.</p></li></ul><p>DORA itself acknowledges this distinction. It recommends CABs transform <strong>&#8220;from gatekeeper to process architect,&#8221;</strong> focusing on strategic business decisions. <strong>The research does not say business risk assessment is unnecessary &#8212; it says the traditional CAB is the wrong mechanism for it.</strong></p><p>For the rare change that genuinely requires business-risk assessment &#8212; a major infrastructure migration, a change with regulatory notification requirements &#8212; a lightweight, asynchronous, expert-driven review is the proportionate response. Not a weekly committee meeting. And certainly not for the 95%+ of changes that are standard code deployments.</p><div><hr></div><h2><strong>9. The Compliance Theater Problem</strong></h2><p>Jez Humble cuts to the heart of it:</p><blockquote><p><strong>&#8220;Most change management is Change Management Theatre. It is not about making things better, it is about covering your ass when things go wrong.&#8221;</strong></p></blockquote><p>CABs create an illusion of control through several mechanisms:</p><p><strong>The paper trail illusion.</strong> Change request forms create documentation. But documentation of a bad change does not prevent the bad change &#8212; it merely records that someone approved it.</p><p><strong>The expert review illusion.</strong> CAB members are positioned as experts who will catch problems. But <strong>CAB members rarely have the codebase knowledge, context, or time to meaningfully evaluate individual changes.</strong> As Humble notes, heavyweight processes often result in forms being sent to approvers in another country who &#8220;don&#8217;t know anything about what they&#8217;re approving.&#8221;</p><p><strong>The approval rate illusion.</strong> If a CAB approves 90%+ of changes &#8212; as the FCA found across UK financial services &#8212; the process is not filtering. It is rubber-stamping. <strong>A 90% approval rate means the CAB either lacks the ability to identify bad changes or lacks the courage to reject them.</strong> Either way, it is not performing its stated function.</p><p><strong>The accountability illusion.</strong> The CAB provides someone to blame when things go wrong (&#8221;the CAB approved it&#8221;), but this diffused accountability reduces the incentive for teams to take ownership of quality.</p><p>The honest motivation behind many CABs is not risk reduction but <strong>liability protection.</strong> The real purpose is ensuring that when a failure occurs, someone can point to a signed-off change request form and say &#8220;I followed the process.&#8221; This protects individuals from blame but does nothing to prevent the failure itself.</p><div><hr></div><h2><strong>10. The Psychology: Why CABs Refuse to Die</strong></h2><p>If the evidence is this clear, why do CABs persist? The answer lies not in technology or process, but in psychology, power, and identity.</p><h3><strong>Identity threat</strong></h3><p>When organizational changes threaten the value of professional roles, employees experience what researchers call <strong>identity threat</strong> &#8212; &#8220;an experience appraised as indicating potential harm to the value, meanings, or enactment of an identity&#8221; (Petriglieri, 2011). For CAB members who have spent years building their identity around being &#8220;the person who approves changes to production,&#8221; any suggestion that this role is unnecessary strikes at their core self-concept.</p><p>Research shows that identity threat leads to <strong>&#8220;exemplification behaviors&#8221;</strong> &#8212; performative actions designed to demonstrate one&#8217;s value. This maps directly to CAB members who, when threatened, <em>intensify</em> their gatekeeping &#8212; demanding more documentation, asking more questions, slowing down approvals &#8212; to demonstrate their indispensability.</p><h3><strong>Loss aversion and status quo bias</strong></h3><p>Kahneman and Tversky&#8217;s Prospect Theory established that <strong>losses loom larger than gains.</strong> Even when presented with DORA research showing that CABs make things <em>worse</em>, the cognitive calculus for those invested in the CAB is asymmetric: <strong>the potential loss of role, status, and identity weighs more heavily than the potential gain of organizational performance.</strong></p><p>For someone who has championed the CAB for years, admitting it does not work also means admitting their past efforts were wasted (sunk cost), that they were wrong (cognitive dissonance), and that they will lose influence (control).</p><h3><strong>The ivory tower: authority without expertise</strong></h3><p>French and Raven&#8217;s framework on the Bases of Power is directly applicable. CAB members exercise <strong>legitimate power</strong> (authority from position) rather than <strong>expert power</strong> (authority from knowledge). The CAB inverts the natural power hierarchy: <strong>those with the least domain knowledge hold the most decision-making authority over the change.</strong></p><p>Nicole Forsgren put it simply: <strong>&#8220;Knowledge is power, and you should give power to those who have the knowledge.&#8221;</strong> The CAB does the opposite.</p><h3><strong>The bus factor</strong></h3><p>CABs concentrate decision-making authority in a small group &#8212; often 3 to 7 people. When key members are unavailable, the entire deployment pipeline stalls. <strong>This is precisely the fragility that CABs claim to prevent.</strong> If the CAB chair is on holiday when a critical security patch needs to go out, the organization faces a choice: bypass the process or wait. Either way, the CAB has failed.</p><h3><strong>Theory X incarnate</strong></h3><p>Douglas McGregor&#8217;s Theory X assumes workers &#8220;dislike their work, avoid responsibility, and need constant direction.&#8221; Theory Y assumes work is natural and people will exercise self-direction when committed to objectives. <strong>CABs are a pure Theory X mechanism.</strong> They exist because of an implicit assumption that developers cannot be trusted to deploy responsibly.</p><p>McGregor identified the self-fulfilling prophecy: if you treat engineers as people who need external supervision, they internalize that belief and stop developing their own judgment. <strong>The CAB creates the dependency it then uses to justify its own existence.</strong></p><h3><strong>The meta-irony</strong></h3><p>The Change Advisory Board resists change to itself. Bureaucratic inertia &#8212; &#8220;the tendency of bureaucratic organizations to perpetuate established procedures, resisting change irrespective of shifts in the environment&#8221; &#8212; ensures that every production incident becomes evidence that <em>more</em> oversight is needed, never that the oversight itself is contributing to the problem.</p><p>Conway&#8217;s Law compounds this: a CAB creates organizational roles (change manager, release coordinator, CAB secretary) that create constituencies whose job survival depends on the CAB&#8217;s continued existence. <strong>The process creates the roles. The roles defend the process.</strong></p><p>David Graeber&#8217;s taxonomy of meaningless work applies directly. CAB reviewers who rubber-stamp changes they cannot evaluate are <strong>box tickers</strong> &#8212; performing roles &#8220;meant to make a firm appear to be doing something it&#8217;s not actually doing.&#8221; CABs that supervise change processes which would perform better without supervision are <strong>taskmasters</strong> &#8212; &#8220;people who supervise people who don&#8217;t need supervision.&#8221;</p><div><hr></div><h2><strong>11. What Defenders Say &#8212; and Why They&#8217;re Mostly Wrong</strong></h2><p>It would be intellectually dishonest to pretend that no reasonable person defends CABs.</p><p><strong>Gareth Davies of Lloyds Banking Group</strong> claims the CAB has kept them off the BBC front page. <strong>James Finister</strong>, an ITIL co-author, argues that &#8220;CABs stop the corridor conversations and the changes that bring an organization to its knees.&#8221; <strong>Stuart Rance</strong>, an ITIL practitioner, offers a nuanced middle ground: daily code deployments should bypass the CAB, but major infrastructure changes &#8220;like migrating cloud providers&#8221; may warrant broad stakeholder discussion.</p><p>These are fair points. But they conflate several different functions. The question is not whether organizations need risk assessment, cross-team coordination, or change visibility &#8212; they clearly do. <strong>The question is whether a weekly committee meeting is the best mechanism for delivering those functions.</strong> The DORA data says it is not.</p><p>CABs may serve functions beyond filtering: risk communication, stakeholder alignment, regulatory evidence, escalation. But <strong>a weekly meeting that adds up to 10 days of lead time, reduces deployment frequency, encourages larger batches, creates learned helplessness, and has no measurable effect on change fail rate is an expensive mechanism for functions that lighter alternatives can serve more effectively.</strong></p><p>The TSB bank disaster of 2018 &#8212; where a botched IT migration left 2 million customers locked out &#8212; is frequently cited by CAB defenders. But the independent review revealed that TSB had governance structures; it <em>failed to follow them</em>. The root causes were inadequate testing, unrealistic timelines, and testing in production rather than pre-production. <strong>This is evidence for better engineering practices, not for more committees.</strong></p><div><hr></div><h2><strong>Conclusions</strong></h2><p>The evidence is not ambiguous. Across the largest study of software delivery performance ever conducted, across a regulator&#8217;s analysis of over 1 million production changes, across case studies from ING to Amazon to Barclays &#8212; the conclusion is consistent:</p><p><strong>Change Advisory Boards do not do what they are supposed to do.</strong> They do not reduce change failure rates. They slow delivery. They force larger, riskier deployments. They erode team ownership. They create learned helplessness. They are compliance theater &#8212; providing the appearance of risk management without the substance.</p><p>The alternative is not a free-for-all. It is a rigorous, evidence-based approach to change management:</p><ul><li><p><strong>Peer review</strong> by engineers who actually understand the code</p></li><li><p><strong>Test-Driven Development</strong> that validates every change at the atomic level</p></li><li><p><strong>Automated deployment pipelines</strong> that provide consistent, objective, comprehensive quality gates</p></li><li><p><strong>Feature flags and progressive delivery</strong> that separate deployment from release</p></li><li><p><strong>Observability and fast rollback</strong> that make production behavior visible and recoverable</p></li><li><p><strong>Risk-based differentiation</strong> that applies proportionate scrutiny &#8212; lightweight for routine changes, focused expert review for the rare high-risk change</p></li><li><p><strong>Team ownership</strong> &#8212; &#8220;you build it, you run it&#8221; &#8212; because the people closest to the code are the people best equipped to assess its risk</p></li></ul><p>Even ITIL 4 &#8212; the framework that created the CAB &#8212; now supports this approach. It renamed &#8220;Change Management&#8221; to &#8220;Change Enablement,&#8221; made CABs optional, and aligned with Agile and DevOps practices.</p><p><strong>The best change advisory board is not a room full of people skimming change request forms once a week.</strong> It is a deployment pipeline powered by automated tests written before the code, running thousands of checks in minutes, providing feedback a hundred times faster than any committee ever could &#8212; and backed by the professional discipline of engineers who own their code from commit to production.</p><p>The data has spoken. The question is whether your organization will listen &#8212; or whether the CAB will continue to do what CABs do best: resist change.</p><div><hr></div><h2><strong>References</strong></h2><h3><strong>Books</strong></h3><ol><li><p>Forsgren, N., Humble, J., Kim, G. (2018). <em>Accelerate: The Science of Lean Software and DevOps.</em> IT Revolution Press.</p></li><li><p>Humble, J., Farley, D. (2010). <em>Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation.</em> Addison-Wesley.</p></li><li><p>Kim, G., Humble, J., Debois, P., Willis, J., Forsgren, N. (2021). <em>The DevOps Handbook, 2nd Edition.</em> IT Revolution Press.</p></li><li><p>Kim, G., Behr, K., Spafford, G. (2013). <em>The Phoenix Project.</em> IT Revolution Press.</p></li><li><p>Skelton, M., Pais, M. (2019). <em>Team Topologies.</em> IT Revolution Press.</p></li><li><p>Schwartz, M. (2020). <em>The (Delicate) Art of Bureaucracy.</em> IT Revolution Press.</p></li><li><p>Smart, J. (2020). <em>Sooner Safer Happier.</em> IT Revolution Press.</p></li><li><p>Majors, C., Fong-Jones, L., Miranda, G. (2022). <em>Observability Engineering.</em> O&#8217;Reilly.</p></li><li><p>Beck, K. (2003). <em>Test-Driven Development: By Example.</em> Addison-Wesley.</p></li><li><p>Farley, D. (2022). <em>Modern Software Engineering.</em> Addison-Wesley.</p></li><li><p>Graeber, D. (2018). <em>Bullshit Jobs: A Theory.</em> Simon &amp; Schuster.</p></li></ol><h3><strong>Research Reports</strong></h3><ol start="12"><li><p>DORA/Puppet. State of DevOps Reports (2014-2019). Multiple years of research on external approval and software delivery performance.</p></li><li><p>DORA/Google Cloud. (2019). <em>Accelerate State of DevOps Report.</em> Heavyweight approval 2.6x more likely in low performers.</p></li><li><p>DORA/Google Cloud. (2024). <em>Accelerate State of DevOps Report.</em> Continued emphasis on streamlining change approval.</p></li><li><p>Financial Conduct Authority. (2021). <em>Implementing Technology Change.</em> Multi-firm review. <a href="https://www.fca.org.uk/publications/multi-firm-reviews/implementing-technology-change">https://www.fca.org.uk/publications/multi-firm-reviews/implementing-technology-change</a></p></li><li><p>Nagappan, N., Maximilien, E.M., Bhat, T., Williams, L. (2008). &#8220;Realizing quality improvement through test driven development.&#8221; <em>Empirical Software Engineering.</em> <a href="https://www.microsoft.com/en-us/research/wp-content/uploads/2009/10/Realizing-Quality-Improvement-Through-Test-Driven-Development-Results-and-Experiences-of-Four-Industrial-Teams-nagappan_tdd.pdf">https://www.microsoft.com/en-us/research/wp-content/uploads/2009/10/Realizing-Quality-Improvement-Through-Test-Driven-Development-Results-and-Experiences-of-Four-Industrial-Teams-nagappan_tdd.pdf</a></p></li><li><p>IEEE. (2016). &#8220;Continuous Deployment at Facebook and OANDA.&#8221; <em>FSE 2016.</em> <a href="https://ieeexplore.ieee.org/document/7883285/">https://ieeexplore.ieee.org/document/7883285/</a></p></li></ol><h3><strong>Online Resources</strong></h3><ol start="18"><li><p>DORA. &#8220;Streamlining Change Approval.&#8221; <a href="https://dora.dev/capabilities/streamlining-change-approval/">https://dora.dev/capabilities/streamlining-change-approval/</a></p></li><li><p>DORA. &#8220;Test Automation.&#8221; <a href="https://dora.dev/capabilities/test-automation/">https://dora.dev/capabilities/test-automation/</a></p></li><li><p>Octopus Deploy. &#8220;Change Advisory Boards Don&#8217;t Work.&#8221; <a href="https://octopus.com/blog/change-advisory-boards-dont-work">https://octopus.com/blog/change-advisory-boards-dont-work</a></p></li><li><p>ThinkingLabs. &#8220;Jez Humble: Continuous Delivery Sounds Great But It Won&#8217;t Work Here.&#8221; <a href="https://thinkinglabs.io/notes/2021/12/11/agiletd-continuous-delivery-sounds-but-it-wont-work-here-jez-humble.html">https://thinkinglabs.io/notes/2021/12/11/agiletd-continuous-delivery-sounds-but-it-wont-work-here-jez-humble.html</a></p></li><li><p>Latham &amp; Watkins/JD Supra. &#8220;Implementing Technology Change &#8212; Successes and Pitfalls.&#8221; <a href="https://www.jdsupra.com/legalnews/implementing-technology-change-2874927/">https://www.jdsupra.com/legalnews/implementing-technology-change-2874927/</a></p></li><li><p>Rance, S. &#8220;IT Change Management Doesn&#8217;t Always Need a CAB.&#8221; Optimal Service Management. <a href="https://www.optimalservicemanagement.com/blog/it-change-management-doesnt-always-need-a-cab/">https://www.optimalservicemanagement.com/blog/it-change-management-doesnt-always-need-a-cab/</a></p></li><li><p>InvGate. &#8220;Change Advisory Board Best Practices: 15+ Industry Leaders Weigh In.&#8221; <a href="https://blog.invgate.com/do-we-still-need-the-change-advisory-board">https://blog.invgate.com/do-we-still-need-the-change-advisory-board</a></p></li><li><p>Martin, R.C. &#8220;Professionalism and TDD.&#8221; Clean Coder Blog. <a href="https://blog.cleancoder.com/uncle-bob/2014/05/02/ProfessionalismAndTDD.html">https://blog.cleancoder.com/uncle-bob/2014/05/02/ProfessionalismAndTDD.html</a></p></li><li><p>Farley, D. &#8220;Continuous Compliance.&#8221; </p></li></ol><p>https://www.davefarley.net/?p=285</p><ol start="18"><li><p>Farley, D. &#8220;Test Driven Development.&#8221; </p></li></ol><p>https://www.davefarley.net/?p=220</p><ol start="18"><li><p>Fowler, M. &#8220;Feature Toggles (aka Feature Flags).&#8221; <a href="https://martinfowler.com/articles/feature-toggles.html">https://martinfowler.com/articles/feature-toggles.html</a></p></li><li><p>Fowler, M. &#8220;Deployment Pipeline.&#8221; <a href="https://martinfowler.com/bliki/DeploymentPipeline.html">https://martinfowler.com/bliki/DeploymentPipeline.html</a></p></li><li><p>Humble, J. &#8220;Continuous Delivery and ITIL Change Management.&#8221; <a href="https://continuousdelivery.com/2010/11/continuous-delivery-and-itil-change-management/">https://continuousdelivery.com/2010/11/continuous-delivery-and-itil-change-management/</a></p></li></ol><h3><strong>ITIL 4 Sources</strong></h3><ol start="31"><li><p>ITSM.tools. &#8220;Change Enablement in ITIL 4.&#8221; <a href="https://itsm.tools/change-enablement/">https://itsm.tools/change-enablement/</a></p></li><li><p>Joe The IT Guy. &#8220;What&#8217;s Changed with Change in ITIL 4?&#8221; <a href="https://www.joetheitguy.com/whats-changed-with-change-in-itil-4/">https://www.joetheitguy.com/whats-changed-with-change-in-itil-4/</a></p></li><li><p>Beyond20. &#8220;Understanding Change Enablement Practice in ITIL 4.&#8221; <a href="https://www.beyond20.com/blog/understanding-change-enablement-practice-itil-4/">https://www.beyond20.com/blog/understanding-change-enablement-practice-itil-4/</a></p></li></ol><h3><strong>Case Studies</strong></h3><ol start="34"><li><p>InfoQ. &#8220;ING Netherlands&#8217; Transition to DevOps.&#8221; <a href="https://www.infoq.com/news/2014/06/ing-transtition-to-devops/">https://www.infoq.com/news/2014/06/ing-transtition-to-devops/</a></p></li><li><p>The Enterprisers Project. &#8220;Target CIO Explains How DevOps Took Root.&#8221; <a href="https://enterprisersproject.com/article/2017/1/target-cio-explains-how-devops-took-root-inside-retail-giant">https://enterprisersproject.com/article/2017/1/target-cio-explains-how-devops-took-root-inside-retail-giant</a></p></li><li><p>InfoQ. &#8220;Benefits of Agile Transformation at Barclays.&#8221; <a href="https://www.infoq.com/news/2016/09/benefits-agile-barclays/">https://www.infoq.com/news/2016/09/benefits-agile-barclays/</a></p></li><li><p>Barclays. &#8220;Insights: Jonathan Smart.&#8221; <a href="https://home.barclays/news/2018/02/insights-jonathan-smart/">https://home.barclays/news/2018/02/insights-jonathan-smart/</a></p></li><li><p>Meta Engineering. &#8220;Rapid Release at Massive Scale.&#8221; <a href="https://engineering.fb.com/2017/08/31/web/rapid-release-at-massive-scale/">https://engineering.fb.com/2017/08/31/web/rapid-release-at-massive-scale/</a></p></li><li><p>Panorama Consulting. &#8220;4 Lessons from the TSB Software Failure.&#8221; <a href="https://panorama-consulting.com/tsb-software-failure/">https://panorama-consulting.com/tsb-software-failure/</a></p></li><li><p>Bank of England. &#8220;TSB fined &#163;48.65m for operational resilience failings.&#8221; <a href="https://www.bankofengland.co.uk/news/2022/december/tsb-fined-for-operational-resilience-failings">https://www.bankofengland.co.uk/news/2022/december/tsb-fined-for-operational-resilience-failings</a></p></li></ol><h3><strong>Psychological Frameworks</strong></h3><ol start="41"><li><p>Petriglieri, J.L. (2011). &#8220;Under Threat: Responses to and the Consequences of Threats to Individuals&#8217; Identities.&#8221; <em>Academy of Management Review.</em></p></li><li><p>Kahneman, D., Tversky, A. (1979). &#8220;Prospect Theory: An Analysis of Decision under Risk.&#8221; <em>Econometrica.</em></p></li><li><p>Seligman, M.E.P. (1972). <em>Learned Helplessness.</em> Annual Review of Medicine.</p></li><li><p>McGregor, D. (1960). <em>The Human Side of Enterprise.</em> McGraw-Hill.</p></li><li><p>Edmondson, A.C. (1999). &#8220;Psychological Safety and Learning Behavior in Work Teams.&#8221; <em>Administrative Science Quarterly.</em></p></li><li><p>French, J.R.P., Raven, B. (1959). &#8220;The Bases of Social Power.&#8221; In Cartwright, D. (Ed.), <em>Studies in Social Power.</em></p></li><li><p>Westrum, R. (1988). &#8220;Organizational and Inter-Organizational Thought.&#8221; World Bank Conference on Safety Control and Risk Management.</p></li><li><p>Corporate Rebels. &#8220;Learned Helplessness at Work.&#8221; <a href="https://www.corporate-rebels.com/blog/learned-helplessness-at-work">https://www.corporate-rebels.com/blog/learned-helplessness-at-work</a></p></li></ol>]]></content:encoded></item><item><title><![CDATA[Expectation-Driven Development: A Validation Framework for the Age of AI Agents]]></title><description><![CDATA[The Problem Nobody Wants to Talk About]]></description><link>https://a4al6a.substack.com/p/expectation-driven-development-a</link><guid isPermaLink="false">https://a4al6a.substack.com/p/expectation-driven-development-a</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Sat, 21 Mar 2026 00:30:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><strong>The Problem Nobody Wants to Talk About</strong></h2><p>Here&#8217;s an uncomfortable truth about working with AI coding agents: <strong>we can&#8217;t review their output.</strong></p><p>Not &#8220;we choose not to.&#8221; We <em>can&#8217;t</em>. Not meaningfully.</p><p>An AI agent can produce hundreds of lines of working code in seconds. A human reviewing that code line-by-line needs minutes, sometimes hours. The economics don&#8217;t work. And it&#8217;s getting worse &#8212; agents are getting faster, and the code they produce is getting more complex.</p><p>So what do teams actually do? They glance at the diff, run the existing tests, maybe squint at the architecture, and merge. We&#8217;ve replaced &#8220;trust but verify&#8221; with &#8220;trust and hope the CI is green.&#8221;</p><p>This isn&#8217;t sustainable. But I don&#8217;t think the answer is &#8220;just write more tests&#8221; either.</p><h2><strong>Where TDD and BDD Leave a Gap</strong></h2><p>Let me be clear: I&#8217;m not here to bury Test-Driven Development. TDD is one of the most important ideas in software engineering, and it works well with AI agents as well.</p><p>But TDD has a scope problem. It excels at specifying precise, deterministic behaviors &#8212; this input produces that output. It&#8217;s less natural for expressing the kind of requirements that live in the gaps: the relationship between multiple behaviors, the qualitative expectations (&#8221;the error message should be helpful, not cryptic&#8221;), the systemic properties (&#8221;this must hold even under high load&#8221;). These requirements are real. They matter to users. And they tend to live in the developer&#8217;s head, unwritten, because there&#8217;s no natural place to put them in a test file.</p><p>There&#8217;s also a phase problem. Before you can write a test, you need to know <em>what</em> to test. The design-time exploration &#8212; &#8220;what should this feature actually do, in all its edge cases?&#8221; &#8212; happens before TDD begins. It usually happens informally: in Slack, in a meeting, in someone&#8217;s head. It&#8217;s rarely captured.</p><p>BDD (Behavior-Driven Development) gets closer. Its Given/When/Then syntax forces you to think about behavior from the outside in, and its natural-language layer bridges the gap between business intent and executable code. But BDD&#8217;s formalism &#8212; specific frameworks (Cucumber, SpecFlow, Behave), step definitions, glue code &#8212; is also its strength: it&#8217;s what makes BDD scenarios executable and rerunnable. That rigor has a cost, though. It constrains how you express requirements. Some expectations are hard to fit into a three-line template.</p><p>What if we could keep the <em>intent</em> of BDD &#8212; specify behavior, then verify it &#8212; while trading some of its formalism for the flexibility to capture the full picture?</p><h2><strong>Enter Expectation-Driven Development</strong></h2><p>The idea is simple, maybe deceptively so:</p><p><strong>1. Formulate expectations in plain text.</strong></p><p>Not in a formal Given/When/Then template, though you can use that structure if it helps. Not in a programming language. In the same natural language you&#8217;d use to explain the feature to a colleague.</p><p>For example:</p><blockquote><p><strong>Expectation: Cart total calculation</strong></p><p>When a user adds multiple items to their cart, the total should reflect the sum of all item prices multiplied by their quantities. If a discount code is applied, the discount should be calculated on the pre-tax subtotal, not on individual items. Tax is applied after the discount. If the cart is empty, the total should be zero, not an error.</p></blockquote><p>Notice what this captures that a unit test wouldn&#8217;t: the <em>relationship</em> between discount and tax ordering, the empty cart edge case framed as an expectation about behavior (zero, not error), and the implicit requirement that this should work with &#8220;multiple items&#8221; &#8212; not just the two items in your test fixture.</p><p>You can go further:</p><blockquote><p><strong>Expectation: Race condition in concurrent booking</strong></p><p>If two users attempt to book the last available slot at the same time, exactly one should succeed and receive a confirmation. The other should receive a clear rejection &#8212; not a timeout, not a double booking, not a corrupted state. This must hold even under high load.</p></blockquote><p>Try writing that as a unit test. You&#8217;ll end up with a page of setup code that obscures the actual expectation. To be fair, nobody writes concurrency verification as a unit test &#8212; this would be an integration or stress test. But that&#8217;s the point. EDD&#8217;s job here isn&#8217;t to replace the concurrency test. It&#8217;s to make sure someone writes one, and that it tests the right thing. The expectation captures the requirement precisely; the test implements the verification.</p><p><strong>2. Hand the expectations to the AI agent.</strong></p><p>The agent implements the feature. It has the full context of your codebase, your expectations, and its own understanding of best practices.</p><p><strong>3. Ask the agent to prove it.</strong></p><p>This is the key step. Don&#8217;t just ask &#8220;did you implement it?&#8221; Ask: <strong>&#8220;Prove to me that each expectation is fulfilled. Show me specific scenarios with concrete inputs and the outputs the system produces.&#8221;</strong></p><p>The agent should respond with evidence like:</p><blockquote><p><strong>Evidence for: Cart total calculation</strong></p><p><em>Scenario 1</em>: Added &#8220;Widget&#8221; ($10.00, qty 2) and &#8220;Gadget&#8221; ($25.00, qty 1). Subtotal: $45.00. Applied discount code &#8220;SAVE10&#8221; (10% off). Discount: -$4.50. After discount: $40.50. Tax (8%): $3.24. <strong>Total: $43.74.</strong></p><p><em>Scenario 2</em>: Empty cart. GET /api/cart/total returns <code>{ "total": 0.00, "items": [] }</code> with HTTP 200. <strong>No error thrown.</strong></p><p><em>Scenario 3</em>: Added &#8220;Widget&#8221; ($10.00, qty 1). No discount. Tax (8%): $0.80. <strong>Total: $10.80.</strong> Tax applied to full subtotal, not discounted amount (no discount to apply).</p></blockquote><p><strong>4. Evaluate, challenge, iterate.</strong></p><p>Read the evidence critically. Does it actually cover the expectation? Did the agent dodge the hard part? Are the numbers right? Push back:</p><blockquote><p>&#8220;You showed the discount applied to the subtotal, but you didn&#8217;t show what happens if someone tries to apply two discount codes. What happens then?&#8221;</p></blockquote><p>The agent revises. New evidence is produced. You iterate until convergence.</p><p><strong>5. The evidence becomes the documentation.</strong></p><p>When you&#8217;re done, you have something valuable: a set of expectations paired with concrete proof that the system meets them. This isn&#8217;t a test suite that requires a framework to run. It&#8217;s not a wiki page that was outdated the day it was written. It&#8217;s a living record of what the system does and why, backed by specific examples.</p><h2><strong>The Workflow, Visualized</strong></h2><pre><code><code>Human: Formulates expectations (plain text)
  &#9474;
  &#9660;
AI Agent: Implements the feature
  &#9474;
  &#9660;
Human: "Prove it meets the expectations"
  &#9474;
  &#9660;
AI Agent: Produces evidence (concrete scenarios, inputs, outputs)
  &#9474;
  &#9660;
Human: Reviews evidence &#9472;&#9472;&#9472;&#9472; Satisfied? &#9472;&#9472;&#9472;&#9472; YES &#9472;&#9472;&#8594; Document &amp; ship
  &#9474;
  NO
  &#9474;
  &#9660;
Human: Challenges, adds expectations
  &#9474;
  &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472; Loop back to AI Agent &#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
</code></code></pre><h2><strong>This Isn&#8217;t a New Idea (And That&#8217;s a Good Thing)</strong></h2><p>If you&#8217;ve been in software long enough, you&#8217;re probably thinking: &#8220;This sounds like Specification by Example.&#8221; You&#8217;re right. It does.</p><p>EDD stands on the shoulders of a long lineage of ideas that share the same core insight &#8212; that concrete examples in human-readable form are the best way to specify and verify software behavior:</p><ul><li><p><strong>Specification by Example</strong> (Gojko Adzic, 2011) &#8212; uses concrete examples collaboratively authored by the team as both specifications and tests. The bible for this way of thinking.</p></li><li><p><strong>FIT / FitNesse</strong> (Ward Cunningham, ~2002) &#8212; executable acceptance tables written in near-plain-English, verified automatically against the system.</p></li><li><p><strong>Concordion</strong> &#8212; specifications written in natural language that become executable through instrumentation.</p></li><li><p><strong>Design by Contract</strong> (Bertrand Meyer, Eiffel) &#8212; preconditions, postconditions, and invariants as formal specifications embedded in code.</p></li><li><p><strong>Property-based testing</strong> (QuickCheck and descendants) &#8212; addresses the &#8220;not just the two items in your test fixture&#8221; problem by generating many inputs from declared properties.</p></li></ul><p>So what&#8217;s actually new here? Not the idea of specifying behavior in natural language. Not the idea of verifying with concrete examples. What&#8217;s new is the <em>execution context</em>.</p><p>In Specification by Example, a human team collaborates to write examples, then a developer writes glue code (step definitions) to make them executable. The bottleneck is the glue code &#8212; it&#8217;s tedious, it breaks when the code changes, and it requires a framework.</p><p>In EDD, the LLM <em>is</em> the glue code. It interprets natural-language expectations directly, without step definitions, without a framework, without the ceremony. And the verification loop is conversational &#8212; you challenge, the agent responds, you push back &#8212; rather than pass/fail binary.</p><p>That&#8217;s a meaningful difference. But I want to be honest that it&#8217;s an evolutionary one, not a revolutionary one. If you&#8217;ve read Adzic&#8217;s work, you&#8217;ll feel at home here. If you haven&#8217;t, go read it &#8212; it will make you better at EDD.</p><p>There&#8217;s one more thing EDD inherits from Specification by Example: the question of <em>who writes the expectations</em>. In the Specification by Example tradition, examples are collaboratively authored in &#8220;specification workshops&#8221; &#8212; conversations between developers, testers, and business stakeholders. EDD as I&#8217;ve described it is a solo workflow: one human, one agent. But most software is built by teams. If different team members write expectations for the same feature, they may express different &#8212; or contradictory &#8212; assumptions. If only one person writes them, you&#8217;ve concentrated a single point of failure in that person&#8217;s understanding. I don&#8217;t have a neat answer for this yet. My instinct is that expectations should be written collaboratively, then the conversation with the agent happens individually. But this is an open question.</p><h2><strong>Why This Might Actually Work</strong></h2><p><strong>It plays to human strengths.</strong> Humans are better at judging than producing. We&#8217;re excellent critics and mediocre typists. EDD lets us do what we do best: specify intent, evaluate outcomes, spot what&#8217;s missing.</p><p><strong>It plays to AI strengths.</strong> AI agents are fast, tireless implementers that can generate both code and evidence at scale. The bottleneck was never &#8220;can the AI write the code?&#8221; It was &#8220;can we trust the code the AI wrote?&#8221; EDD creates a structured trust-building process.</p><p><strong>Natural language captures what formal tests can&#8217;t.</strong> &#8220;The error message should be helpful, not cryptic.&#8221; &#8220;The response time should feel instant for typical queries.&#8221; &#8220;The fallback behavior should be graceful, not surprising.&#8221; These are real requirements that matter to users. They&#8217;re nearly impossible to express in <code>assert</code> statements, but an LLM can interpret them. A caveat: for subjective qualities like &#8220;helpful&#8221; or &#8220;graceful,&#8221; the LLM will tend to judge its own output favorably &#8212; it&#8217;s the fox-guarding-the-henhouse problem again. For these expectations, the human&#8217;s judgment in Step 4 becomes especially critical. Don&#8217;t outsource taste.</p><p><strong>It forces you to think before coding &#8212; the best part of TDD, without the ceremony.</strong> The discipline of TDD was never really about the tests. It was about the act of specifying behavior before implementation. EDD preserves that discipline. You still think first. You just express your thinking in a more natural medium.</p><p><strong>The evidence trail creates accountability.</strong> Every expectation has a proof artifact. Six months from now, when someone asks &#8220;does the system handle concurrent bookings correctly?&#8221;, you don&#8217;t grep through test files. You read the expectation and its evidence.</p><h2><strong>Why This Might Not Work (The Honest Part)</strong></h2><p>I believe in this idea, but I&#8217;d be dishonest if I didn&#8217;t confront its weaknesses head-on. There are real problems here.</p><h3><strong>The Fox Guarding the Henhouse (And the Deeper Problem Beneath It)</strong></h3><p>This is the big one, and it has two layers.</p><p><strong>Layer 1: Bias.</strong> You&#8217;re asking the same AI that wrote the code to produce evidence that the code works. That&#8217;s like asking a student to both take the exam and grade it. The AI has every structural incentive to produce evidence that confirms its implementation. If it made a subtle error in the discount calculation, it might generate scenarios that avoid the precise inputs that would reveal the bug. Not maliciously. Just because the same reasoning flaw that caused the bug will also cause it to overlook the bug in its evidence.</p><p><strong>Layer 2: Execution vs. narration.</strong> This is the deeper problem, and it&#8217;s one we need to confront directly. When the agent produces &#8220;evidence,&#8221; what actually happened? There are two very different possibilities:</p><ul><li><p><strong>Executed evidence</strong>: The agent actually ran the code &#8212; made an API call, executed a function, queried the database &#8212; and is showing you real output from a real system.</p></li><li><p><strong>Generative evidence</strong>: The agent <em>described</em> what it believes would happen, based on its understanding of the code it wrote. It narrated a plausible verification without executing anything.</p></li></ul><p>These are not the same thing. Executed evidence is evidence. Generative evidence is a second assertion by the same entity that made the first assertion. It&#8217;s the difference between &#8220;I tested it and here are the results&#8221; and &#8220;I&#8217;m pretty sure it would work like this.&#8221;</p><p><strong>EDD requires executed evidence.</strong> If the agent can&#8217;t run the code and show you real outputs, you don&#8217;t have verification &#8212; you have a shared hallucination dressed up in scenario format. This means EDD depends on AI agents with tool use: the ability to execute code, call APIs, run scripts, and capture actual output. Fortunately, this is where agents are headed &#8212; modern coding agents already have shell access, can run test suites, and can interact with running systems. But you must insist on it. When reviewing evidence, ask: &#8220;Did you actually run this, or are you telling me what you think would happen?&#8221; If the agent can&#8217;t answer that clearly, the evidence is worthless.</p><p><strong>But let&#8217;s be honest: not all code is executable in an agent loop.</strong> In practice, evidence falls into three categories:</p><ul><li><p><strong>Directly executable</strong>: Functions, APIs, scripts, database queries. The agent can run these and show real output. This is the gold standard.</p></li><li><p><strong>Partially verifiable</strong>: Infrastructure code (Terraform plan output), build configurations (dry runs), schema migrations (against a test database). The agent can&#8217;t deploy to production, but it can show you what <em>would</em> happen. This is weaker but still useful &#8212; a Terraform plan is better than a guess.</p></li><li><p><strong>Not executable in the loop</strong>: UI rendering, production-only behavior, third-party API integrations with rate limits or authentication, code that requires manual user interaction. Here, you&#8217;re back to generative evidence whether you like it or not.</p></li></ul><p>For that third category, the mitigations below become critical, and you should be clear-eyed that your confidence is lower. If most of your expectations fall into category three, EDD&#8217;s value proposition degrades &#8212; and you should invest more in the &#8220;Stabilize&#8221; step (converting expectations to automated tests that <em>can</em> run in CI).</p><p><strong>Mitigations for both layers:</strong></p><ul><li><p><strong>Be an adversarial reviewer.</strong> Your job isn&#8217;t to passively receive evidence. It&#8217;s to actively challenge it. Run adversarial scenarios. Ask &#8220;what about...?&#8221; questions. Spot-check the numbers manually.</p></li><li><p><strong>Demand execution receipts.</strong> Ask the agent to show the actual command it ran and the raw output. Not a summary. The output.</p></li><li><p><strong>Use a different agent to audit.</strong> Have a second AI agent (or a different model) independently verify the evidence against the code. The fox can guard the henhouse if there&#8217;s a different fox checking the first fox&#8217;s work.</p></li><li><p><strong>Spot-check yourself.</strong> For critical expectations, pick one scenario and run it manually. If the agent&#8217;s evidence matches reality for the one you checked, you have higher confidence in the rest.</p></li></ul><h3><strong>The Reproducibility Question</strong></h3><p>Here&#8217;s a hard question: are the expectations rerunnable?</p><p>If I come back next week, after the code has changed, can I re-verify the expectations? If yes &#8212; how? If you&#8217;re asking an AI to re-run the evidence each time, you&#8217;re relying on LLM interpretation, which is non-deterministic. The same expectation might produce different evidence on different runs. The same model might interpret an ambiguous expectation differently after an update.</p><p>If the expectations <em>aren&#8217;t</em> rerunnable, then what you have is a snapshot, not a safety net. Documentation, not regression protection.</p><p><strong>Mitigation:</strong> EDD should complement automated tests, not replace them. The expectations drive the initial implementation and verification. But the critical paths should <em>also</em> be captured in traditional automated tests for regression. Think of expectations as the <em>design-time</em> validation tool, and automated tests as the <em>runtime</em> safety net. The expectations are the &#8220;why.&#8221; The tests are the &#8220;what, forever.&#8221;</p><p>There&#8217;s a related versioning problem worth flagging. If you ship feature v1 with its evidence, then modify the feature in v2, the v1 evidence is now potentially misleading. Do you re-run evidence for every change? If so, you&#8217;re paying the EDD cost on every iteration. If not, the documentation rots like any other documentation. Automated tests don&#8217;t have this problem &#8212; they fail loudly when the code changes in ways they don&#8217;t expect. EDD evidence is silent when it becomes stale. This is another reason the &#8220;Stabilize&#8221; step matters: the automated tests are your regression alarm. The expectations are your design-time conversation.</p><h3><strong>Natural Language Ambiguity</strong></h3><p>I praised the freedom of natural language earlier. But freedom has a cost. Consider:</p><blockquote><p>&#8220;The system should handle large uploads efficiently.&#8221;</p></blockquote><p>What&#8217;s large? 10MB? 10GB? What&#8217;s efficiently? Under 5 seconds? Without running out of memory? Without blocking other requests?</p><p>In BDD, the formalism forces you to be specific: <code>Given a file of 500MB / When the user uploads it / Then the upload completes within 30 seconds</code>. The rigidity is a feature. It prevents hand-waving.</p><p><strong>Mitigation:</strong> Write expectations as specifically as you can. The freedom of natural language doesn&#8217;t mean you should be vague &#8212; it means you can be specific in ways that formal syntax doesn&#8217;t allow. Instead of &#8220;handle large uploads efficiently,&#8221; write &#8220;a 500MB upload should complete without timeout and without consuming more than 2x the file size in memory. A 5GB upload should use streaming and never hold the full file in memory.&#8221; Still natural language. But precise.</p><h3><strong>The Evidence Scalability Problem</strong></h3><p>If you have 5 expectations, you can carefully review 5 sets of evidence. If you have 50, you&#8217;ll start skimming. If you have 200, you&#8217;ll rubber-stamp.</p><p>And yet, more expectations means better coverage. There&#8217;s a tension between thoroughness and human attention span.</p><p><strong>Mitigation:</strong> Prioritize. Not all expectations are equal. Some protect critical business logic. Some cover edge cases that matter but don&#8217;t need forensic review. Categorize expectations by risk and allocate your review attention accordingly.</p><h2><strong>EDD vs. TDD vs. BDD: Not a Replacement, a Complement</strong></h2><p>Let me be blunt: if you read this and think &#8220;great, I can stop writing tests,&#8221; you&#8217;ve missed the point.</p><p><strong>TDD</strong> speaks to developers, in code. It&#8217;s always executable, gives you strong regression protection, and works well as a design-time thinking tool. But it captures limited nuance &#8212; an <code>assert</code> can only say so much &#8212; and it wasn&#8217;t designed with AI agents in mind (though &#8220;make these tests pass&#8221; works surprisingly well).</p><p><strong>BDD</strong> speaks to the whole team, in structured natural language. It&#8217;s executable via step definitions, gives you strong regression protection, and captures more nuance than raw code. But it still constrains how you express requirements, and it wasn&#8217;t designed for AI agents either.</p><p><strong>EDD</strong> speaks to the human-AI pair, in free-form natural language. The expectations themselves aren&#8217;t executable &#8212; they&#8217;re text &#8212; but the protocol demands execution in Step 3. Regression protection is weak without automation (which is why the Stabilize step matters). Where EDD wins is nuance: natural language can express the qualitative, relational, and systemic requirements that formal test syntax struggles with. And it&#8217;s designed from the ground up for the AI-agent workflow.</p><p>The three aren&#8217;t competitors. They&#8217;re layers. EDD is the <em>conversation layer</em> between human intent and AI implementation &#8212; where you say what you mean, in all its messy, nuanced, edge-case-laden glory. Then TDD or BDD takes over for the parts that need to be deterministic and repeatable.</p><h2><strong>A Practical Protocol</strong></h2><p>If you want to try EDD today, here&#8217;s a concrete protocol:</p><h3><strong>Step 1: Write expectations before touching code</strong></h3><p>Spend 15 minutes writing expectations for the feature. Be specific. Cover:</p><ul><li><p>The happy path</p></li><li><p>Edge cases you&#8217;ve been burned by before</p></li><li><p>Non-functional requirements (performance, error handling, security)</p></li><li><p>Behaviors that should explicitly <em>not</em> happen</p></li></ul><p><strong>How big should an expectation be?</strong> Think &#8220;one behavior you&#8217;d explain in a single breath to a colleague.&#8221; The cart total calculation example is about right &#8212; it covers one coherent concern (pricing math) with its key edge cases. If you find yourself writing a page-long expectation, you&#8217;re probably bundling multiple concerns. Split them. If you find yourself writing a one-liner with no edge cases, you&#8217;re probably being too vague. Expand it.</p><h3><strong>Step 2: Hand expectations to the AI agent with your codebase</strong></h3><p>Give the agent your expectations and let it implement. Don&#8217;t hover. Let it work.</p><h3><strong>Step 3: Request executed evidence</strong></h3><p>Prompt: <em>&#8220;For each expectation, show me concrete evidence that the system fulfills it. Actually run the code &#8212; execute the function, call the API, run the script &#8212; and show me the real inputs and outputs. Don&#8217;t describe what you think would happen; show me what actually happened. If an expectation cannot be fulfilled, explain why.&#8221;</em></p><p>The word &#8220;actually&#8221; is doing heavy lifting here. You want execution receipts, not narration.</p><h3><strong>Step 4: Review adversarially</strong></h3><p>For each piece of evidence:</p><ul><li><p>Do the numbers add up?</p></li><li><p>Did the agent test the edge case or dodge it?</p></li><li><p>What input would break this?</p></li><li><p>Is there a scenario the expectation didn&#8217;t cover that it should?</p></li></ul><h3><strong>Step 5: Challenge and iterate</strong></h3><p>Add new expectations based on what you discover. Tighten vague ones. Ask &#8220;what if?&#8221; until you&#8217;re satisfied.</p><h3><strong>Step 6: Stabilize</strong></h3><p>For critical paths, convert the expectations and evidence into automated tests (unit, integration, or e2e). The evidence gives you the test cases for free &#8212; you just need to make them executable and deterministic.</p><h3><strong>Step 7: Archive</strong></h3><p>The final set of expectations + evidence becomes your feature documentation. It answers &#8220;what does this do?&#8221; and &#8220;how do we know?&#8221; in one artifact.</p><h2><strong>A Disclaimer, and the Deeper Point</strong></h2><p>I should be upfront: I haven&#8217;t battle-tested EDD on a real project. This is a proposed framework, not a field report. I&#8217;m sharing it because I think the underlying problem &#8212; how do we validate AI-generated code at the speed AI generates it? &#8212; is urgent enough that we should be thinking out loud about solutions, even imperfect ones. If you try this and it works, I want to hear about it. If you try it and it falls apart, I want to hear about that even more.</p><p>That said, I think EDD points at something deeper than a specific technique. It reflects a shift in the role of the human developer.</p><p>In the pre-AI world, we were <em>authors</em>. We wrote the code. We wrote the tests. We wrote the documentation. Our value was in production.</p><p>In the AI-agent world, we&#8217;re becoming <em>editors</em>. We specify intent. We evaluate output. We challenge evidence. Our value is in judgment.</p><p>Authors need tools for writing &#8212; IDEs, compilers, debuggers. Editors need tools for <em>evaluating</em> &#8212; ways to express what they want, inspect what they got, and close the gap between the two. TDD is an author&#8217;s tool. EDD is an attempt at an editor&#8217;s tool.</p><p>We&#8217;re early. The tooling isn&#8217;t there yet. The methodology needs pressure-testing by real teams on real projects. But the direction feels right: humans specifying intent, machines producing implementations, and a structured conversation in between to build justified confidence that the two actually match.</p><div><hr></div><p><em>The question isn&#8217;t whether AI agents will write most of our code. They already do. The question is whether we&#8217;ll find a way to stay confident that the code does what we intended. Expectation-Driven Development is one bet on how we get there. I&#8217;m looking for others.</em></p>]]></content:encoded></item><item><title><![CDATA[Stop Using Pull Requests]]></title><description><![CDATA[Your team&#8217;s code review process is probably an expensive illusion of quality. Here&#8217;s what the evidence says, and what to do instead.]]></description><link>https://a4al6a.substack.com/p/stop-using-pull-requests</link><guid isPermaLink="false">https://a4al6a.substack.com/p/stop-using-pull-requests</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Thu, 19 Mar 2026 12:34:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><em>&#8220;Inspection is too late. The quality, good or bad, is already in the product. Cease dependence on inspection to achieve quality. Eliminate the need for inspection on a mass basis by building quality into the product in the first place.&#8221;</em></p><p>-- W. Edwards Deming, <em>Out of the Crisis</em> (1982)</p></blockquote><div><hr></div><h3><strong>Abstract</strong></h3><p>Pull requests have become the default code review mechanism in software teams everywhere. But the pull request was invented for open source projects, where strangers contribute code to repositories maintained by people who don&#8217;t know them. When private teams of trusted colleagues adopt the same process, they import a model designed for low trust into an environment that should operate on high trust. The result is a post-development inspection system that the evidence says catches few bugs, introduces enormous waiting time, incentivises large batches, and fractures teams into isolated individuals. Academic research, large-scale industry data from DORA, and a growing practitioner consensus all point in the same direction: there are better ways to build quality into software than inspecting it after the fact. This article examines the evidence and proposes an alternative &#8212; T*D, the union of Test-Driven Development, Trunk-Based Development, and Team-focused Development &#8212; as a path from slow, inspection-heavy workflows to fast-flowing teams that produce high-quality, safe systems.</p><div><hr></div><h3><strong>TL;DR</strong></h3><ul><li><p>Pull requests were designed for open source contributions from untrusted strangers. Applying them to trusted teams is a category error.</p></li><li><p>Peer-reviewed research shows code review&#8217;s primary value is knowledge transfer, not bug detection. Less than 15% of review comments relate to actual bugs.</p></li><li><p>Async PR workflows mean your code spends 86-99% of its lead time <em>waiting</em>. One organisation spent 130,000 hours in a single year waiting on PRs that received zero comments.</p></li><li><p>DORA research across 36,000+ professionals shows trunk-based development correlates with dramatically higher software delivery performance, and faster code reviews alone improve performance by 50%.</p></li><li><p>The alternative is T*D: Test-Driven Development (build quality in), Trunk-Based Development (integrate continuously), and Team-focused Development (review during creation, not after).</p></li><li><p>The transition is gradual: optimise PRs first, adopt Ship/Show/Ask, then move to pairing and trunk-based development as trust and automation mature.</p></li></ul><div><hr></div><p>Deming wrote those words about manufacturing in 1982, but they describe what happens in most software teams today with uncanny precision. A developer writes code in isolation, on a branch, for hours or days. Then they open a pull request. The code sits in a queue. Someone eventually looks at it, leaves a few comments -- mostly about naming and formatting -- clicks approve, and the code merges. The whole process feels thorough. It feels like quality control. But the evidence suggests it is mostly theatre.</p><p>This article is not against code review. <strong>Code review has real value -- but that value is knowledge transfer and shared understanding, not catching bugs.</strong> And a blocking, asynchronous pull request is one of the worst possible mechanisms for achieving it.</p><p>What follows is a synthesis of peer-reviewed academic research, large-scale industry data, and practitioner experience spanning two decades. The picture that emerges is consistent and, for many teams, uncomfortable: the way most organisations practice code review is an expensive ritual that slows delivery, encourages large batches, fractures team cohesion, and produces a false sense of security. Better alternatives exist, and they aren&#8217;t theoretical -- they are practised successfully by high-performing teams around the world.</p><div><hr></div><h2><strong>A Solution Designed for Strangers</strong></h2><p>Pull requests have a specific origin story, and understanding it explains why they are a poor fit for most teams.</p><p>Linus Torvalds created the <code>git-request-pull</code> command in 2005, shortly after releasing Git. Before that, open source projects used email-based patch workflows: contributors mailed diffs to a mailing list, and the maintainer decided whether to apply them. In 2008, GitHub launched the web-based pull request, making it dramatically easier to accept contributions from the outside world -- from people the maintainers did not know.</p><p>This was a genuine innovation for open source. <strong>The pull request was designed as a gatekeeper mechanism for untrusted contributors submitting code to repositories they don&#8217;t own.</strong> It solved a real problem: how do you safely accept code from strangers?</p><p>But then something happened. Git became the dominant version control system. GitHub became the dominant platform. And teams everywhere adopted the pull request as their default workflow -- not because they had evaluated it against alternatives, but because it was there and everyone else was using it. As Martin Fowler observed: <em>&#8220;I suspect that since pull requests are so popular, a lot of teams are using them by default when they would do better without them.&#8221;</em></p><p>The category error should be obvious. <strong>In an open source project, you are reviewing code from someone you may never have met, working in a codebase they may not fully understand, with no shared context about the team&#8217;s conventions or direction.</strong> A gatekeeper model makes sense. In a private team, you are reviewing code from a colleague who sits in the same stand-up, shares the same goals, and (presumably) has been hired precisely because you trust their competence. Yet the process is the same.</p><p>Thierry de Pauw puts it bluntly: <em>&#8220;Pull requests are designed to make it easier to accept contributions from the outside world, from untrusted people we do not know about.&#8221;</em> When your team adopts this model internally, you are importing a trust assumption that does not match your reality.</p><p>Even in open source, pull requests cause friction. Contributions sit unreviewed for weeks or months. Maintainer burnout is well documented. The process works <em>better</em> than emailing patches, but it is hardly frictionless. And the friction that is merely inconvenient in open source becomes genuinely destructive in a team that needs to integrate and deploy multiple times a day.</p><div><hr></div><h2><strong>What Code Review Actually Finds</strong></h2><p>If you ask developers why they do code review, the most common answer is: to catch bugs. The evidence says this is largely a myth.</p><p>The most significant data comes from Microsoft Research. A 2015 IEEE study titled <em>&#8220;Code Reviews Do Not Find Bugs: How the Current Code Review Best Practice Slows Us Down&#8221;</em> found that only a very small percentage of code review comments had anything to do with bugs. Most were about structural issues and style. A landmark 2013 paper by Bacchelli and Bird -- <em>&#8220;Expectations, Outcomes, and Challenges of Modern Code Review&#8221;</em> -- studied code review at Microsoft through observation, interviews, and surveys and reached the same conclusion: <strong>while finding defects remains the main motivation for review, reviews are less about defects than expected. Less than 15% of issues discussed in code reviews relate directly to bugs.</strong> The primary benefits are knowledge transfer, increased team awareness, and the creation of alternative solutions.</p><p>A separate Microsoft Research study analysed 1.5 million review comments across five projects. The more files in a change, the lower the proportion of useful comments. Developers spent on average six hours per week reviewing others&#8217; changes. Reviews containing more than 20 files were already too big for effective review.</p><p>Large-scale industry data tells the same story. One major technology company, after examining nine million code reviews internally, cited <strong>knowledge transfer as the primary source of code-review ROI</strong>, not bug finding. Up to 75% of code review comments affect software evolvability and maintainability rather than functionality.</p><p>None of this means code review is worthless. It means its value lies in a different place than most people assume. <strong>If your organisation justifies pull requests primarily as a bug-catching mechanism, you are optimising for a benefit that peer-reviewed research says is marginal.</strong> The knowledge transfer benefit is real -- but a blocking asynchronous queue is one of the worst ways to achieve it.</p><div><hr></div><h2><strong>The Staggering Cost of Waiting</strong></h2><p>Here is a calculation that should disturb any engineering leader.</p><p>If a code change takes 10 minutes to make but waits 1 hour for review, it is waiting for 86% of its total lead time. If review takes 4 hours: 96% waiting. <strong>If review takes one working day -- which is common -- the code spends 99% of its existence waiting for a human to look at it.</strong></p><p>Martin Fowler cites a client that spent <strong>130,000 hours in 2020 waiting for 7,000 pull requests that had no comments</strong>. Ninety-one percent of their PRs received no comments at all. The vast majority were rubber-stamped without substantive review, yet the process still imposed enormous delay.</p><p>The damage compounds through context switching. Research shows that developers wait an average of four days for a pull request review, that 86% of pull requests are handled under context-switching conditions, and that context-switching pull requests are 223% slower than non-context-switching ones. When a developer opens a PR and moves on to something else, rebuilding the mental context of the original work takes 30-60 minutes -- if they rebuild it at all.</p><p>Don Reinertsen&#8217;s <em>The Principles of Product Development Flow</em> provides the theoretical lens: batch size is an economic tradeoff between holding cost and transaction cost, and halving batch sizes halves queues and halves cycle time. <strong>Pull requests create a perverse incentive toward larger batches</strong>: because the transaction cost of getting a review is high (waiting, context switching, reviewer availability), developers batch more changes into each PR to amortise the review cost. Larger PRs take longer to review, reviews are less effective on large changes, and the cycle reinforces itself.</p><p>Charity Majors frames the cost financially. A 6-person team requiring days to deploy would need 24 people to match the output of a team deploying continuously -- roughly $3.6 million in unnecessary salary costs. A 10-person team shipping weekly would need 80 people: $14 million in waste. Her target: <em>&#8220;Any merge triggers automatic deploy to production, completed in 15 minutes or less with no human intervention.&#8221;</em></p><div><hr></div><h2><strong>Async Reviews: The Vicious Cycle</strong></h2><p>Dragan Stepanovic examined async code review workflows across more than 30 active repositories and observed a counter-intuitive pattern: <strong>teams using small PRs with async code reviews tended to have lower throughput than teams using large PRs</strong>. The reason is a feedback loop. Waiting for review leads to high work-in-progress (developers pull in new work while waiting). High WIP means less availability for reviewing. Less availability means more async handoffs. More async handoffs mean more waiting.</p><p>Instead of forming one team, developers become what Stepanovic calls <em>&#8220;N teams of one person, often with different engineering cultures and coding practices.&#8221;</em> The isolation is the opposite of what high-performing teams need.</p><p>Chelsea Troy identifies the root cause as insufficient shared context. Asynchronous code review places a massive demand on the reviewer&#8217;s time because they lack the context the author built over hours or days. Pairing is more efficient precisely because the context transfer is immediate -- the work of understanding and reviewing happens in a condensed amount of time relative to the amount of work being done.</p><p>A caveat is warranted: these are practitioner observations, not controlled experiments. The causal mechanism is plausible but not empirically proven. It is possible that teams with lower throughput default to async review for other reasons. But the observations are consistent across multiple independent practitioners, and the logic of WIP accumulation under queuing theory is well established.</p><div><hr></div><h2><strong>What DORA Tells Us</strong></h2><p>If the practitioner arguments are directional, the DORA data provides the scale.</p><p><em>Accelerate</em>, by Nicole Forsgren, Jez Humble, and Gene Kim, is based on four years of research across 23,000 surveys from over 2,000 organisations. Their findings on trunk-based development are unambiguous: <strong>developing off trunk rather than long-lived feature branches correlated with higher delivery performance</strong>. High-performing teams had fewer than three active branches at any time, branches lasted less than a day, and there were no code freezes or stabilisation periods.</p><p>The DORA capabilities framework identifies trunk-based development as a key technical capability. Elite performers who meet reliability targets are 2.3 times more likely to use trunk-based development. Crucially, the framework explicitly calls out <em>&#8220;heavyweight code review processes&#8221;</em> as a barrier -- they push developers toward larger batches and delay merges.</p><p>The 2023 State of DevOps Report, based on 36,000+ professionals worldwide, found that <strong>accelerating the code review process alone can lead to a 50% improvement in software delivery performance</strong>. Note what this says: the problem is not review itself, but the speed of review. Make review instant -- through pairing, rapid turnaround, or automation -- and you keep the benefit without the cost.</p><p>An important caveat: the DORA research is correlational, not causal. It shows that high-performing teams tend to use trunk-based development, not that trunk-based development causes high performance. Teams that adopt TBD may also be better-funded, more skilled, or have stronger engineering cultures. But the consistency of the finding across years and across tens of thousands of respondents makes it the strongest industry evidence available.</p><p>The Minimum CD initiative, co-created by Jez Humble, is blunt about the implication: <strong>daily integration to trunk is non-negotiable. If your team is not integrating to trunk daily, you are not doing continuous integration.</strong> Pull requests, as typically practised, fail this test.</p><div><hr></div><h2><strong>A Growing Practitioner Consensus</strong></h2><p>ThoughtWorks placed <em>&#8220;Peer review equals pull request&#8221;</em> in the <strong>HOLD</strong> ring of their Technology Radar in April 2021 -- meaning &#8220;proceed with caution.&#8221; They noted that PRs create significant team bottlenecks, degrade review quality as overloaded reviewers begin rubber-stamping, and in one case, a client&#8217;s regulatory audit found that pull requests <em>did not satisfy compliance</em> because there was no evidence the code was actually read.</p><p>Kief Morris, also of ThoughtWorks and author of <em>Infrastructure as Code</em>, argues that pull requests add overhead designed for low-trust situations and that <strong>pull requests are not continuous integration -- CI is an alternative to pull requests, not a complement</strong>.</p><p>Dave Farley, co-author of <em>Continuous Delivery</em>, argues that pull requests are an artifact of branch-based development which deliberately isolates changes from mainline -- the opposite of what CI requires.</p><p>Jason Gorman asks the most piercing question: <em>&#8220;Ask not so much &#8216;How do we do Pull Requests?&#8217; but rather &#8216;Why do we need to do Pull Requests?&#8217;&#8221;</em> His answer: <strong>pull requests are a symptom of low trust, not a solution for low quality.</strong> Teams should address the root causes -- skills development, professional standards, pair programming -- rather than institutionalising slow inspection.</p><p>And perhaps the most quotable line comes from Jessica Kerr, amplified by Kent Beck: <em>&#8220;Pull requests are an improvement on working alone. But not on working together.&#8221;</em></p><div><hr></div><h2><strong>T*D: Build Quality In</strong></h2><p>If pull requests are Deming&#8217;s inspection, what is the alternative? What does it look like to build quality into the process itself?</p><p>The answer is a union of three practices I&#8217;ll call <strong>T*D</strong> -- a deliberate echo of TDD, because all three components share the same initials and the same philosophy of shifting quality left.</p><h3><strong>Test-Driven Development</strong></h3><p>Write a failing test. Write the minimum code to pass it. Refactor. Repeat.</p><p>Microsoft and IBM studies found that <strong>TDD reduced pre-release defect density by 40-90%</strong> compared to projects not using it, at a cost of 15-35% more development time. Kent Beck&#8217;s insight was that automated testing replaces fear with confidence: when stress increases, developers run tests rather than skipping them.</p><p>TDD&#8217;s role in replacing PRs is straightforward: when every line of code is written to satisfy a test, the code has already been verified before it reaches anyone. The automated safety net catches regressions immediately. You do not need a human gatekeeper to tell you whether the code works -- the tests do.</p><h3><strong>Trunk-Based Development</strong></h3><p>All developers merge to the main branch at least daily, working in small batches. Feature flags manage incomplete work.</p><p>The evidence, covered extensively above, is unambiguous: trunk-based development is correlated with higher delivery performance across every DORA metric. Paul Hammant&#8217;s trunkbaseddevelopment.com establishes that the ideal branch duration is one day maximum, with the smallest pieces being a quarter of a day. Feature flags eliminate the argument that &#8220;we need branches because features aren&#8217;t complete.&#8221;</p><p>TBD&#8217;s role in replacing PRs is structural: when changes are tiny (hours, not days), the review burden is minimal and can be handled through pairing or post-commit review. There is nothing to gate because there is nothing large enough to fear.</p><h3><strong>Team-Focused Development</strong></h3><p>Two or more developers work on the same code at the same time, providing continuous real-time review.</p><p>This is the component with the most nuance. <strong>The peer-reviewed evidence for pair programming&#8217;s quality benefits is substantial.</strong> Williams et al. (2000) found pairs produced 15% fewer defects. Hannay et al.&#8217;s 2009 meta-analysis -- the most comprehensive to date -- found a <em>&#8220;small significant positive overall effect on quality&#8221;</em>, strongest on complex tasks. Arisholm et al.&#8217;s 2007 study of 295 professional developers found a 48% increase in correct solutions on complex systems. Junior developers showed the strongest benefit: 149% improvement on complex tasks.</p><p>Most critically for our argument, Muller and Tichy (2005) directly compared pair programming to solo programming with peer review. <strong>Their central finding: pairs and solo developers with peer review achieve equivalent cost when quality is held constant.</strong> The review phase for solo developers adds enough overhead to match the cost of pairing. This means pair programming does not cost more than solo-plus-review -- it simply moves the quality assurance from after development to during it.</p><p>Ensemble (mob) programming extends this: the whole team works on the same thing, at the same time, on the same computer. Review happens instantly, after every line. The academic evidence for mob programming is still preliminary, but the practitioner consensus is strong: when the whole team creates together, there is no need for post-hoc inspection.</p><p>An honest caveat: <strong>no peer-reviewed study has directly proven that teams can safely eliminate post-hoc review when practising pair programming.</strong> The evidence shows cost-equivalence and comparable quality, not that one fully replaces the other. The argument that pairing eliminates the need for PRs is practitioner consensus extending academic evidence -- well-reasoned, but not empirically proven.</p><h3><strong>The Synthesis</strong></h3><p>When combined, these three practices create a system where quality is built in at every stage:</p><ul><li><p><strong>TDD</strong> means every line of code is verified by an automated test before it exists</p></li><li><p><strong>TBD</strong> means changes are tiny, integrated frequently, and never far from mainline</p></li><li><p><strong>TFD</strong> means human review happens continuously, during creation, not after</p></li></ul><p>Fear of bugs? TDD catches them at creation time. Need for review? Pair programming reviews continuously. Integration risk? TBD integrates frequently in small batches. Knowledge sharing? Co-creation shares knowledge in real time.</p><p>The result is what Dave Farley describes as <em>&#8220;Extreme Programming practices of ensemble programming and continuous code review that eliminate all waiting and waste.&#8221;</em></p><div><hr></div><h2><strong>The Quality Journey: From Fear to Flow</strong></h2><p>I want to address the hardest question honestly: what if your team genuinely isn&#8217;t ready to drop pull requests?</p><p>This is a real concern. Jason Gorman acknowledges that the need for PRs often indicates a real skills gap. Thierry de Pauw argues that distrust-driven PR adoption signals deeper problems -- legacy code comprehension issues, blame cultures, or absent trust in engineers. <strong>No process fixes fundamental cultural dysfunction.</strong> But the answer is not to permanently institutionalise a slow inspection process. It is to address the root causes directly while gradually reducing dependence on gating review.</p><p>Here is a transition path, drawn from practitioner experience:</p><p><strong>Stage 1: Optimise PRs.</strong> If you&#8217;re not ready to eliminate them, make them less harmful. Enforce smaller PRs (under 200 lines). Set review SLAs (under 4 hours). Automate all style and lint checks. Reduce required approvers to one. This alone can dramatically improve flow.</p><p><strong>Stage 2: Ship/Show/Ask.</strong> Rouan Wilsenach&#8217;s model, published on martinfowler.com, categorises changes into three types. <em>Ship</em>: merge directly to mainline without review -- for routine changes where the developer is confident. <em>Show</em>: open a PR, merge immediately without waiting, and let review happen post-merge. <em>Ask</em>: open a PR and wait for discussion -- for complex or uncertain changes. This reduces the proportion of work that needs blocking review and starts building the muscle of trust.</p><p><strong>Stage 3: Pair programming plus trunk-based development.</strong> Pair on all production code. Merge to trunk multiple times daily. Automated tests run on every commit. Post-commit review covers anything that wasn&#8217;t paired on.</p><p><strong>Stage 4: Ensemble programming plus trunk-based development.</strong> The full team collaborates on code. Continuous review during development. Direct trunk commits. Feature flags for incomplete work.</p><p><strong>Stage 5: Full T*D.</strong> All three practices operating together. Quality built in through TDD. Continuous integration through TBD. Continuous review through pair or ensemble. A fast automated deployment pipeline. Feature flags managing incomplete work. This is not a fantasy -- it is how the highest-performing teams in the industry operate.</p><p>Thierry de Pauw documents a specific intermediate approach: <strong>non-blocking continuous code review</strong>. Reviews happen on mainline after merging, on a per-feature level. Teams add a &#8220;To Review&#8221; column to their board. Developers check for pending reviews before starting new work. No hierarchical requirements -- seniors review juniors, juniors review seniors. Unreviewed code may reach production, but only if it passed automated testing. Notably, internal auditors in one organisation found this model <em>superior</em> to traditional checkbox-based compliance.</p><p><strong>The key insight from Charity Majors is that speed and safety are not trade-offs -- speed IS safety.</strong> <em>&#8220;Ship a single changeset by a single dev at a time, making it easy to isolate the owner of any problem, preventing the blast radius from expanding, and making it easy to fix while the intended effects of the code are fresh in their mind.&#8221;</em></p><div><hr></div><h2><strong>Trust Is a Gradient</strong></h2><p>One final nuance. The argument that PRs are designed for &#8220;untrusted strangers&#8221; and therefore wrong for &#8220;trusted teams&#8221; implies a binary. In reality, trust is a gradient.</p><p>A small, co-located team of five people who pair daily has very high trust. The case against async PRs is strongest here. A large organisation with hundreds of engineers across multiple teams has lower cross-team trust. Within a team, pairing and trunk-based development may be ideal; across team boundaries, lightweight PRs with fast SLAs may still serve a purpose. For inner-source or platform teams that accept contributions from semi-external developers, a gatekeeper model may remain appropriate -- but even then, the goal should be fast, non-blocking review, not multi-day async queues.</p><p>For distributed teams with significant timezone overlap, pair programming via screen sharing works during overlap windows, with non-blocking post-commit review for work done outside overlap. For teams with minimal overlap, async PR workflows may be the least-bad option, but should be optimised aggressively: small changes, 4-hour review SLAs, automated quality gates, and a single required reviewer.</p><p><strong>The antipattern diagnosis applies most strongly to same-team, same-context work where async blocking PRs replace trust that should already exist.</strong> As contributor trust decreases, some form of gatekeeping becomes more defensible. But it should always be as fast and as non-blocking as possible.</p><div><hr></div><h2><strong>Conclusions</strong></h2><p>The evidence, while not without gaps, points in a clear and consistent direction.</p><p><strong>Pull requests, as commonly practised -- large changes sitting in async queues for hours or days -- are an antipattern for private software teams.</strong> They were designed for a context (open source, untrusted contributors) that does not apply to most organisations. Peer-reviewed research shows they catch few bugs. Industry data shows the waiting time they impose constitutes almost all of a change&#8217;s lead time. Practitioner experience across dozens of independent voices confirms that better alternatives exist.</p><p>The alternative is not &#8220;no review.&#8221; It is <em>better</em> review -- built into the process rather than bolted on at the end. Test-driven development builds quality into every line. Trunk-based development keeps changes small and integration continuous. Pair and ensemble programming provide review that is immediate, contextual, and collaborative rather than delayed, decontextualised, and adversarial.</p><p>This is not about being reckless. It is about being rigorous in the right way. Deming&#8217;s insight, now more than 40 years old, still holds: <strong>you cannot inspect quality into a product. You must build it in.</strong> The teams that ship the fastest, with the fewest defects, are not the ones with the most elaborate gating processes. They are the ones that invested in the practices, skills, and trust that make gating unnecessary.</p><p>The path from where you are to where you want to be is gradual. Start by making your PRs smaller and faster. Then start asking which ones you don&#8217;t need at all. Then start pairing. Then start driving with tests. At each stage, the fear recedes a little, the flow increases, and the quality -- paradoxically, to those who equate inspection with safety -- gets better.</p><p><strong>The question is not &#8220;How do we do pull requests better?&#8221; The question is &#8220;Why do we still need them?&#8221;</strong></p><div><hr></div><h2><strong>References</strong></h2><h3><strong>Peer-Reviewed Academic Research</strong></h3><ol><li><p>Microsoft Research. <em>&#8220;Code Reviews Do Not Find Bugs: How the Current Code Review Best Practice Slows Us Down.&#8221;</em> IEEE, 2015. <a href="https://www.microsoft.com/en-us/research/publication/code-reviews-do-not-find-bugs-how-the-current-code-review-best-practice-slows-us-down/">https://www.microsoft.com/en-us/research/publication/code-reviews-do-not-find-bugs-how-the-current-code-review-best-practice-slows-us-down/</a></p></li><li><p>Bacchelli, A. &amp; Bird, C. <em>&#8220;Expectations, Outcomes, and Challenges of Modern Code Review.&#8221;</em> ICSE 2013. <a href="https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/ICSE202013-codereview.pdf">https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/ICSE202013-codereview.pdf</a></p></li><li><p>Bosu, A. et al. <em>&#8220;Characteristics of Useful Code Reviews: An Empirical Study at Microsoft.&#8221;</em> Microsoft Research, 2015. <a href="https://www.microsoft.com/en-us/research/publication/characteristics-of-useful-code-reviews-an-empirical-study-at-microsoft/">https://www.microsoft.com/en-us/research/publication/characteristics-of-useful-code-reviews-an-empirical-study-at-microsoft/</a></p></li><li><p>Williams, L. et al. <em>&#8220;Strengthening the Case for Pair Programming.&#8221;</em> IEEE Software, Vol. 17, No. 4, 2000. <a href="https://ieeexplore.ieee.org/document/854064/">https://ieeexplore.ieee.org/document/854064/</a></p></li><li><p>Hannay, J.E. et al. <em>&#8220;The Effectiveness of Pair Programming: A Meta-Analysis.&#8221;</em> Information and Software Technology, Vol. 51, No. 7, 2009. <a href="https://www.sciencedirect.com/science/article/abs/pii/S0950584909000123">https://www.sciencedirect.com/science/article/abs/pii/S0950584909000123</a></p></li><li><p>Arisholm, E. et al. <em>&#8220;Evaluating Pair Programming with Respect to System Complexity and Programmer Expertise.&#8221;</em> IEEE Transactions on Software Engineering, Vol. 33, No. 2, 2007. <a href="https://ieeexplore.ieee.org/document/4052584/">https://ieeexplore.ieee.org/document/4052584/</a></p></li><li><p>Muller, M.M. &amp; Tichy, W.F. <em>&#8220;Two Controlled Experiments Concerning the Comparison of Pair Programming to Peer Review.&#8221;</em> Journal of Systems and Software, Vol. 78, No. 2, 2005. <a href="https://www.sciencedirect.com/science/article/abs/pii/S0164121205000038">https://www.sciencedirect.com/science/article/abs/pii/S0164121205000038</a></p></li><li><p>Madeyski, L. <em>&#8220;An Empirical Study on Design Quality Improvement from Best-Practice Inspection and Pair Programming.&#8221;</em> Springer LNCS Vol. 4034, 2006. <a href="https://link.springer.com/chapter/10.1007/11767718_27">https://link.springer.com/chapter/10.1007/11767718_27</a></p></li><li><p>di Bella, E. et al. <em>&#8220;Pair Programming and Software Defects -- A Large, Industrial Case Study.&#8221;</em> IEEE TSE, Vol. 39, No. 7, 2013. <a href="https://ieeexplore.ieee.org/document/6331491/">https://ieeexplore.ieee.org/document/6331491/</a></p></li><li><p>Nagappan, N. et al. <em>&#8220;Realizing Quality Improvement Through Test Driven Development.&#8221;</em> Microsoft/IBM study. <a href="https://www.researchgate.net/publication/258126622_How_Effective_is_Test_Driven_Development">https://www.researchgate.net/publication/258126622_How_Effective_is_Test_Driven_Development</a></p></li><li><p>Edmondson, A. <em>&#8220;Psychological Safety and Learning Behavior in Work Teams.&#8221;</em> Administrative Science Quarterly, 1999.</p></li></ol><h3><strong>Industry Reports and Guides</strong></h3><ol start="12"><li><p>Forsgren, N., Humble, J., Kim, G. <em>Accelerate: The Science of Lean Software and DevOps.</em> IT Revolution, 2018.</p></li><li><p>DORA/Google. <em>&#8220;Capabilities: Trunk-based Development.&#8221;</em> <a href="https://dora.dev/capabilities/trunk-based-development/">https://dora.dev/capabilities/trunk-based-development/</a></p></li><li><p>Google Cloud. <em>&#8220;Accelerate State of DevOps Report 2023.&#8221;</em> <a href="https://dora.dev/research/2023/dora-report/">https://dora.dev/research/2023/dora-report/</a></p></li><li><p>ThoughtWorks. <em>&#8220;Peer Review Equals Pull Request.&#8221;</em> Technology Radar, April 2021. <a href="https://www.thoughtworks.com/radar/techniques/peer-review-equals-pull-request">https://www.thoughtworks.com/radar/techniques/peer-review-equals-pull-request</a></p></li><li><p>SmartBear/Cisco. <em>&#8220;Code Review at Cisco Systems.&#8221;</em> 2006. <a href="https://static0.smartbear.co/support/media/resources/cc/book/code-review-cisco-case-study.pdf">https://static0.smartbear.co/support/media/resources/cc/book/code-review-cisco-case-study.pdf</a></p></li></ol><h3><strong>Practitioner Sources</strong></h3><ol start="17"><li><p>Deming, W.E. <em>Out of the Crisis.</em> MIT Press, 1982. Also: Deming Institute, <em>&#8220;Dr. Deming&#8217;s 14 Points for Management.&#8221;</em> <a href="https://deming.org/explore/fourteen-points/">https://deming.org/explore/fourteen-points/</a></p></li><li><p>Deming Institute. <em>&#8220;Software Code Reviews from a Deming Perspective.&#8221;</em> <a href="https://deming.org/software-code-reviews-from-a-deming-perspective/">https://deming.org/software-code-reviews-from-a-deming-perspective/</a></p></li><li><p>de Pauw, T. <em>&#8220;The Good and the Dysfunctional of Pull Requests.&#8221;</em> ThinkingLabs, 2024. <a href="https://thinkinglabs.io/articles/2024/02/22/the-good-and-the-dysfunctional-of-pull-requests.html">https://thinkinglabs.io/articles/2024/02/22/the-good-and-the-dysfunctional-of-pull-requests.html</a></p></li><li><p>de Pauw, T. <em>&#8220;Non-Blocking, Continuous Code Reviews -- A Case Study.&#8221;</em> ThinkingLabs, 2023. <a href="https://thinkinglabs.io/articles/2023/05/02/non-blocking-continuous-code-reviews-a-case-study.html">https://thinkinglabs.io/articles/2023/05/02/non-blocking-continuous-code-reviews-a-case-study.html</a></p></li><li><p>Stepanovic, D. <em>&#8220;Async Code Reviews Are Killing Your Company&#8217;s Throughput.&#8221;</em> 2021-2023. <a href="https://www.slideshare.net/kobac/async-code-reviews-are-killing-your-companys-throughput-248758692">https://www.slideshare.net/kobac/async-code-reviews-are-killing-your-companys-throughput-248758692</a></p></li><li><p>Troy, C. <em>&#8220;Reviewing Pull Requests.&#8221;</em> chelseatroy.com, 2019. <a href="https://chelseatroy.com/2019/12/18/reviewing-pull-requests/">https://chelseatroy.com/2019/12/18/reviewing-pull-requests/</a></p></li><li><p>Morris, K. <em>&#8220;Why Your Team Doesn&#8217;t Need to Use Pull Requests.&#8221;</em> infrastructure-as-code.com, 2021. <a href="https://infrastructure-as-code.com/posts/pull-requests.html">https://infrastructure-as-code.com/posts/pull-requests.html</a></p></li><li><p>Fowler, M. <em>&#8220;bliki: Pull Request.&#8221;</em> <a href="https://martinfowler.com/bliki/PullRequest.html">https://martinfowler.com/bliki/PullRequest.html</a></p></li><li><p>Fowler, M. <em>&#8220;Continuous Integration.&#8221;</em> Updated 2024. <a href="https://martinfowler.com/articles/continuousIntegration.html">https://martinfowler.com/articles/continuousIntegration.html</a></p></li><li><p>Farley, D. <em>&#8220;Continuous Integration and Feature Branching&#8221;</em> (blog) </p><p>https://www.davefarley.net/?p=247</p></li><li><p> / <em>&#8220;You NEED to Stop Using Pull Requests&#8221;</em> (YouTube video).</p></li><li><p>Gorman, J. <em>&#8220;Pull Requests, Defensive Programming -- It&#8217;s All About Trust.&#8221;</em> Codemanship, 2020. <a href="https://codemanship.wordpress.com/2020/09/12/pull-requests-defensive-programming-its-all-about-trust/">https://codemanship.wordpress.com/2020/09/12/pull-requests-defensive-programming-its-all-about-trust/</a></p></li><li><p>Vocke, H. <em>&#8220;You Might Be Better Off Without Pull Requests.&#8221;</em> <a href="https://hamvocke.com/blog/better-off-without-pull-requests/">https://hamvocke.com/blog/better-off-without-pull-requests/</a></p></li><li><p>Beck, K. (via Twitter/X): <em>&#8220;Pull requests are an improvement on working alone but not on working together.&#8221;</em> </p></li><li><p>Kerr, J. <em>&#8220;Those Pesky Pull Request Reviews.&#8221;</em> jessitron.com, 2021. <a href="https://jessitron.com/2021/03/27/those-pesky-pull-request-reviews/">https://jessitron.com/2021/03/27/those-pesky-pull-request-reviews/</a></p></li><li><p>Majors, C. <em>&#8220;How Much Is Your Fear of Continuous Deployment Costing You?&#8221;</em> charity.wtf, 2021. <a href="https://charity.wtf/2021/02/19/how-much-is-your-fear-costing-you/">https://charity.wtf/2021/02/19/how-much-is-your-fear-costing-you/</a></p></li><li><p>Reinertsen, D. <em>The Principles of Product Development Flow.</em> Celeritas Publishing, 2009.</p></li><li><p>Wilsenach, R. <em>&#8220;Ship / Show / Ask.&#8221;</em> Published on martinfowler.com. <a href="https://martinfowler.com/articles/ship-show-ask.html">https://martinfowler.com/articles/ship-show-ask.html</a></p></li><li><p>Hammant, P. <em>&#8220;Trunk Based Development.&#8221;</em> </p><p>https://trunkbaseddevelopment.com/</p></li><li><p>Minimum CD. <em>&#8220;Minimum Viable Continuous Delivery.&#8221;</em> </p><p>https://minimumcd.org/</p></li><li><p>Meek, J. <em>&#8220;Pull Requests are an Anti-Pattern.&#8221;</em> Substack.</p></li><li><p>Zuill, W. <em>&#8220;Mob Programming: A Whole Team Approach.&#8221;</em> </p><p>https://mobprogramming.org/</p></li><li><p>Beck, K. <em>Test-Driven Development: By Example.</em> Addison-Wesley, 2002.</p></li><li><p>Humble, J. &amp; Farley, D. <em>Continuous Delivery.</em> Addison-Wesley, 2010.</p></li><li><p>Bogard, J. <em>&#8220;Trunk-Based Development or Pull Requests: Why Not Both?&#8221;</em> <a href="https://www.jimmybogard.com/trunk-based-development-or-pull-requests-why-not-both/">https://www.jimmybogard.com/trunk-based-development-or-pull-requests-why-not-both/</a></p></li></ol>]]></content:encoded></item><item><title><![CDATA[What Makes a High-Performing Team?]]></title><description><![CDATA[What a decade of research across 39,000 professionals tells us about building teams that actually deliver]]></description><link>https://a4al6a.substack.com/p/what-makes-a-high-performing-team</link><guid isPermaLink="false">https://a4al6a.substack.com/p/what-makes-a-high-performing-team</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Wed, 18 Mar 2026 13:41:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Over the years, I&#8217;ve been asked countless times what a high-performing team actually looks like and how to build or coach one. As a consultant who has helped many organisations embrace a better way of working, I have a pretty good idea of what the hallmarks of a high-performing team are &#8212; both cultural and technical. But personal experience, however hard-won, only takes you so far. There is also a vast body of research around this topic, and I wanted to surface it all in one place.</p><p>Over the past decade, rigorous evidence &#8212; spanning Google&#8217;s Project Aristotle, the DORA research programme, academic studies, and hard-won lessons from companies like Spotify, Amazon, and GitLab &#8212; has converged on a surprisingly consistent set of themes. These studies don&#8217;t all measure the same things or use the same methods, and they don&#8217;t agree on everything. What follows is not a unified theory of team performance &#8212; it&#8217;s a curated synthesis of the strongest evidence available. Some of the conclusions are intuitive. Others are deeply counterintuitive. All of them are grounded in data, though the strength of that grounding varies, and I&#8217;ll try to be honest about where.</p><div><hr></div><h2><strong>It&#8217;s Not About the Stars &#8212; It&#8217;s About the Dynamics</strong></h2><p>One of the most influential studies in this space is Google&#8217;s Project Aristotle, a two-year internal study of 180 teams examining hundreds of attributes. <strong>The single most critical factor in team success was not individual talent, technical skill, or seniority &#8212; it was psychological safety</strong>, defined by <a href="https://www.linkedin.com/in/amycedmondson/">Harvard professor Amy Edmondson</a> as &#8220;a shared belief held by members of a team that the team is safe for interpersonal risk taking.&#8221;</p><p>Google identified five key dynamics, in decreasing order of importance: psychological safety, dependability, structure and clarity, meaning, and impact. What did not matter <em>in that study, with those measures</em> was equally revealing: colocation, consensus decision-making, individual extroversion, individual performance, workload size, seniority, team size, and tenure showed no significant correlation with team effectiveness. That doesn&#8217;t mean these factors never matter &#8212; anyone who&#8217;s managed a distributed team or scaled a 30-person department knows better. But within Google&#8217;s context, they weren&#8217;t the differentiators.</p><p>This finding challenges the persistent myth of the &#8220;10x engineer.&#8221; The evidence says that <em>who</em> is on the team matters far less than <em>how</em> the team works together. Edmondson&#8217;s foundational 1999 study &#8212; one of the most cited papers in organisational psychology &#8212; established that psychological safety enables learning behaviour, which in turn drives performance improvement. The DORA research programme validated the connection: generative cultures that include psychological safety predict both software delivery performance and organisational performance.</p><p>An important nuance: psychological safety <em>enables</em> performance, but it doesn&#8217;t <em>cause</em> it on its own. Plenty of teams feel safe and still don&#8217;t ship anything useful. Safety without standards, direction, or healthy pressure can produce very comfortable mediocrity. That&#8217;s why Google&#8217;s research found it was the foundation &#8212; the other four dynamics (dependability, structure, meaning, impact) stack on top. Safety alone is not the goal. Safety <em>in service of high expectations</em> is.</p><p>The practical implication is profound. <strong>Hiring alone won&#8217;t get you there. You have to build the conditions for high performance, not just recruit for it.</strong></p><div><hr></div><h2><strong>Culture Is Not a Perk &#8212; It&#8217;s a Performance Predictor</strong></h2><p>Sociologist Ron Westrum identified three types of organisational culture: pathological (power-oriented, fear-based), bureaucratic (rule-oriented, turf-protecting), and generative (performance-oriented, mission-focused). <strong>The DORA research demonstrated that generative culture predicts software delivery performance, organisational performance, and job satisfaction.</strong></p><p>Generative cultures are characterised by good information flow, high cooperation and trust, bridging between teams, and conscious inquiry. But here&#8217;s what makes this finding actionable: the research shows that adopting technical practices &#8212; continuous integration, test automation, trunk-based development &#8212; actively shifts culture over time. Organisations do not need to &#8220;fix culture first.&#8221; The relationship is bidirectional: culture enables technical practices, and technical practices shape culture.</p><p>A word of caution, though. This bidirectional claim is easier to write than to live. In messy organisations &#8212; the ones with blame-heavy post-mortems, territorial middle managers, and approval chains that strangle momentum &#8212; trying to introduce trunk-based development or continuous delivery can backfire. The practice gets adopted in name only, or it surfaces conflicts that the culture isn&#8217;t ready to handle. The research says the path exists; it doesn&#8217;t say the path is clean. In reality, it&#8217;s often political, uncomfortable, and slow. But the evidence that it <em>can</em> work, even if it doesn&#8217;t always, is worth something.</p><div><hr></div><h2><strong>Size Matters &#8212; But Not the Way You Think</strong></h2><p>Amazon&#8217;s two-pizza teams, Hackman and Vidmar&#8217;s research, Scrum guidance, and Dunbar&#8217;s number all point in the same direction: the optimal team size is somewhere around 5 to 9 members. These sources come from different domains &#8212; organisational psychology, military anthropology, practitioner heuristics, corporate operating models &#8212; and they aren&#8217;t studying the same phenomenon in the same way. But the convergence is suggestive. The underlying mechanism is well-understood: communication channels grow quadratically &#8212; a 10-person team has three times the channels of a 6-person team. Beyond that threshold, productivity per person declines as coordination costs overwhelm individual contributions.</p><p>And when a project is already late? Fred Brooks&#8217;s 1975 observation remains validated by five decades of experience: &#8220;adding manpower to a late software project makes it later.&#8221; The ramp-up time for new members, combined with the explosion in communication overhead and the non-divisibility of many software tasks, means that the intuitive response &#8212; throw more people at the problem &#8212; is usually the wrong one.</p><p>This applies to team composition as well &#8212; but with an essential caveat. <strong>Cognitive diversity drives innovation only when combined with psychological safety and inclusive practices. Without those conditions, diverse teams may actually underperform homogeneous ones.</strong> When the conditions are right, however, the effects are striking: research published by the National Institutes of Health found that cognitively diverse teams solve complex tasks significantly faster and produce more innovative solutions. The implication is that hiring for diversity must be accompanied by investing in inclusive team dynamics &#8212; one without the other can backfire.</p><div><hr></div><h2><strong>Your Software Architecture Is Your Org Chart (and Vice Versa)</strong></h2><p>In 1967, Melvin Conway observed that organisations design systems that mirror their communication structures. Nearly sixty years later, MIT and Harvard Business School researchers found &#8220;strong evidence to support the mirroring hypothesis&#8221; &#8212; products built by loosely-coupled organisations are significantly more modular than those from tightly-coupled ones.</p><p>This has moved from observation to strategic tool. Martin Fowler endorses the &#8220;Inverse Conway Manoeuvre&#8221; &#8212; deliberately structuring teams to produce the desired architecture. Matthew Skelton and Manuel Pais built the entire Team Topologies framework on this principle, defining four team types: stream-aligned (delivering direct customer value), platform (reducing cognitive load via self-service), enabling (coaching other teams temporarily), and complicated-subsystem (handling specialist technical components).</p><p><strong>The key insight from Team Topologies is that the primary benefit of a platform team is not efficiency &#8212; it&#8217;s reducing the cognitive load on stream-aligned teams.</strong> When cognitive load is managed, teams can focus on their core mission. When it isn&#8217;t, every team becomes a bottleneck for every other team.</p><div><hr></div><h2><strong>Measure What Matters &#8212; And Only What Matters</strong></h2><p>The DORA research programme originally identified four key metrics that predict software delivery performance: deployment frequency, lead time for changes, change failure rate, and time to restore service. A fifth metric &#8212; reliability &#8212; was added in subsequent years, and the 2024 report further refined the framework by replacing &#8220;mean time to recovery&#8221; with failed deployment recovery time (moving it from stability to throughput) and introducing rework rate as a new stability measure.</p><p>From 2018 through 2024, DORA grouped teams into four performance clusters &#8212; elite, high, medium, and low &#8212; derived each year via cluster analysis of survey respondents. The table below, based on the original <em>Accelerate</em> benchmarks, illustrates the scale of the gap between top and bottom performers:</p><p>ClusterDeploy FrequencyLead TimeFailure RateRecovery TimeEliteOn demand&lt; 1 hour~5%&lt; 1 hourHighDaily to weekly1 day - 1 week~10%&lt; 1 dayMediumWeekly to monthly1 week - 1 month~15%1 day - 1 weekLowMonthly+1 month+~64%1 week+</p><p>In 2025, DORA retired these tiers entirely and replaced them with seven team archetypes &#8212; such as &#8220;Harmonious High Achiever,&#8221; &#8220;Stable and Methodical,&#8221; and &#8220;Legacy Bottleneck&#8221; &#8212; that blend delivery metrics with human factors like burnout and friction. The shift reflects a recognition that a single linear ranking oversimplifies how teams actually perform: a team can have high throughput but crushing burnout, or modest deployment frequency but exceptional stability. For practitioners, the implication is significant: &#8220;high-performing&#8221; is no longer a single target to aim for. It means understanding <em>which</em> archetype your team resembles and addressing its specific weaknesses, rather than chasing a universal ideal. The field is moving from &#8220;are we elite?&#8221; toward &#8220;what kind of team are we, and what&#8217;s holding us back?&#8221;</p><p>Regardless of how the clusters are drawn, one finding has remained consistent across every year of the research: <strong>throughput and stability are not trade-offs &#8212; the highest-performing teams excel at both.</strong> This demolishes the common excuse that moving fast requires sacrificing quality. If you&#8217;re deploying monthly, aim for weekly first &#8212; not &#8220;on demand.&#8221;</p><p>But beware the wrong metrics. McKinsey&#8217;s 2023 attempt to measure individual developer productivity generated fierce backlash from the engineering community, with critics including Gergely Orosz and Kent Beck arguing the framework would &#8220;do far more harm than good.&#8221; <strong>Among engineering leaders, the prevailing view is that individual-level productivity metrics are counterproductive and damage trust.</strong> Measure teams and systems, not individuals. The SPACE framework &#8212; Satisfaction, Performance, Activity, Communication, and Efficiency &#8212; formalises the principle that developer productivity cannot be reduced to a single dimension or metric. At least three of the five dimensions should be measured simultaneously to avoid misleading conclusions.</p><div><hr></div><h2><strong>The Practices That Separate the Best From the Rest</strong></h2><p>Technical practices are not just implementation details &#8212; they are culture shapers. The Accelerate research found that teams practising trunk-based development deploy over two hundred times more frequently with over a hundred times faster lead times than the lowest-performing teams. Teams that work off trunk or branches lasting less than a day have significantly higher performance.</p><p>A necessary caveat: most of this evidence is correlational, not causal. High-performing teams practice trunk-based development, but high-performing teams also tend to be the ones <em>capable</em> of adopting these practices well. The arrow of causation isn&#8217;t always clear. Treat these practices as strongly associated with high performance rather than as guaranteed recipes for it. And note that these practices done badly can actively hurt: trunk-based development without solid CI and feature flags leads to broken mainlines; test automation without discipline produces brittle suites that slow everything down and erode trust in the tests themselves.</p><p>Code review velocity matters too. Google&#8217;s code review practices demand a maximum one-business-day response time, and the overwhelming majority of their engineers report satisfaction with the process. The primary purpose is not gatekeeping &#8212; it&#8217;s &#8220;making sure that the overall code health of Google&#8217;s code base is improving over time.&#8221; Speed and quality in code review are not trade-offs. Fast turnaround prevents context-switching costs and unblocks developers, while the review itself maintains code quality and spreads knowledge.</p><p>Test automation is another proven driver. When Google Web Server hit a crisis in 2005 &#8212; 80% of pushes causing user-impacting bugs &#8212; mandatory test automation and the &#8220;Test Certified&#8221; programme resolved the crisis within two years. Mike Cohn&#8217;s test pyramid (many unit tests, fewer integration tests, few end-to-end tests) remains the most validated testing strategy.</p><p>Pair programming, by contrast, shows mixed results. A meta-analysis found a small positive effect on quality and a medium positive effect on duration, but a medium negative effect on effort. Quality benefits are strongest for complex tasks. The evidence supports pair programming as a tool to deploy selectively &#8212; for complex tasks, knowledge transfer, and onboarding &#8212; not as a mandatory practice.</p><div><hr></div><h2><strong>Autonomy With Alignment &#8212; The Balancing Act</strong></h2><p>High-performing teams need both freedom and direction. OKRs (Objectives and Key Results) provide the mechanism: at Google, the majority of OKRs are developed bottom-up and then aligned to company goals. Fully top-down OKRs reduce autonomy; fully bottom-up OKRs risk misalignment. This balance enables what the DORA research calls &#8220;loosely coupled architecture with tightly aligned goals.&#8221;</p><p>Spotify&#8217;s organisational model &#8212; squads, tribes, chapters, guilds &#8212; offers a cautionary tale about the limits of structure. Spotify themselves have publicly acknowledged the model never fully worked as described. Tribes became siloed, cross-tribe coordination proved difficult, guilds lost effectiveness at scale, and high autonomy without collaboration processes led to duplicated effort. Companies that adopted the structure without the underlying culture of trust and psychological safety consistently failed.</p><p>The lesson: <strong>organisational structure is necessary but insufficient. It works only when underlaid by generative culture.</strong> You can&#8217;t copy another company&#8217;s org chart and expect their results.</p><div><hr></div><h2><strong>Developer Experience Is a Strategic Lever</strong></h2><p>The DevEx framework, published in ACM Queue in 2023, distils developer experience to three core dimensions: feedback loops (speed and quality of responses to developer actions), cognitive load (mental processing required for tasks), and flow state (energised focus, full involvement, and enjoyment).</p><p>All three are actionable. Feedback loops can be measured through build times, CI/CD pipeline duration, and code review turnaround. Cognitive load can be reduced through better documentation, simpler architecture, and fewer context switches. Flow state can be enabled by reducing interruptions and providing blocks of uninterrupted time.</p><p>Asynchronous communication plays a role here. GitLab, a fully remote company with 2,000+ employees, operates on an &#8220;async-first&#8221; basis. Async communication minimises distractions and enables deeper focus, but success requires deliberate investment in documentation and clear communication guidelines. The trade-off is real: async can slow decision-making and reduce personal connection if not managed carefully.</p><div><hr></div><h2><strong>Invest in Beginnings and Belonging</strong></h2><p>According to widely-cited research, organisations with strong onboarding processes see dramatically higher new hire retention and productivity &#8212; though these figures originate from a single industry study and should be treated as directional rather than precise. More robustly sourced: Google found that new hires paired with buddies reached full efficiency significantly faster, and one company reduced time-to-productivity from six weeks to ten days through structured onboarding with automation.</p><p>The most effective onboarding programmes share common elements: hardware and access configured before day one, a buddy or mentor pairing, a 30/60/90-day milestone framework, a first meaningful commit within the first week, and architecture documentation that explains <em>why</em>, not just <em>how</em>.</p><p>Mentorship extends beyond onboarding. Deloitte found that millennials with mentors are roughly twice as likely to stay at their organisation long-term, and employees involved in mentoring programmes have markedly higher retention rates. The evidence broadly supports mentorship as beneficial, though specific, evidence-based guidance for implementing it in software engineering contexts remains a real gap. Most of the research is cross-industry, and what works for management consulting or sales may not transfer directly to a codebase with steep domain complexity. Teams should treat mentorship as a high-probability bet worth making, while measuring retention and ramp-up outcomes locally rather than assuming the general statistics apply.</p><div><hr></div><h2><strong>The Anti-Patterns That Destroy Performance</strong></h2><p>Knowing what to avoid is as important as knowing what to do.</p><p><strong>Burnout is not an individual weakness &#8212; it&#8217;s an organisational design problem.</strong> The Accelerate research identifies six organisational risk factors: work overload, lack of control, insufficient rewards, breakdown of community, absence of fairness, and value conflicts. The vast majority of developers report experiencing work-related burnout. The 2024 DORA report found a specific, actionable intervention: teams with stable priorities face significantly less of it.</p><p>Measuring developers by lines of code, commit count, or individual output metrics creates perverse incentives and damages team culture. Organisations with generative culture &#8212; which avoids such metrics &#8212; consistently outperform those that don&#8217;t. The SPACE framework was explicitly designed as an alternative to these reductive metrics.</p><p>Platform engineering can backfire. The 2024 DORA report found measurable decreases in both throughput and change stability when teams were required to exclusively use internal platforms. The benefit comes from voluntary adoption of well-designed platforms, not mandated use.</p><p>And AI? The early data is more nuanced than the hype suggests. The 2024 DORA report found that as AI adoption increases, delivery stability drops measurably &#8212; even as developers <em>feel</em> more productive. One plausible interpretation &#8212; though not yet proven &#8212; is that AI-assisted coding tends to produce larger changesets, and larger changesets introduce more risk. Without AI-specific code review practices and batch size discipline, teams ship more code but break more things. This is not an argument against AI adoption &#8212; it&#8217;s an argument for treating it as a practice that requires the same measurement rigour as any other. Monitor your change failure rate and batch sizes before and after AI rollout. If stability drops, slow down and adjust.</p><div><hr></div><h2><strong>The Bottom Line</strong></h2><p>Before listing what the research points to, it&#8217;s worth naming what it largely leaves out: the external forces that often dominate outcomes more than any internal practice. Funding models, leadership churn, regulatory pressure, market dynamics, reorgs &#8212; these shape what teams can actually do far more than whether they&#8217;ve adopted trunk-based development. The research tends to study teams in relatively stable contexts. If your organisation is in the middle of a merger, a layoff, or a pivot, the advice here still applies in principle, but the path to applying it is much harder and much more political than the research alone suggests.</p><p>With that caveat, there are clear patterns across these studies, even if the details vary. High-performing software teams are not built by assembling the most talented individuals. <strong>They are built by creating the conditions in which ordinary professionals can do extraordinary work together.</strong></p><p>Those conditions are:</p><ol><li><p><strong>Psychological safety</strong> as the foundation &#8212; people must feel safe to take risks, ask questions, and admit mistakes</p></li><li><p><strong>Generative culture</strong> &#8212; mission-focused, high-trust, with good information flow</p></li><li><p><strong>Small, cognitively diverse teams</strong> &#8212; 5 to 9 members with managed cognitive load</p></li><li><p><strong>Technical excellence</strong> &#8212; CI/CD, trunk-based development, test automation, fast code reviews</p></li><li><p><strong>Autonomy with alignment</strong> &#8212; bottom-up goals connected to organisational mission</p></li><li><p><strong>Great developer experience</strong> &#8212; fast feedback loops, low cognitive load, protected flow state</p></li><li><p><strong>Investment in people</strong> &#8212; structured onboarding, mentorship, inclusive practices</p></li><li><p><strong>The right metrics</strong> &#8212; team-level outcomes, not individual activity</p></li></ol><p>The relationship between these elements is not linear. Technical practices shape culture. Culture enables technical practices. Structure constrains architecture. Architecture constrains structure. <strong>The organisations that succeed are the ones that treat all of these as a system, not a checklist.</strong></p><p>If I had to distil it even further, the research points to what I think of as <strong>T*D &#8212; three practices that, taken together, capture the essence of high-performing teams. Trunk-based development: integrate continuously, keep branches short-lived, and ship in small increments. Test-driven development: build quality in from the start rather than inspecting it in at the end. And team-focused development &#8212; what some call social programming &#8212; pairing, mobbing, and ensemble work that spreads knowledge, reduces bus-factor risk, and turns code review from a bottleneck into a conversation. Each of these reinforces the others. Trunk-based development demands good tests. Good tests demand shared understanding. Shared understanding comes from working together. T*D is where technical excellence and team culture meet.</strong></p><p>The good news? You don&#8217;t have to fix everything at once. Pick one practice. Implement it well. Measure the result. The evidence says that even small changes &#8212; adopting trunk-based development, speeding up code reviews, running genuine retrospectives with follow-through &#8212; create virtuous cycles that build momentum over time.</p><p>The evidence is strong, but it&#8217;s suggestive rather than prescriptive. Context matters enormously. The hard part was never knowing the practices &#8212; it&#8217;s navigating the trade-offs in your specific environment, with your specific constraints, and your specific people. The research won&#8217;t make those decisions for you. But it can tell you which directions have worked for others, and which ones haven&#8217;t. That&#8217;s worth paying attention to.</p><div><hr></div><h2><strong>References</strong></h2><ol><li><p>AWS Executive Insights. &#8220;Amazon&#8217;s Two Pizza Teams.&#8221; aws.amazon.com.</p></li><li><p>Martin Fowler. &#8220;Two Pizza Team.&#8221; martinfowler.com.</p></li><li><p>Mountain Goat Software. &#8220;The Just-Right Size for Agile Teams.&#8221; mountaingoatsoftware.com.</p></li><li><p>Robin Dunbar. &#8220;Dunbar&#8217;s Number.&#8221; Referenced via psychsafety.com and Wikipedia.</p></li><li><p>Fred Brooks. &#8220;The Mythical Man-Month: Essays on Software Engineering.&#8221; Addison-Wesley, 1975.</p></li><li><p>Alan MacCormack, John Rusnak, Carliss Baldwin. &#8220;Exploring the Duality between Product and Organizational Architectures.&#8221; Harvard Business School.</p></li><li><p>Nagappan, Murphy, Basili. &#8220;The Influence of Organizational Structure on Software Quality.&#8221; University of Maryland.</p></li><li><p>&#8220;Conway&#8217;s Law Revisited: The Evidence for a Task-Based Perspective.&#8221; IEEE Software.</p></li><li><p>Matthew Skelton, Manuel Pais. &#8220;Team Topologies: Organizing Business and Technology Teams for Fast Flow.&#8221; IT Revolution, 2019.</p></li><li><p>Nicole Forsgren, Jez Humble, Gene Kim. &#8220;Accelerate: The Science of Lean Software and DevOps.&#8221; IT Revolution, 2018.</p></li><li><p>Amy Edmondson. &#8220;Psychological Safety and Learning Behavior in Work Teams.&#8221; Administrative Science Quarterly, 44(2), 1999.</p></li><li><p>Google re:Work. &#8220;Guide: Understand team effectiveness.&#8221; rework.withgoogle.com.</p></li><li><p>DORA. &#8220;Capabilities: Generative organizational culture.&#8221; dora.dev.</p></li><li><p>Ron Westrum. &#8220;A typology of organisational cultures.&#8221; BMJ Quality &amp; Safety, 2004.</p></li><li><p>IT Revolution. &#8220;Westrum&#8217;s Organizational Model in Technology Organizations.&#8221; itrevolution.com.</p></li><li><p>Boston Consulting Group. &#8220;How Diverse Leadership Teams Boost Innovation.&#8221;</p></li><li><p>PMC/NIH. &#8220;When and how is team cognitive diversity beneficial?&#8221; pmc.ncbi.nlm.nih.gov.</p></li><li><p>ACM Digital Library. &#8220;Diversity and Teamwork in Student Software Teams.&#8221; dl.acm.org.</p></li><li><p>DORA. &#8220;DORA&#8217;s software delivery performance metrics.&#8221; dora.dev.</p></li><li><p>DORA. &#8220;Accelerate State of DevOps Report 2024.&#8221; dora.dev.</p></li><li><p>Google Cloud Blog. &#8220;Use Four Keys metrics to measure your DevOps performance.&#8221; cloud.google.com.</p></li><li><p>Forsgren, Storey, Maddila, Zimmermann, Houck, Butler. &#8220;The SPACE of Developer Productivity.&#8221; ACM Queue, Vol 19(1), 2021.</p></li><li><p>Gergely Orosz. &#8220;Measuring developer productivity? A response to McKinsey.&#8221; newsletter.pragmaticengineer.com.</p></li><li><p>McKinsey. &#8220;Developer Velocity: How software excellence fuels business performance.&#8221; mckinsey.com.</p></li><li><p>InfoQ. &#8220;Trunk Based Development as a Cornerstone for Continuous Delivery.&#8221; infoq.com.</p></li><li><p>Google Engineering Practices. &#8220;The Standard of Code Review.&#8221; google.github.io.</p></li><li><p>Google. &#8220;Code Review - Software Engineering at Google.&#8221; abseil.io.</p></li><li><p>Hannay, Dyba, Arisholm. &#8220;The effectiveness of pair programming: A meta-analysis.&#8221; Information and Software Technology, 2009.</p></li><li><p>Mike Cohn. &#8220;Succeeding with Agile: Software Development Using Scrum.&#8221; Addison-Wesley, 2009.</p></li><li><p>Google re:Work. &#8220;Guide: Set goals with OKRs.&#8221; rework.withgoogle.com.</p></li><li><p>John Doerr. &#8220;Measure What Matters.&#8221; Portfolio/Penguin, 2018.</p></li><li><p>Henrik Kniberg, Anders Ivarsson. &#8220;Scaling Agile @ Spotify.&#8221; Crisp, 2012.</p></li><li><p>Noda, Storey, Forsgren, Greiler. &#8220;DevEx: What Actually Drives Productivity.&#8221; ACM Queue, 2023.</p></li><li><p>GitLab Handbook. &#8220;How to embrace asynchronous communication for remote work.&#8221; handbook.gitlab.com.</p></li><li><p>Cortex. &#8220;Developer Onboarding Guide.&#8221; cortex.io.</p></li><li><p>Deloitte. Mentorship and retention research. Referenced via guider-ai.com.</p></li><li><p>OSTI. &#8220;An Exploration of the Mentorship Needs of Research Software Engineers.&#8221; osti.gov.</p></li><li><p>ScienceDirect. &#8220;Burnout in software engineering: A systematic mapping study.&#8221; 2022.</p></li><li><p>Ellahi, Rehman, Javed, Sultan, Rehman. &#8220;Impact of Servant Leadership on Project Success.&#8221; SAGE Open, 2022.</p></li><li><p>Spotify Engineering. &#8220;Squad Health Check model.&#8221; engineering.atspotify.com, 2014.</p></li><li><p>Patrick Lencioni. &#8220;The Five Dysfunctions of a Team.&#8221; Jossey-Bass, 2002.</p></li><li><p>ADR GitHub. &#8220;Architectural Decision Records.&#8221; adr.github.io.</p></li><li><p>Spotify Engineering. &#8220;When Should I Write an Architecture Decision Record.&#8221; 2020.</p></li><li><p>PMC. &#8220;Perceived diversity in software engineering: a systematic literature review.&#8221; 2021.</p></li><li><p>Martin Fowler. &#8220;Team Topologies.&#8221; martinfowler.com.</p></li><li><p>Inc. &#8220;Google Spent 2 Years Studying 180 Teams.&#8221; inc.com.</p></li><li><p>LeaderFactor. &#8220;Project Aristotle Psychological Safety.&#8221; leaderfactor.com.</p></li><li><p>LaunchDarkly. &#8220;Elite Performance with Trunk-based Development.&#8221; launchdarkly.com.</p></li><li><p>Dr. Michaela Greiler. &#8220;Code Reviews at Google are lightweight and fast.&#8221; michaelagreiler.com.</p></li><li><p>LeadDev. &#8220;What McKinsey got wrong about developer productivity.&#8221; leaddev.com.</p></li><li><p>GetDX. &#8220;Highlights from the 2024 DORA State of DevOps Report.&#8221; getdx.com.</p></li><li><p>ScienceDirect. &#8220;Burnout in software engineering: A systematic mapping study.&#8221; 2022.</p></li><li><p>DORA. &#8220;State of AI-assisted Software Development 2025.&#8221; dora.dev.</p></li><li><p>Kent Beck. &#8220;Measuring developer productivity? A response to McKinsey.&#8221; Substack, 2023.</p></li></ol>]]></content:encoded></item><item><title><![CDATA[Stop Counting Tickets, Start Watching the Work]]></title><description><![CDATA[A practical guide for managers who want to understand whether their developers are thriving, struggling, or somewhere in between]]></description><link>https://a4al6a.substack.com/p/stop-counting-tickets-start-watching</link><guid isPermaLink="false">https://a4al6a.substack.com/p/stop-counting-tickets-start-watching</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Sun, 15 Mar 2026 14:17:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Picture a mob programming session. A manager joins, not to check on anyone, just genuinely interested in how the team is approaching a tricky migration. But about twenty minutes in, they notice something: one of the developers is letting others drive every decision, never pushing back, never offering an alternative. He is present but passive. In a stand-up or a sprint review, he would be invisible &#8212; his tickets are getting done, his pull requests are merging. By every metric the team tracks, he is fine.</p><p><strong>He is not fine. He is stuck. And no dashboard would ever show it.</strong></p><p>[If you have lived a version of this moment &#8212; and if you have managed engineers for any length of time, you almost certainly have &#8212; then you already know what this article is about.]</p><p>That is the core problem this article is about. Not that measurement is bad &#8212; measurement is essential &#8212; but that most of what we measure in software engineering tells us almost nothing about the thing we actually need to understand: whether our people are growing, struggling, or coasting. The question is not whether to assess your developers. It is whether your assessment tools are capable of seeing what matters.</p><h2><strong>The measurement problem</strong></h2><p>Most managers, when asked how they evaluate their developers, will point to some combination of activity metrics: tickets closed, story points delivered, lines of code written, pull requests merged. These feel scientific. They feel objective. And as tools for evaluating individuals, they are actively counterproductive.</p><p>There is a name for this. Goodhart&#8217;s Law states that when a measure becomes a target, it ceases to be a good measure. In software engineering, this plays out predictably: if lines of code become a performance indicator, developers produce verbose, unnecessarily complex code. If story points become the yardstick, estimates get inflated. If pull requests merged is the metric, people split work into trivially small changes to pump the numbers. When McKinsey published a paper in 2023 claiming they could measure individual developer productivity through activity metrics, Dan North called the thinking &#8220;absurdly naive.&#8221; Kent Beck warned that such measurements create perverse incentives &#8212; managers pressure for better scores, engineers game the system, and the actual work suffers.</p><p><strong>But the answer is not to stop measuring. It is to measure the right things, at the right level, for the right purpose.</strong></p><p>Metrics like deployment frequency, lead time for changes, change failure rate, and mean time to recovery &#8212; the DORA metrics, validated by Forsgren, Humble, and Kim across 23,000 data points over four years of research &#8212; are genuinely useful. They tell you about the health of the system: how quickly value flows from idea to production, how often things break, and how fast the team recovers. What they do not tell you is which individuals are thriving and which are struggling. That requires a different kind of assessment entirely &#8212; one that is richer, more contextual, and necessarily more human.</p><p><strong>The distinction matters enormously. Metrics for understanding how a system is performing are productive. Metrics for evaluating individual people are destructive.</strong> Everything in this article sits on that line.</p><h2><strong>Watch people work</strong></h2><p>The richest source of signal about a developer is watching them work. Not in a surveillance sense, but through practices like pair programming and mob programming. When you are pairing with someone, or watching them pair with others, you learn things within a few sessions that no dashboard could ever reveal. You see how they think through problems, how they break down complexity, whether they write tests first or bolt them on as an afterthought, how they respond when their assumptions are wrong, and how they communicate with the person next to them. Academic studies on pair programming confirm it reveals individual skill levels, communication ability, and problem-solving approaches in ways no other assessment method matches.</p><p>There is a catch, and it is worth naming upfront. Harvard researcher Ethan Bernstein demonstrated through a rigorous field experiment that being observed can counterintuitively make people worse at their jobs &#8212; not because they are lazy and get caught, but because they hide the creative, experimental behaviour that produces the best work. I will return to this tension later, because it is real and it does not have a tidy resolution. But the alternative &#8212; assessing people from a distance through thin numerical proxies &#8212; is not just less effective. It is actively misleading.</p><p>Pay attention to how quickly code gets from someone&#8217;s machine to production. If a developer is sitting on branches for days, that tells you something. If their work integrates smoothly and frequently, that tells you something else entirely. The ability to work in small, safe increments is one of the most reliable markers of engineering maturity. The Accelerate research confirms this: elite teams deploy code multiple times per day, while low performers deploy once per month or less. Working in small batches reduces cycle times, accelerates feedback, reduces risk, and improves efficiency. And here is the finding that surprises many managers: speed and stability are not trade-offs. Teams that deploy more frequently also have lower failure rates and faster recovery times.</p><p>One caveat matters here. A developer sitting on branches for days might not be signalling personal weakness. The CI/CD pipeline might be broken. Code review might be bottlenecked. The team might lack trunk-based development practices. <strong>Before judging the individual, check the system.</strong> This is a principle that runs through the entire article: whenever you see individual behaviour that looks like a problem, ask whether the environment is causing it before concluding that the person is.</p><p>Watch how people behave when something breaks. Good developers do not panic and do not blame. They have a systematic approach to diagnosis. They narrow things down. They ask good questions. And crucially, after they fix the problem, they ask &#8220;how do we make sure this class of problem cannot happen again?&#8221; Google&#8217;s Project Aristotle &#8212; a two-year study of over 180 teams &#8212; found that psychological safety was the single strongest predictor of team performance, accounting for 43 percent of the variance. Teams where people blame and panic are teams where psychological safety is low. Teams where people diagnose calmly and think about prevention are teams where it is high.</p><p>Then there is the question of learning. Good developers are visibly curious. They read, they experiment, they ask &#8220;why&#8221; a lot. But here is the nuance: you are not looking for people who chase every shiny new framework. You are looking for people who deepen their understanding of fundamentals. Someone who properly understands testing, design, and feedback loops will pick up any new tool in a week. Someone who only knows tools but not principles will struggle every time the landscape shifts. Daniel Pink&#8217;s research on motivation calls this mastery &#8212; the desire to continually improve at something that matters. It is one of three core intrinsic motivators for knowledge work, alongside autonomy and purpose. The developers who chase depth are driven by mastery. The ones who chase breadth without depth are often driven by anxiety.</p><p>That said, this is not a binary. In a fast-moving field, some exploration of new tools is adaptive and necessary. The signal is not whether someone looks at new things, but whether they can evaluate them critically against fundamentals. Curiosity about new tools combined with deep understanding of principles is the strongest combination.</p><h2><strong>Practical actions for managers</strong></h2><p>Let us get concrete about what a manager can actually do day to day to understand their team, not in a way that punishes or controls, but in a way that genuinely helps.</p><p>Sit with your teams regularly. Not in a &#8220;let me observe you&#8221; way, but genuinely participate. Join a mob session. Pair with someone on a real task. Even if you are not coding yourself, being in the room while work happens gives you an enormous amount of information. You will see who drives the design conversation, who asks clarifying questions, who just waits to be told what to do, and who challenges assumptions constructively. Do this often enough and you will build a mental model of every person on the team that no spreadsheet could ever give you.</p><p>Pay attention to what happens after someone&#8217;s code merges. Does it cause problems downstream? Do other people frequently have to rework it? Or does it land cleanly and become something the team builds on with confidence? You do not need a fancy tool for this. Just ask the team during retrospectives: &#8220;What slowed us down this week?&#8221; If the same person&#8217;s work keeps coming up as a source of friction, that is a signal. If someone&#8217;s work is invisible because it just works, that is also a signal, and you should notice it.</p><p>Give people a small, ambiguous problem and see what they do with it. Not as a test, but as real work. Something like &#8220;we have had three customer complaints about this area of the product, can you investigate and come back with a recommendation?&#8221; A strong developer will dig in, ask smart questions, maybe prototype something, and come back with options. Someone less experienced, or someone who has not been given the space to develop this skill, will either freeze, ask you to define every detail upfront, or jump straight to coding without understanding the problem. This tells you far more than any technical interview ever could, because it is real. Will Larson&#8217;s research on staff engineering identifies exactly this &#8212; the capacity to self-direct through uncertainty rather than waiting for someone to remove it &#8212; as the defining characteristic that separates mid-level from senior engineers, and senior from staff.</p><p>Look at how people handle knowledge gaps. Give someone a task that is slightly outside their comfort zone and watch what happens. Do they go and learn what they need? Do they ask for help at the right moment, not too early and not after they have wasted three days? Do they share what they learned afterwards? The ability to self-direct learning and then bring it back to the team is one of the strongest indicators of someone who will grow into the role versus someone who has plateaued.</p><p>Have honest one-to-ones where you ask open questions and then actually listen. Not &#8220;how are your tickets going&#8221; but things like &#8220;what is the hardest technical decision you have made recently and why?&#8221; or &#8220;what part of our codebase worries you most?&#8221; or &#8220;if you could change one thing about how we work, what would it be?&#8221; The quality of the answers tells you a lot. Someone who engages thoughtfully with these questions is someone who cares about the craft. Someone who consistently has nothing to say may not be thinking deeply about the work &#8212; or they may not yet trust that it is safe to say what they think. Some people are internal processors. Some come from cultures where questioning how the team works is not done casually. The manager&#8217;s job is to figure out which it is, not to assume the worst. If someone goes quiet in one-to-ones, the first question should be &#8220;have I made this space safe enough?&#8221; not &#8220;does this person care?&#8221;</p><p>And get feedback from peers. Not formal 360 reviews with scores and ratings used for appraisal &#8212; research shows those tend to produce manipulated data, with gaming documented at companies including GE, IBM, and Amazon. But do not dismiss structured peer feedback entirely. When 360-degree feedback is used for development rather than evaluation, and when responses are anonymous, research shows it genuinely increases communication, productivity, and team effectiveness. The best approach combines informal peer conversations &#8212; &#8220;What is it like working with Sarah?&#8221; &#8212; with structured but anonymous developmental feedback that lets people say things they would not say to someone&#8217;s face. Some of the most important signals, especially upward feedback about a manager&#8217;s own behaviour, only surface when anonymity is guaranteed.</p><h2><strong>Being honest about scale</strong></h2><p>Everything I have described so far sounds reasonable for a manager with five or six direct reports. But what if you have twelve? What if you manage two teams across different time zones? The honest answer is that you cannot pair with every engineer every week, attend every mob session, and have deep one-to-ones with everyone on a fortnightly cycle. If this approach only works at small scale, it is not a complete approach.</p><p>There are three things that make it practical at scale.</p><p>First, you do not need to be the only observer. Tech leads, senior engineers, and other experienced team members are already watching the work. They see how people pair, how they handle ambiguity, how they respond to feedback. Your job is not to personally observe everything but to build a network of people whose judgement you trust and to triangulate their perspectives with your own. Ask your tech lead: &#8220;How is Maria doing on the migration? What have you noticed?&#8221; <strong>This is not delegation of responsibility. It is distributed observation</strong>, and it produces a richer, less biased picture than a single manager&#8217;s viewpoint ever could.</p><p>Second, you can rotate your depth of focus. Rather than shallowly observing everyone all the time, spend a few weeks deeply engaged with one part of the team &#8212; joining their sessions, reading their pull requests, having longer conversations &#8212; then rotate. Over a quarter, you will have built a detailed picture of everyone. This is not ideal, but it is honest, and it is better than the alternative of seeing no one deeply.</p><p>Third, mob programming is more efficient for observation than pairing, precisely because you see the whole team&#8217;s dynamics at once. A single mob session reveals who drives, who supports, who withdraws, who challenges, and who defers. You do not need to attend many sessions to learn a lot. If you are time-constrained, mobs give you the highest signal per hour invested.</p><h2><strong>The honest tension: observation is still observation</strong></h2><p>There is a tension here that deserves honesty rather than a neat resolution. If you are joining a session partly to learn about the work and partly to assess the people doing it, then dressing it up as pure curiosity is a form of dishonesty. People are not stupid. They can feel when they are being evaluated, regardless of what words you use. And if they later discover that your &#8220;just joining in&#8221; sessions fed into a performance conversation, the trust damage is significant.</p><p>This tension is not just intuition. It has academic backing. Bernstein&#8217;s &#8220;Transparency Paradox&#8221; research, which I mentioned earlier, demonstrated this through a field experiment at a large manufacturing plant. Workers who knew they were being watched concealed innovative approaches, suppressed experimentation, and avoided productive deviation from standard practice. Bernstein found that even a modest increase in group-level privacy sustainably and significantly improved performance. The implication is uncomfortable: the very act of a manager watching can make people worse at their jobs, because they hide the creative, experimental behaviour that produces the best work.</p><p>So the better approach is transparency. Not &#8220;I am here to check on you,&#8221; but something closer to the truth: &#8220;Part of my job is understanding how the team works, what is going well, and where people might need support. I cannot do that if I am never near the work. So I am going to join sessions from time to time. Not to catch anyone out, but because I need to see reality if I am going to be any use to you.&#8221; That is honest. It acknowledges the evaluative dimension without making it adversarial.</p><p>But here is the deeper point, and Bernstein&#8217;s research supports it. If the only time a manager is near the work is when they are assessing people, then of course it feels like surveillance, no matter how it is framed. The real fix is not about how you explain your presence. It is about whether your presence is normal or exceptional. If a manager is routinely involved in technical discussions, regularly joins sessions, frequently asks about design decisions, and has ongoing conversations about the codebase, then there is no single moment that feels like &#8220;the inspection.&#8221; It is just how things work. The observation happens as a side effect of genuine involvement, not as a discrete activity with a hidden agenda.</p><p>The managers who struggle most with this are the ones who are distant from the work ninety percent of the time and then suddenly show up. In that context, no amount of friendly framing will stop people from feeling watched. <strong>The discomfort is not caused by the words. It is caused by the pattern.</strong></p><p>There is also something worth saying about the broader culture. In teams where feedback is frequent and open, where people regularly talk about what is going well and what is not, evaluation loses its menacing quality. People feel micromanaged when assessment is something that happens to them in secret and then gets revealed in a formal review. If instead the manager is consistently saying &#8220;I noticed this went well&#8221; and &#8220;I think this area needs work&#8221; as part of normal conversation, then the evaluative aspect of being present becomes something people are used to rather than something they dread.</p><p>None of this fully resolves the tension. A manager who is present will, by definition, be forming judgements. People who are being observed will, by definition, behave somewhat differently. You cannot eliminate that dynamic entirely. But you can make it honest, make it routine, and make it part of a relationship where people trust that the judgements being formed are fair and in their interest. That is the best you can do, and it is a lot better than the alternative, which is forming judgements from a distance based on terrible data and then surprising people with them twice a year.</p><h2><strong>The biases you carry into the room</strong></h2><p>If observation is your primary assessment tool &#8212; and I am arguing it should be &#8212; then you need to be honest about the fact that you are not an objective instrument. Decades of research on performance evaluation show that managers are systematically biased in ways they are rarely aware of.</p><p>Recency bias means you overweight what happened in the last few weeks and forget the previous months. SHRM research shows that recency and central tendency errors affect nearly 40 percent of annual appraisals, and they do not disappear just because you are observing in person rather than reading a spreadsheet. That developer who had a rough week when you happened to join the mob session? You will remember that disproportionately.</p><p>Similarity bias means you unconsciously favour people who remind you of yourself &#8212; same background, same communication style, same approach to problem-solving. The developer who thinks the way you do will seem &#8220;stronger&#8221; than the one who thinks differently but equally well.</p><p>The halo and horns effects mean that one strong or weak impression colours everything else. If someone impressed you with an excellent design decision, you will be more generous when evaluating their testing practices, even if those are mediocre. If someone fumbled an incident response, you may underrate their day-to-day coding, even if it is excellent.</p><p>And then there is proximity bias, which matters more than ever. Managers form more favourable impressions of people they see frequently. In hybrid or remote teams, this creates systematic unfairness: research shows employees may receive better review outcomes simply because they work in the same office as their manager. The person you pair with every week will feel more &#8220;known&#8221; &#8212; and therefore rated more positively &#8212; than the remote team member you interact with only in stand-ups.</p><p>One of the largest studies on feedback found that more than half of the variance in performance ratings had more to do with the quirks of the person giving the rating than the person being rated. <strong>That means the biggest variable in your evaluation is you, not them.</strong></p><p>What do you do about this? You cannot eliminate bias, but you can discipline your observation. Keep running notes. Not a surveillance dossier, but a simple habit: after a pairing session or a notable interaction, write down what you actually saw, not what you felt about it. Over time, patterns emerge from evidence rather than from impressions. Seek disconfirming evidence actively &#8212; if you think someone is weak, look specifically for moments where they are strong, and vice versa. Calibrate with peers: ask other managers or tech leads who work with the same people whether they see what you see. And be especially deliberate about equalising your attention across the team. If you are pairing more often with some people than others, you are building a biased dataset whether you mean to or not.</p><h2><strong>When the team is not in the room</strong></h2><p>Everything I have described so far assumes you can physically or synchronously join a session with your team. But many teams are distributed. Some are fully remote. Some span time zones. If observation-based assessment only works when you can sit next to someone, then it is not a complete approach.</p><p>The principles remain the same, but the channels change. In distributed teams, the work leaves more written traces, and those traces become your primary observation material.</p><p>Code review is the most revealing. Not reviewing code yourself as a gatekeeper, but reading how people engage with each other&#8217;s code. The difference between a thoughtful reviewer and a careless one is immediately visible. Compare &#8220;this is wrong, fix it&#8221; with &#8220;I think this approach might cause issues under concurrency &#8212; have you considered using a lock here? Happy to pair on it if useful.&#8221; The first tells you someone is going through the motions. The second tells you someone understands the system, communicates with care, and is willing to invest time in a colleague&#8217;s growth. How people respond to critique is equally telling: do they engage with the substance, or do they get defensive? Do they ask follow-up questions, or do they silently apply the change without understanding why? The quality of someone&#8217;s code review comments tells you as much about their engineering judgement as watching them code would.</p><p>Look at how people communicate in writing more broadly. In asynchronous teams, the ability to write a clear problem statement, a well-structured RFC, or a concise incident summary is itself a form of engineering skill. People who can articulate their thinking in writing are usually people who think clearly, and that matters regardless of whether you are ever in the same room.</p><p>Pay attention to how distributed team members handle the particular challenges of remote work. Do they proactively communicate blockers or sit silently for days? Do they make themselves available for synchronous collaboration when it matters, or are they always unavailable? Do they contribute to team discussions, or do they disappear between assigned tasks?</p><p>None of this is as rich as pairing with someone in real time. But it is far better than falling back on ticket counts and activity dashboards, which is what most managers of remote teams end up doing by default.</p><h2><strong>Promotions, recognition, and the evidence problem</strong></h2><p>This is where it gets uncomfortable, because most organisations pretend they have a system for promotions when really they have a ritual. Someone fills in a form, a manager writes a justification, a calibration meeting happens where people who have never seen the work argue about ratings, and then a decision gets made that is mostly political. Everyone involved knows it is theatre, but nobody says it out loud.</p><p>The data you need is not data in the traditional sense. It is evidence. And the best evidence comes from the practices already described, but you need to be deliberate about collecting it. Not in a creepy dossier way, but as a habit of noticing and writing things down. When someone handles an incident well, make a note. When someone&#8217;s pairing session lifts the whole team&#8217;s understanding, write it down. When someone repeatedly delivers work that needs reworking, note that too. Over time you build a picture that is grounded in real events, not vibes and not metrics.</p><p>For promotions specifically, the question should not be &#8220;has this person earned a reward?&#8221; It should be &#8220;is this person already operating at the next level?&#8221; That distinction matters enormously. If someone is already doing the work of a senior engineer, meaning they are influencing design decisions, mentoring others, taking ownership of ambiguous problems, thinking about the system rather than just their task, then promoting them is just recognising reality. If you are promoting someone because they have been around long enough or because they will leave otherwise, you are creating problems.</p><p>The evidence for this is observable. You can point to specific moments. &#8220;In the last six months, you led the redesign of the payment service. You brought three junior developers along with you through pairing. You identified the performance issue before it hit production. You pushed back on the product team when the requirements did not make sense and proposed a better alternative.&#8221; That is a promotion case built on things that actually happened, not on a self-assessment form where someone writes &#8220;I demonstrated leadership&#8221; with no context.</p><p>But there are two fairness problems here that deserve honesty.</p><p>The first is access. The &#8220;already operating at the next level&#8221; criterion only works if everyone has equal access to the opportunities that let them demonstrate next-level work. In practice, they often do not. People from underrepresented groups, quieter team members, people in less visible parts of the codebase, and remote workers may all have fewer chances to lead a high-profile redesign or push back on product in a visible way. Research on equitable promotion policies confirms that marginalised groups are promoted at a slower rate, partly because the opportunities to demonstrate next-level capability are unequally distributed.</p><p>This means managers have an active responsibility, not just to observe who steps up, but to deliberately distribute stretch opportunities. <strong>If you only promote people who naturally volunteer for visible work, you are filtering for confidence and political skill, not engineering ability.</strong> Make sure the quiet person on the team gets the chance to lead something. Make sure the remote team member gets the same ambiguous problems as the person sitting next to you.</p><p>The second is compensation. &#8220;Already operating at the next level&#8221; means, in practice, doing a harder job for months while being paid for the easier one. Most companies expect six to twelve months of sustained performance at the next level before promoting. That is six to twelve months of uncompensated labour at a higher level of responsibility. This might sound reasonable in the abstract, but it disproportionately affects people who cannot afford to &#8220;invest&#8221; unpaid effort &#8212; and it is the single most common criticism of this promotion model among the engineers who live inside it.</p><p>I do not have a clean solution for this. But I think the manager&#8217;s obligation is clear: make the gap between &#8220;doing the work&#8221; and &#8220;getting the title&#8221; as short as organisationally possible. <strong>If someone has been operating at the next level for two full review cycles and you still have not promoted them, that is a management failure, not a demonstration of rigour.</strong> And be transparent about the timeline. If someone is doing senior-level work, tell them: &#8220;I see it, I am building the case, and here is when I expect it to happen.&#8221; <strong>Silence on this topic is how you lose your best people.</strong></p><h2><strong>Feedback that actually changes behaviour</strong></h2><p>For feedback conversations, the same principle applies. Be specific and be timely. &#8220;Your code in the checkout service last week had no tests and broke the build twice&#8221; is useful feedback. &#8220;You need to improve your quality&#8221; is not. The first gives someone something to act on. The second just makes them feel bad.</p><p>Here is something most managers get wrong: feedback should not be saved up for a quarterly review. If you see something that needs addressing, address it within days, not months. If you see something excellent, say so immediately. The idea that feedback is a formal event is one of the most damaging conventions in management. It means people spend months not knowing where they stand, which breeds anxiety and kills trust.</p><p>The research on this is overwhelming. Gallup found that employees who receive feedback weekly are 2.7 times more likely to be engaged at work. More strikingly, employees are 3.6 times more likely to be motivated to do outstanding work when their manager provides daily feedback compared to annual feedback. Companies with strong feedback cultures see 14.9 percent lower turnover. When Adobe shifted from annual reviews to continuous check-ins, they saw a 30 percent drop in voluntary turnover within a single year.</p><p>If you are giving feedback regularly, the formal review becomes a summary of things both of you already know, which is exactly what it should be. <strong>Nothing in a performance review should ever be a surprise.</strong> If it is, you have failed as a manager, not because the assessment is wrong, but because you waited too long to share it.</p><p>On recognition, be careful about making it purely individual. Software is a team activity. If you only recognise individual heroes, you incentivise hero behaviour, and that is corrosive. Hero culture leads to knowledge silos, bus factor risks, and burnout among the very people being celebrated. Perhaps most damningly, research suggests that hero culture is itself a sign of low psychological safety &#8212; when only heroes are recognised, others stop taking initiative because the implicit message is that only extraordinary individual acts matter.</p><p>Google&#8217;s Project Aristotle found that the number one predictor of team performance was not individual talent but psychological safety. Teams with high psychological safety showed 19 percent higher productivity, 31 percent more innovation, and 27 percent lower turnover. Individual heroics were not on the list.</p><p>Recognise teams. Recognise collaborative moments. &#8220;The way you and Tom worked through that production issue together was excellent&#8221; reinforces the behaviour you actually want. Individual recognition has its place, but it should be tied to team-enabling behaviours &#8212; mentoring, unblocking others, sharing knowledge &#8212; not solo heroics.</p><p>As for what to base the conversation on practically, keep it simple. For each person, maintain a running list of observations under three headings: things they are doing well, things they need to work on, and situations where you were not sure what to make of their contribution. Review your notes before any one-to-one. Share them openly. Ask the person if they see it the same way. The best feedback conversations are ones where the person mostly agrees with your assessment because nothing in it is a surprise.</p><h2><strong>The environment is the assessment</strong></h2><p>On the &#8220;did I hire the right people&#8221; question, it is worth reframing it entirely. The better question is &#8220;have I created an environment where good people can do good work and where struggling people can improve?&#8221; If you have strong collaborative practices in place, meaning pairing, mobbing, frequent code review as a learning exercise rather than a gate, continuous integration, and short feedback loops, then you will find that most people rise to the level of the environment. The ones who genuinely cannot are not usually a mystery. They become visible quite quickly when the work is transparent.</p><p>This is not a new idea. W. Edwards Deming, whose work on quality management became foundational to both Agile and DevOps, argued decades ago that the system, as designed by leaders, is almost always to blame for problems &#8212; not the individual working within the system. Deming saw the manager&#8217;s role as improving the system in which people work, not judging the people and leaving the system untouched. His insight that &#8220;mistakes typically come from bad systems, not bad workers&#8221; has been directly supported by the Accelerate research, which identified 24 organisational capabilities &#8212; not individual traits &#8212; as the drivers of software delivery performance. </p><p>Pink&#8217;s motivation framework reinforces the point from the individual&#8217;s perspective. Metrics-based evaluation systems undermine autonomy by telling people what to optimise for, distort mastery by rewarding the wrong skills, and corrode purpose by making people focus on numbers rather than outcomes. When you build an environment that gives developers genuine autonomy over how they work, supports their drive for mastery through pairing and challenging problems, and connects their work to a purpose they care about, you create the conditions where good people do their best work &#8212; and where struggling people either rise to the challenge or become honestly visible.</p><p><strong>The real danger is not hiring the wrong person. It is creating an environment where you cannot tell the difference between a good developer in a bad system and a bad developer in a good system.</strong> Make the system transparent and the assessment mostly takes care of itself.</p><p>That said, environment-first thinking does not mean never addressing individual performance. Some people genuinely underperform even in excellent environments. The point is the order of operations: fix the system first, then assess the individual. If someone is still struggling after you have given them good practices, clear expectations, psychological safety, timely feedback, and opportunities to grow, then you have a genuine individual performance problem &#8212; and you have the evidence to address it fairly, because you have been watching the work all along.</p><p>Here is what that conversation sounds like when you have done the work: &#8220;Over the last three months, I have noticed a pattern. When you paired with Laura on the billing service, she drove all the design decisions and you did not push back on any of them. The migration task I gave you came back without tests, and the team spent two days fixing the issues it caused downstream. In our last three one-to-ones, I asked what you would change about how we work and you said you were not sure. I have given you feedback on each of these as they happened, and I have not seen a change. I want to help you get to where you need to be, but I need to see a shift in the next few weeks, and here is specifically what that looks like.&#8221; That is a hard conversation, but it is not an unfair one. Nothing in it is a surprise. Nothing in it is a vibe. Every claim points to something that actually happened, and the person has already heard about each of those moments because you addressed them at the time. <strong>That is the payoff of proximity: when the difficult conversation finally has to happen, it is grounded in evidence both of you recognise.</strong></p><p>The underlying principle running through all of this is proximity. You cannot evaluate people from a distance. Dashboards, velocity charts, and ticket counts are all ways of trying to understand people without actually being close to the work, and they consistently fail. The managers who truly understand their teams are the ones who stay close to the work itself. Not micromanaging, not controlling, just present and paying attention &#8212; while remaining honest about their own biases, deliberate about distributing their attention fairly, and disciplined about grounding their judgements in evidence rather than impressions.</p><div><hr></div><h2><strong>References</strong></h2><p><strong>Books</strong></p><ul><li><p>Forsgren, N., Humble, J., Kim, G. (2018). <em>Accelerate: The Science of Lean Software and DevOps</em>. IT Revolution Press.</p></li><li><p>Pink, D. (2009). <em>Drive: The Surprising Truth About What Motivates Us</em>. Riverhead Books.</p></li><li><p>Larson, W. (2021). <em>Staff Engineer: Leadership Beyond the Management Track</em>.</p></li><li><p>Deming, W.E. <em>Out of the Crisis</em>. MIT Press.</p></li></ul><p><strong>Academic and peer-reviewed research</strong></p><ul><li><p>Bernstein, E. (2012). &#8220;The Transparency Paradox: A Role for Privacy in Organizational Learning and Operational Control.&#8221; <em>Administrative Science Quarterly</em>, 57(2), 181&#8211;216.</p></li><li><p>Forsgren, N., Storey, M-A., Maddila, C., Zimmermann, T., Houck, B., Butler, J. (2021). &#8220;The SPACE of Developer Productivity: There&#8217;s More to It Than You Think.&#8221; <em>ACM Queue</em>, 19(1).</p></li><li><p>Google re:Work. &#8220;Guide: Understand Team Effectiveness&#8221; (Project Aristotle).</p></li><li><p>Hawthorne Effect research (1924&#8211;1932). Western Electric / Elton Mayo.</p></li></ul><p><strong>Industry sources</strong></p><ul><li><p>Orosz, G. (2023). &#8220;Measuring Developer Productivity? A Response to McKinsey.&#8221; <em>The Pragmatic Engineer</em>.</p></li><li><p>Orosz, G. &#8220;Common Performance Review Biases.&#8221; <em>The Pragmatic Engineer</em>.</p></li><li><p>Orosz, G. &#8220;Software Developer Promotions: Advice to Get to That Next Level.&#8221; <em>The Pragmatic Engineer</em>.</p></li><li><p>Swarmia (2025). &#8220;Engineering Metrics Leaders Should Track.&#8221;</p></li><li><p>DORA. &#8220;DORA&#8217;s Software Delivery Performance Metrics.&#8221; dora.dev.</p></li></ul><p><strong>Organisational research</strong></p><ul><li><p>Gallup (2020). Employee engagement and feedback frequency research.</p></li><li><p>SHRM. Research on recency and central tendency errors in annual appraisals.</p></li><li><p>ISACA (2024). &#8220;Examining the Risks of IT Hero Culture.&#8221;</p></li><li><p>Adobe Systems. Continuous check-in programme results.</p></li></ul><p><strong>Behavioural science</strong></p><ul><li><p>Goodhart, C. &#8220;Goodhart&#8217;s Law.&#8221; Popularised as: &#8220;When a measure becomes a target, it ceases to be a good measure.&#8221;</p></li><li><p>Harvard Business Review (2018). &#8220;3 Biases That Hijack Performance Reviews, and How to Address Them.&#8221;</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Developer Productivity Trap]]></title><description><![CDATA[Why Everything We Think We Know About Developer Productivity Is Wrong &#8212; And Why AI Is Making It Worse]]></description><link>https://a4al6a.substack.com/p/the-developer-productivity-trap</link><guid isPermaLink="false">https://a4al6a.substack.com/p/the-developer-productivity-trap</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Sat, 14 Mar 2026 20:59:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3><strong>Abstract</strong></h3><p>For decades, the software industry has been trying to measure developer productivity the way factories measure widget output &#8212; and failing spectacularly. From lines of code to story points, from velocity charts to McKinsey frameworks, every attempt to reduce the complex, creative act of software development to a number has produced the same result: perverse incentives, gaming, and surrogation &#8212; the nonconscious process by which people forget what the metric was supposed to represent and treat the number itself as the goal. Modern frameworks like DORA and SPACE represent genuine progress, yet they too are routinely misapplied. Now, AI coding assistants have entered the picture, promising to make every developer a &#8220;10x engineer.&#8221; But the research tells a far more nuanced and troubling story: AI amplifies whatever is already there &#8212; good practices and bad metrics alike. This article traces the history of productivity measurement in software, dismantles the most persistent myths, examines what the evidence actually says about AI-assisted development, and offers leadership a fundamentally different way to think about the question. The answer, it turns out, is not a better metric. It is a better question.</p><div><hr></div><h2><strong>1. Minus Two Thousand Lines of Code</strong></h2><p>In 1982, Bill Atkinson was rewriting QuickDraw&#8217;s region calculation engine for the Apple Lisa. He replaced a complex algorithm with a simpler, more general one that ran approximately six times faster. The new code was elegant, compact, and superior in every measurable way. It also happened to be 2,000 lines shorter than what it replaced.</p><p>That week, Apple&#8217;s managers required every engineer to submit a form reporting how many lines of code they had written. Atkinson wrote <strong>-2000</strong> on his form.</p><p>Management discontinued the reporting shortly after.</p><p>Atkinson believed that &#8220;lines of code was a silly measure of software productivity&#8221; and that such metrics &#8220;only encouraged writing sloppy, bloated, broken code.&#8221; He was right &#8212; and yet, more than four decades later, the software industry is still searching for the number that captures what a developer is worth. We are still, in essence, asking engineers to fill out the same form. We have just made the form more sophisticated.</p><p><strong>The pursuit itself is the problem.</strong> And now, with AI generating code at unprecedented speed, we are not getting closer to an answer. We are getting further away, faster.</p><div><hr></div><h2><strong>2. The Graveyard of Bad Metrics</strong></h2><p>Every era of software engineering has produced its own favourite way to measure developer productivity. Every one of them has failed. Understanding <em>why</em> they failed is more instructive than understanding what they measured.</p><h3><strong>Lines of Code</strong></h3><p>The most intuitive metric is also the most thoroughly discredited. Bill Gates reportedly said that &#8220;measuring software productivity by lines of code is like measuring progress on an airplane by how much it weighs.&#8221; The problems are well-documented: a task requiring 100 lines in C++ might take 10 in Python, making cross-project comparisons meaningless. Refactoring, deduplication, and simplification &#8212; the hallmarks of good engineering &#8212; all <em>reduce</em> line count. <strong>Under an LOC regime, the best work registers as negative productivity.</strong></p><p>Martin Fowler put it simply: well-designed code is shorter because it eliminates duplication. Copy-paste programming inflates LOC while degrading design. LOC indicates system size, not value created.</p><h3><strong>Story Points and Velocity</strong></h3><p>Story points and velocity were designed as <em>planning tools</em> &#8212; ways for teams to forecast how much work they could take on in a sprint. They were never meant to measure productivity. But in organization after organization, they have been weaponized for exactly that purpose.</p><p>The problems are predictable. When velocity becomes a target, teams inflate their estimates to appear more productive. Individual teams can easily game the system, and when some members start inflating story points, others notice. From there, it is a short path to broken trust, frustration, and culture damage. Story points focus on effort, not value. A team can complete 50 points of features that nobody uses and score higher than a team that delivers 20 points of work that transforms the business.</p><h3><strong>Hours Worked</strong></h3><p>In knowledge work, more hours frequently produce worse outcomes. Context switching after interruptions costs approximately 23 minutes of refocus time per interruption. Attention residue from task switching can reduce cognitive capacity for 10 to 30 minutes. Cal Newport&#8217;s research on &#8220;deep work&#8221; demonstrates that sustained focus periods of 90 or more minutes are necessary for complex problem-solving. <strong>Measuring hours measures presence, not thought.</strong> And in an industry where the most valuable breakthroughs often happen during a walk or in the shower, presence is a particularly poor proxy for contribution.</p><h3><strong>Three Concepts That Explain Everything</strong></h3><p>Three concepts explain why every simple metric fails &#8212; and will always fail.</p><p><strong>The McNamara Fallacy</strong>, named for U.S. Secretary of Defense Robert McNamara, describes a four-step descent into delusion: First, measure whatever can be easily measured. Second, disregard what cannot be easily measured. Third, presume that what cannot be measured is not important. Fourth, declare that what cannot be measured does not exist.</p><p>During the Vietnam War, McNamara used enemy body counts as the primary measure of success. When a general suggested adding a factor for the sentiments of the Vietnamese people, McNamara erased it from the report &#8212; he could not quantify it. The war was lost despite the metrics showing consistent progress. The software industry repeats this pattern with striking fidelity: we measure commits, pull requests, and deployment frequency while ignoring design quality, knowledge sharing, and whether anyone actually uses what was built.</p><p><strong>Goodhart&#8217;s Law</strong> states: &#8220;When a measure becomes a target, it ceases to be a good measure.&#8221; In software engineering, this manifests everywhere: developers rush deployments to meet frequency targets, producing unstable code; teams close easy tickets to inflate resolution rates; engineers write verbose code to increase line counts; managers ship unnecessary features to hit delivery targets. <strong>The metric improves. The system degrades.</strong></p><p><strong>Surrogation</strong> is the most insidious of the three, because unlike gaming, it is invisible to the person it affects. Formally defined by Choi, Hecht, and Tayler in a 2012 paper in <em>The Accounting Review</em>, surrogation describes the nonconscious process by which people stop treating a metric as a proxy for a goal and start treating the metric <em>as</em> the goal. The metric does not just become a target &#8212; it replaces the thing it was supposed to represent in people&#8217;s minds.</p><p>The mechanism is what psychologists call attribute substitution: when the thing you actually care about (developer productivity, code quality, customer satisfaction) is abstract and hard to observe, and the metric (velocity, deployment frequency, NPS score) is concrete and immediately available, your brain quietly swaps one for the other. You do not notice the substitution. You genuinely believe you are pursuing the original goal when you are, in fact, optimising a number.</p><p>A person affected by Goodhart&#8217;s Law may know they are gaming. <strong>A person affected by surrogation does not know they are substituting.</strong> And critically, research by Black, Meservy, Tayler, and Williams (2021) demonstrated that you do not even need to tie compensation to a metric for surrogation to occur. Simply tracking and reporting a number is enough. The mere existence of the metric triggers the substitution.</p><p>Harris and Tayler, writing in <em>Harvard Business Review</em> in 2019, identified three conditions that create surrogation risk: the strategic objective is abstract, the metric is concrete and visible, and people accept the measure as representing the goal. Software engineering meets all three conditions for virtually every metric it uses.</p><p>These three concepts are not bugs in specific metrics. They are features of the relationship between measurement and human cognition. The McNamara Fallacy explains why we ignore what we cannot count. Goodhart&#8217;s Law explains why people game what we do count. Surrogation explains why we forget what we were counting in the first place. <strong>Any single metric used for evaluation will eventually be gamed, surrogated, or both.</strong> This is not cynicism &#8212; it is a predictable consequence of how human minds interact with numbers.</p><div><hr></div><h2><strong>3. The Modern Frameworks: Progress and Pitfalls</strong></h2><p>The good news is that the industry has produced genuinely better thinking about developer productivity. The bad news is that better thinking is routinely applied in the same old ways.</p><h3><strong>DORA Metrics</strong></h3><p>The four DORA metrics &#8212; Deployment Frequency, Lead Time for Changes, Mean Time to Recovery, and Change Failure Rate &#8212; were introduced by Dr. Nicole Forsgren, Jez Humble, and Gene Kim in <em>Accelerate</em> (2018). They represent a significant leap forward because they measure <em>outcomes</em> (how effectively software reaches users) rather than <em>outputs</em> (how much stuff developers produce).</p><p>But DORA&#8217;s creators explicitly warn against using these metrics for team-by-team comparison. Yet that is precisely what happens: leadership teams benchmark teams on them, comparing deployment frequency between a mobile app team and a web service team as if the numbers were comparable. DORA metrics do not capture developer satisfaction, cognitive load, code quality, technical debt, or business value. <strong>A team can score &#8220;elite&#8221; on all four metrics while building a product that nobody wants.</strong></p><p>The 2025 DORA Report acknowledged these limitations by abandoning the low/medium/high/elite performance tiers entirely, replacing them with seven team archetypes &#8212; from &#8220;harmonious high-achievers&#8221; to &#8220;legacy bottleneck&#8221; teams &#8212; reflecting a more nuanced understanding that performance is contextual and multi-dimensional.</p><h3><strong>SPACE Framework</strong></h3><p>SPACE, developed in 2021 by Nicole Forsgren, Margaret-Anne Storey, and colleagues at Microsoft Research, explicitly addresses the multi-dimensional nature of productivity across five dimensions: Satisfaction and well-being, Performance, Activity, Communication and collaboration, and Efficiency and flow. Its key insight is powerful: &#8220;Productivity cannot be reduced to a single dimension (or metric!). Only by examining a constellation of metrics in tension can we understand and influence developer productivity.&#8221;</p><p>Forsgren herself warned against the most common misuse: &#8220;One of the most common myths &#8212; and potentially most threatening to developer happiness &#8212; is the notion that productivity is all about developer activity, things like lines of code or number of commits. More activity can appear for various reasons: working longer hours may signal developers having to &#8216;brute-force&#8217; work to overcome bad systems or poor planning.&#8221;</p><h3><strong>Developer Experience as a Lens</strong></h3><p>A more recent strand of thinking shifts the focus away from what developers produce and toward what developers <em>experience</em>. The core insight is that three dimensions shape how effectively developers can work: feedback loops (how quickly they get information about their work), cognitive load (how much mental effort their environment demands), and flow state (how often they can achieve sustained, focused work). <strong>When these conditions improve, outcomes follow. When they degrade, no amount of measurement or exhortation will compensate.</strong></p><h3><strong>The McKinsey Debacle</strong></h3><p>In August 2023, McKinsey published &#8220;Yes, you can measure software developer productivity,&#8221; claiming their framework was already in use at nearly 20 companies. The response from the engineering community was swift and devastating &#8212; the strongest collective rebuttal in recent memory.</p><p>Kent Beck, the creator of Extreme Programming, called the report &#8220;so absurd and naive that it makes no sense to critique it in detail.&#8221; But then he critiqued it in detail, because &#8220;what they published damages people I care about. I&#8217;m here to help geeks feel safe in the world. This kind of surveillance makes geeks feel less safe.&#8221;</p><p>Dave Farley, co-author of <em>Continuous Delivery</em>, was blunt: &#8220;Apart from the use of DORA metrics in this model, the rest is pretty much astrology.&#8221; He drew a sharp line between DORA&#8217;s evidence-based approach and the rest of McKinsey&#8217;s framework &#8212; &#8220;the difference between astronomy and astrology.&#8221;</p><p>Beck and Gergely Orosz, in a detailed joint response, argued that the McKinsey framework &#8220;only measures effort or output, not outcomes and impact, which misses half of the software developer lifecycle.&#8221; They warned that &#8220;introducing a kind of framework that McKinsey is proposing is wrong-headed and certain to backfire. Such a framework will most likely do far more harm than good to organizations &#8212; and to the engineering culture at companies and the damage could take years to undo.&#8221;</p><p>Beck shared a cautionary tale from Facebook that perfectly illustrates surrogation in action. The company had introduced developer satisfaction surveys &#8212; a reasonable approach. But then managers computed an overall score. Then those scores appeared in performance reviews. Directors pressured managers for better scores. Managers began negotiating with engineers: higher survey scores in exchange for better performance ratings. The result was worse organisational outcomes despite improved metrics. First, surrogation: leadership began treating the score <em>as</em> developer satisfaction, forgetting that it was merely a proxy. Then Goodhart&#8217;s Law kicked in: once the surrogate became a target, people gamed it. <strong>The metric had consumed the thing it was supposed to measure.</strong></p><p>The McKinsey episode matters because it revealed a fault line that runs through every discussion of developer productivity: the tension between what management consulting wants (simple numbers that enable comparison and control) and what software engineering actually is (complex, collaborative, creative knowledge work that resists simplification).</p><div><hr></div><h2><strong>4. Why Software Is Fundamentally Different</strong></h2><p>The reason every productivity metric fails or gets corrupted is not that we haven&#8217;t found the right metric yet. It is that software development is a fundamentally different kind of work than what most measurement frameworks were designed for.</p><h3><strong>Knowledge Work Is Not Manufacturing</strong></h3><p>Peter Drucker, who coined the term &#8220;knowledge work,&#8221; identified the crucial distinction: in manufacturing, the task is defined externally, and the worker&#8217;s job is to execute efficiently. In knowledge work, &#8220;the task of <em>what</em> to do is controlled by knowledge workers, they own the means of production (knowledge not machines), they decide what methods and steps to use, and they focus on <em>what to do</em> and on the right things &#8212; effectiveness.&#8221; Drucker considered knowledge-work productivity &#8220;the greatest management task of this century, just as making manual work productive was the great management task of the last century.&#8221;</p><p>The manufacturing metaphor fails for software because output is not standardised (every piece of software solves a different problem), efficiency is not the primary constraint (<strong>figuring out </strong><em><strong>what</strong></em><strong> to build matters more than </strong><em><strong>how fast</strong></em><strong> you build it</strong>), quality is not independently measurable (a feature that works but nobody needs has negative value), and the work itself is fundamentally creative &#8212; requiring &#8220;anticipatory imagination, problem solving, problem seeking, and generating ideas.&#8221;</p><h3><strong>Output vs. Outcome</strong></h3><p>Martin Fowler articulates the core principle: &#8220;If a team delivers lots of functionality... that functionality doesn&#8217;t matter if it doesn&#8217;t help the user improve their activity.&#8221; Output is what is produced: features shipped, code written, pull requests merged. Outcome is the change that the output creates: increased revenue, reduced support tickets, improved user satisfaction.</p><p>The insight is disarmingly simple: &#8220;Your job is to minimize output, and maximize outcome and impact.&#8221; The best solution is often the smallest change, the feature <em>not</em> built, the code <em>deleted</em>. Under any output-based metric, these acts of productive restraint are invisible or actively penalised.</p><p>Fowler dismisses the objection that outcomes are hard to measure: &#8220;We are very good at measuring financial outcomes.&#8221; The problem is not that outcomes cannot be measured, but that organisations prefer the comfort of output metrics because they are immediate and controllable &#8212; even when they measure the wrong thing.</p><h3><strong>The Invisible Work Problem</strong></h3><p>Charity Majors, CTO of Honeycomb, offers perhaps the most uncomfortable truth: &#8220;Some of the hardest and most impactful engineering work will be all but invisible on any set of individual metrics.&#8221;</p><p>Consider the work that keeps a team effective: pairing and mobbing that improve quality and spread knowledge across the team &#8212; but show up as &#8220;two people doing one person&#8217;s job&#8221; in output metrics. Mentoring that multiplies team capacity &#8212; but reduces the mentor&#8217;s personal output. Architectural decisions that save months of future work &#8212; but take hours that produce no visible deliverable. Cross-team coordination that unblocks others. Documentation that prevents future confusion. Technical debt reduction that improves future velocity but produces zero features.</p><p><strong>None of this registers on any output metric. All of it is essential.</strong></p><p>Majors pushes further: &#8220;Metrics are for easy problems &#8212; discrete, self-contained, well-understood problems. The more challenging and novel a problem, the less reliable these metrics will be.&#8221; And the most provocative observation: &#8220;To the extent you can reduce a job to a set of metrics, that job can be automated away.&#8221; <strong>If developer productivity could truly be captured in a number, developers would already be obsolete.</strong></p><h3><strong>The Individual Measurement Trap</strong></h3><p>Dave Farley puts it plainly: &#8220;Measuring software development in terms of individual developer productivity is a terrible idea. Being smart in how we structure and organize our work is much more important than the level of individual genius.&#8221;</p><p>The harms are well-documented. Individual measurement distorts behaviour: people optimise for their own numbers at the expense of the team. It creates misaligned incentives: short-term outputs over long-term value. It renders collaborative work invisible: the developer who writes little code but dramatically improves team performance through guidance and knowledge sharing is undervalued. And it damages culture: &#8220;Monitoring individual performance can cause unnecessary anxiety, even for top-performing contributors. It can also lead to overworking, an overly competitive work culture, and developer burnout.&#8221;</p><div><hr></div><h2><strong>5. Enter AI: The Amplifier of Everything</strong></h2><p>Into this already confused landscape, artificial intelligence has arrived with the promise of making every developer dramatically more productive. The reality is far more complex &#8212; and for leaders relying on traditional metrics, far more dangerous.</p><h3><strong>We Do Not Know What We Think We Know</strong></h3><p>Studies on AI-assisted coding productivity contradict each other wildly &#8212; some report massive speed gains, others find slowdowns and more bugs. But the honest takeaway is not that &#8220;it depends.&#8221; It is that <strong>we simply do not have reliable data yet</strong>. Every developer uses AI differently &#8212; different tools, different prompting styles, different levels of trust and scrutiny, different types of work. The studies measure wildly different things under wildly different conditions. They are not two sides of a debate. They are measurements of entirely different activities that happen to share the label &#8220;coding with AI.&#8221;</p><p>Until the industry converges on how AI is actually used in practice, the numbers should be treated with deep scepticism. Leaders who cite any of these studies to justify or reject AI investment are building on sand.</p><p>What we <em>can</em> observe, however, is a persistent perception gap: developers consistently believe AI is helping them more than it measurably does. This matters for leadership. If developers cannot accurately assess whether AI tools are making them more effective, then the comfortable feedback loop &#8212; &#8220;we bought the tool, developers say they like it, therefore it&#8217;s working&#8221; &#8212; may be an illusion.</p><p>We can also observe quality trends that should concern anyone paying attention. Code duplication is rising. Refactoring is collapsing. Code churn is increasing. Even the CEO of Cursor &#8212; a company that <em>sells</em> an AI coding tool &#8212; has warned publicly that developers accept AI-generated code &#8220;simply because it appears to work, without properly reviewing its structure, logic and long-term impact.&#8221; <strong>When the people selling the tools are urging caution, it is worth listening.</strong></p><h3><strong>AI Makes Bad Metrics Worse</strong></h3><p>Every problem with metrics described in this article &#8212; gaming, perverse incentives, the McNamara Fallacy, Goodhart&#8217;s Law, surrogation &#8212; becomes dramatically worse when AI enters the picture.</p><p>A human developer gaming a velocity target has natural friction: there are only so many hours in a day, only so much code a person can write. AI has none of these constraints. If the metric is lines of code, AI will generate mountains of it. If the metric is pull requests merged, AI will create dozens. If the metric is deployment frequency, AI will ship constantly. All while the codebase bloats, technical debt compounds, and the actual product &#8212; the thing users need &#8212; remains unchanged or gets worse.</p><p>AI also deepens surrogation. When a human writes code, reviewers can draw on shared context to assess whether the work actually serves the original goal. AI-generated output lacks that shared context, making it harder to notice the gap between the metric and the thing the metric was supposed to represent. And AI tool adoption metrics are themselves susceptible to surrogation: &#8220;AI adoption rate&#8221; or &#8220;percentage of code generated by AI&#8221; can quietly replace the actual goal &#8212; effective use of AI to improve engineering outcomes &#8212; in leaders&#8217; minds. Adoption <em>becomes</em> the goal, regardless of whether it improves anything.</p><p>Goodhart&#8217;s Law becomes more dangerous as optimisation power increases. AI is a step-function increase in optimisation power. <strong>Applied to bad metrics, it does not merely fail &#8212; it fails spectacularly, at speed, and at scale.</strong></p><p>The 2025 DORA Report puts it simply: &#8220;AI doesn&#8217;t fix a team; it amplifies what&#8217;s already there.&#8221; <strong>AI is not a productivity solution. It is a productivity </strong><em><strong>amplifier</strong></em><strong>. And amplifiers do not care what signal they are boosting.</strong></p><div><hr></div><h2><strong>6. What Actually Works</strong></h2><p>If simple metrics fail, modern frameworks are misapplied, and AI amplifies dysfunction, what should leadership actually do? The evidence points toward a fundamentally different approach &#8212; one that requires more organisational maturity but produces genuinely better results.</p><h3><strong>Measure Experience, Not Output</strong></h3><p>The most actionable path forward is to stop measuring what developers produce and start measuring the conditions that enable production:</p><p><strong>Feedback loops.</strong> How quickly do developers get information about their work? How long does a CI pipeline take? How fast do code reviews come back? How soon after deployment do they know if something broke? Every minute of waiting is a minute of lost context, and context &#8212; not typing speed &#8212; is the primary constraint on developer effectiveness.</p><p><strong>Cognitive load.</strong> How much mental effort does the development environment demand? How many systems must a developer understand to make a change? How clear is the documentation? How usable are the tools? Teams drowning in complexity do not need productivity metrics. They need simpler systems.</p><p><strong>Flow state.</strong> How often can developers achieve sustained, uninterrupted focus? Research shows it takes approximately 23 minutes to refocus after each interruption, and minimum effective deep work periods are around 90 minutes. <strong>An organisation that interrupts developers every 30 minutes for status meetings and Slack messages is not suffering from a productivity problem. It is suffering from a management problem.</strong></p><p>Almost half of tech managers now report that their companies measure developer productivity, developer experience, or both. Google, Microsoft, and Spotify have long relied on developer surveys to understand the conditions their teams work in. The shift is happening &#8212; but too slowly, and too often alongside the old metrics rather than replacing them.</p><h3><strong>Focus on Team Outcomes, Not Individual Metrics</strong></h3><p>Kent Beck&#8217;s recommendation is elegant in its simplicity: focus on &#8220;producing at least one customer-facing thing per team, per week.&#8221; This is not a metric to be gamed. It is a practice that aligns incentives: the whole team collaborates to deliver something a customer can see and respond to. Value is delivered. Feedback is gathered. The cycle continues.</p><p>Gergely Orosz and Abi Noda studied 17 major tech companies and found that &#8220;rather than wholesale adoption of frameworks like DORA, leading teams use a mix of org-specific qualitative and quantitative metrics.&#8221; <strong>The best organisations do not adopt a framework off the shelf. They build their own understanding of what &#8220;good&#8221; looks like in their specific context.</strong></p><h3><strong>Use Qualitative Approaches Seriously</strong></h3><p>Google&#8217;s approach to developer productivity measurement reveals a counterintuitive truth: qualitative metrics &#8212; &#8220;measurements comprised of data provided by humans&#8221; &#8212; capture what automated systems cannot. Flow state, codebase navigability, technical debt perception, satisfaction &#8212; these are real and consequential, and they can only be measured by asking people.</p><p>Google&#8217;s own analysis of 117 metrics for technical debt found that <em>none</em> were valid indicators. Human judgement about the gap between ideal and actual state proved essential. The numbers, on their own, told the wrong story.</p><p>Practical recommendations from the research: start with qualitative baselines to identify opportunities, then deploy targeted quantitative metrics for deeper analysis. Segment by team and persona rather than aggregating company-wide. Prioritise free-text comments &#8212; developers suggest improvements and identify gaps that structured questions miss. Use transactional surveys at workflow touchpoints for granular, timely feedback. And above all, act on what you learn. Nothing kills a survey programme faster than asking for feedback and visibly ignoring it.</p><h3><strong>Think in Systems, Not Individuals</strong></h3><p>W. Edwards Deming&#8217;s insight from 1993 remains foundational: &#8220;Left to themselves, system components become selfish, competitive, independent profit centres and thus destroy the system. The secret is cooperation between components toward the aim of the organisation.&#8221;</p><p>The practical application is straightforward: use Theory of Constraints thinking. Identify the bottleneck in the system and focus improvement efforts there. <strong>Individual developer speed is rarely the actual constraint.</strong> More often, it is slow code reviews, unclear requirements, flaky CI pipelines, cumbersome deployment processes, or poor cross-team communication. Making developers type faster &#8212; whether through training or AI &#8212; does nothing to address these systemic bottlenecks.</p><h3><strong>Use AI Wisely</strong></h3><p>The evidence does not say &#8220;don&#8217;t use AI.&#8221; It says &#8220;use AI with open eyes.&#8221; AI coding assistants are genuinely valuable for boilerplate, repetitive tasks, and exploration. They are genuinely dangerous when used uncritically on complex, context-rich work &#8212; and when their output is measured by the same broken metrics that have always failed.</p><p>The 2025 DORA finding bears repeating: AI amplifies what is already there. <strong>Before investing in AI tools, invest in the practices that AI will amplify.</strong> Fix the feedback loops, reduce the cognitive load, protect the flow state, clarify the team outcomes. Then introduce AI into a healthy system &#8212; where it will make good things better rather than making bad things faster.</p><div><hr></div><h2><strong>7. The Question Behind the Question</strong></h2><p>When an organisation asks &#8220;How do we measure developer productivity?&#8221;, it is worth pausing to ask why.</p><p>Sometimes the answer is benign: we want to understand where our bottlenecks are so we can remove them. Sometimes it is less benign: we want to identify which engineers to fire. Often, it is anxious: we are spending a lot on engineering and we do not know if we are getting value.</p><p>Each of these motivations leads to a different approach. The first calls for systems thinking and understanding what developers actually experience. The second calls for an honest conversation about whether the problem is individual performance or organisational dysfunction. The third calls for outcome measurement &#8212; connecting engineering work to business results, accepting that the feedback loop is long, and resisting the temptation to substitute easy output metrics for hard outcome questions.</p><p>The real question is not &#8220;How productive are our developers?&#8221; The real question is: <strong>&#8220;How do we create the conditions where developers can do their best work?&#8221;</strong></p><p>This is a harder question. It does not produce a single number. It cannot be answered by a dashboard. It requires leadership to engage with the actual work of software development &#8212; its complexity, its creativity, its inherent resistance to simplification.</p><p>But it is the right question. And in the age of AI, where the temptation to optimise bad metrics at machine speed has never been greater, asking the right question has never mattered more.</p><p>Bill Atkinson understood this in 1982. He made software six times faster by writing 2,000 fewer lines of code. By every metric except the ones that matter, he had a terrible week.</p><p>By the ones that matter, it was one of the most productive weeks in the history of software engineering.</p><div><hr></div><h2><strong>References</strong></h2><ol><li><p>Folklore.org, &#8220;Negative 2000 Lines Of Code&#8221; &#8212; <a href="https://www.folklore.org/Negative_2000_Lines_Of_Code.html">https://www.folklore.org/Negative_2000_Lines_Of_Code.html</a></p></li><li><p>Martin Fowler, &#8220;CannotMeasureProductivity&#8221; (2003) &#8212; <a href="https://martinfowler.com/bliki/CannotMeasureProductivity.html">https://martinfowler.com/bliki/CannotMeasureProductivity.html</a></p></li><li><p>Martin Fowler, &#8220;OutcomeOverOutput&#8221; &#8212; <a href="https://martinfowler.com/bliki/OutcomeOverOutput.html">https://martinfowler.com/bliki/OutcomeOverOutput.html</a></p></li><li><p>Martin Fowler, &#8220;Measuring Developer Productivity via Humans&#8221; &#8212; <a href="https://martinfowler.com/articles/measuring-developer-productivity-humans.html">https://martinfowler.com/articles/measuring-developer-productivity-humans.html</a></p></li><li><p>Nicole Forsgren et al., &#8220;The SPACE of Developer Productivity,&#8221; ACM Queue (2021) &#8212; <a href="https://queue.acm.org/detail.cfm?id=3454124">https://queue.acm.org/detail.cfm?id=3454124</a></p></li><li><p>Nicole Forsgren, Jez Humble, Gene Kim, <em>Accelerate</em> (2018)</p></li><li><p>DORA Metrics Guide &#8212; <a href="https://dora.dev/guides/dora-metrics/">https://dora.dev/guides/dora-metrics/</a></p></li><li><p>DORA Report 2025, Google Cloud &#8212; <a href="https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report">https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report</a></p></li><li><p>Kent Beck &amp; Gergely Orosz, &#8220;Measuring developer productivity? A response to McKinsey,&#8221; The Pragmatic Engineer (2023) &#8212; </p></li><li><p>Gergely Orosz, &#8220;Measuring Developer Productivity: Real-World Examples,&#8221; The Pragmatic Engineer &#8212; </p></li><li><p>Dave Farley, &#8220;What McKinsey got wrong about developer productivity,&#8221; LeadDev &#8212; <a href="https://leaddev.com/process/what-mckinsey-got-wrong-about-developer-productivity">https://leaddev.com/process/what-mckinsey-got-wrong-about-developer-productivity</a></p></li><li><p>Charity Majors, &#8220;Questionable Advice: Can Engineering Productivity Be Measured?&#8221; &#8212; <a href="https://charity.wtf/2020/07/07/questionable-advice-can-engineering-productivity-be-measured/">https://charity.wtf/2020/07/07/questionable-advice-can-engineering-productivity-be-measured/</a></p></li><li><p>Peng et al., &#8220;The Impact of AI on Developer Productivity: Evidence from GitHub Copilot,&#8221; ArXiv (2023) &#8212; <a href="https://arxiv.org/abs/2302.06590">https://arxiv.org/abs/2302.06590</a></p></li><li><p>GitHub Blog, &#8220;Research: Quantifying GitHub Copilot&#8217;s Impact on Developer Productivity and Happiness&#8221; &#8212; <a href="https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness/">https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness/</a></p></li><li><p>METR, &#8220;Measuring the Impact of Early 2025 AI on Experienced Open-Source Developer Productivity&#8221; (2025) &#8212; <a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/">https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/</a></p></li><li><p>Uplevel Data Labs, &#8220;A Data-Driven Look at Gen AI for Coding&#8221; (2024) &#8212; <a href="https://resources.uplevelteam.com/gen-ai-for-coding">https://resources.uplevelteam.com/gen-ai-for-coding</a></p></li><li><p>GitClear, &#8220;AI Copilot Code Quality Research 2025&#8221; &#8212; <a href="https://www.gitclear.com/ai_assistant_code_quality_2025_research">https://www.gitclear.com/ai_assistant_code_quality_2025_research</a></p></li><li><p>Fortune, &#8220;Cursor CEO Michael Truell warns vibe coding builds &#8216;shaky foundations&#8217;&#8221; (2025) &#8212; <a href="https://fortune.com/2025/12/25/cursor-ceo-michael-truell-vibe-coding-warning-generative-ai-assistant/">https://fortune.com/2025/12/25/cursor-ceo-michael-truell-vibe-coding-warning-generative-ai-assistant/</a></p></li><li><p>Peter Drucker, &#8220;Knowledge-Worker Productivity: The Biggest Challenge&#8221; &#8212; </p></li><li><p>Cal Newport, <em>Deep Work: Rules for Focused Success in a Distracted World</em></p></li><li><p>W. Edwards Deming, <em>The New Economics for Industry, Government, Education</em> (1993)</p></li><li><p>Wikipedia, &#8220;McNamara Fallacy&#8221; &#8212; <a href="https://en.wikipedia.org/wiki/McNamara_fallacy">https://en.wikipedia.org/wiki/McNamara_fallacy</a></p></li><li><p>Wikipedia, &#8220;Goodhart&#8217;s Law&#8221; &#8212; <a href="https://en.wikipedia.org/wiki/Goodhart%27s_law">https://en.wikipedia.org/wiki/Goodhart&#8217;s_law</a></p></li><li><p>Scrum.org, &#8220;Velocity, the False Metric of Productivity&#8221; &#8212; <a href="https://www.scrum.org/resources/blog/velocity-false-metric-productivity">https://www.scrum.org/resources/blog/velocity-false-metric-productivity</a></p></li><li><p>LinearB, &#8220;Why Agile Velocity is the Most Dangerous Metric&#8221; &#8212; <a href="https://linearb.io/blog/why-agile-velocity-is-the-most-dangerous-metric-for-software-development-teams">https://linearb.io/blog/why-agile-velocity-is-the-most-dangerous-metric-for-software-development-teams</a></p></li><li><p>Aviator, &#8220;Everything Wrong with DORA Metrics&#8221; &#8212; <a href="https://www.aviator.co/blog/everything-wrong-with-dora-metrics/">https://www.aviator.co/blog/everything-wrong-with-dora-metrics/</a></p></li><li><p>Matt Hopkins, &#8220;Goodhart&#8217;s Law for AI Agents&#8221; &#8212; <a href="https://matthopkins.com/business/goodharts-law-ai-agents/">https://matthopkins.com/business/goodharts-law-ai-agents/</a></p></li><li><p>InfoWorld, &#8220;Software development meets the McNamara Fallacy&#8221; &#8212; <a href="https://www.infoworld.com/article/4010318/software-development-meets-the-mcnamara-fallacy.html">https://www.infoworld.com/article/4010318/software-development-meets-the-mcnamara-fallacy.html</a></p></li><li><p>Microsoft Research, &#8220;The SPACE of Developer Productivity&#8221; &#8212; <a href="https://www.microsoft.com/en-us/research/publication/the-space-of-developer-productivity-theres-more-to-it-than-you-think/">https://www.microsoft.com/en-us/research/publication/the-space-of-developer-productivity-theres-more-to-it-than-you-think/</a></p></li><li><p>ArXiv, &#8220;Leveraging Creativity in Software Engineering&#8221; &#8212; <a href="https://arxiv.org/html/2502.03280v1">https://arxiv.org/html/2502.03280v1</a></p></li><li><p>Choi, J., Hecht, G., and Tayler, W.B. (2012). &#8220;Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation.&#8221; <em>The Accounting Review</em>, 87(4), 1135-1164.</p></li><li><p>Black, P., Meservy, T., Tayler, W.B., and Williams, J.O. (2021). &#8220;Surrogation Fundamentals: Measurement and Cognition.&#8221; <em>Journal of Management Accounting Research</em>, 34(1), 9-28.</p></li><li><p>Harris, M. and Tayler, W.B. (2019). &#8220;Don&#8217;t Let Metrics Undermine Your Business.&#8221; <em>Harvard Business Review</em>, September-October 2019.</p></li><li><p>Kahneman, D. and Frederick, S. (2002). &#8220;Representativeness Revisited: Attribute Substitution in Intuitive Judgment.&#8221; In <em>Heuristics and Biases: The Psychology of Intuitive Judgment</em>. Cambridge University Press.</p></li></ol>]]></content:encoded></item><item><title><![CDATA[“How do you handle conflict between two team members?”]]></title><description><![CDATA[What the interview question doesn&#8217;t tell you, and what psychology, philosophy, and organisational science do]]></description><link>https://a4al6a.substack.com/p/how-do-you-handle-conflict-between</link><guid isPermaLink="false">https://a4al6a.substack.com/p/how-do-you-handle-conflict-between</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Tue, 10 Mar 2026 00:12:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most organisations treat conflict like a burst pipe. Something has gone wrong, someone needs to fix it fast, and the goal is to get back to normal as quickly as possible. Bring in HR. Have the difficult conversation. Move people to different teams if you have to. Patch the leak.</p><p>This instinct is understandable, but it misses something fundamental. Conflict doesn&#8217;t appear out of nowhere. It surfaces because something underneath is unresolved: competing values, unclear expectations, unmet needs, a power dynamic nobody has named, or simply two people with genuinely different views on what good looks like. Treat the symptom without understanding the cause, and the conflict will come back. Usually louder.</p><p>The first shift worth making is conceptual. <strong>Conflict is not the opposite of a healthy team. Avoidance is</strong>. A team where nobody ever disagrees is not a harmonious team. It&#8217;s a team where people have learned that expressing disagreement is unsafe or pointless. That is a much more serious problem, because it&#8217;s invisible.</p><p>Patrick Lencioni, in his work on team dysfunction, placed absence of trust at the very base of what makes teams fall apart. But trust here doesn&#8217;t mean &#8220;we all get on well.&#8221; It means psychological safety: the shared belief that you can speak honestly without being punished for it. <a href="https://www.linkedin.com/in/amycedmondson/">Prof. Amy Edmondson</a> at Harvard has spent decades studying this, and the research is consistent. <strong>Teams with high psychological safety don&#8217;t avoid conflict. They have more of it. What they&#8217;re better at is making it productive.</strong></p><p>So the question isn&#8217;t how to remove conflict from a team. The question is how to raise the quality of it.</p><div><hr></div><h2>The three patterns underneath most team conflicts</h2><p>Underneath most team conflicts, if you look closely enough, you&#8217;ll find one of a small number of recurring patterns.</p><p>The first is competing legitimate goods. Two people are both right, but they&#8217;re optimising for different things. The engineer wants to do it properly; the product manager wants to ship it now. The team lead wants consistency; the senior developer wants autonomy. Neither of them is wrong. <strong>They&#8217;re just prioritising different values that are genuinely in tension</strong>. Aristotle would have called this a problem of <em>phronesis</em>, practical wisdom: knowing not just what is good, but how to navigate situations where several good things pull in different directions. Resolving this kind of conflict requires surfacing both sets of values explicitly and deciding together which one takes precedence in this context. That&#8217;s not a compromise. It&#8217;s a choice made consciously.</p><p>The second pattern is attribution asymmetry. We judge ourselves by our intentions, but we judge others by their behaviour. I was blunt because I was under pressure. You were blunt because you don&#8217;t respect me. This cognitive bias, well documented in social psychology, turns interpersonal friction into moral judgement faster than almost anything else. Once someone has been labelled as difficult, dismissive, or political, the label tends to stick, and the actual content of the disagreement gets buried under a layer of personal narrative. The antidote is <strong>perspective-taking: a genuine attempt to understand what the other person is trying to accomplish, and what pressures they&#8217;re navigating</strong>. This sounds obvious. It&#8217;s surprisingly rare in practice.</p><p>The third pattern is structural conflict. This is the one teams most often misdiagnose as personal. Two people seem unable to get along, but actually their roles are set up in a way that makes them adversarial. They&#8217;re rewarded for different things, they answer to different people, and the organisation has never clearly decided who owns what. Chris Argyris called this &#8220;organisational defensive routines&#8221;: the ways in which the system itself generates stress and then buries the evidence. You can coach both individuals to within an inch of their lives and it won&#8217;t help, because <strong>the conflict is in the structure, not in the people.</strong></p><div><hr></div><h2>Behind every accusation, there is an unmet need</h2><p>Marshall Rosenberg&#8217;s Nonviolent Communication framework offers one of the most practically useful lenses for working through interpersonal conflict. Its core insight is that behind every accusation, there is an unmet need. When someone says &#8220;you never listen to me,&#8221; they&#8217;re not really talking about listening. They&#8217;re expressing something about recognition, respect, or belonging. When someone says &#8220;this team has no standards,&#8221; they&#8217;re usually expressing a need for quality, predictability, or professional pride.</p><p>NVC asks you to separate observations from interpretations, and to connect behaviour to feelings and needs rather than to character judgements. It sounds clinical when described, but in practice it does something powerful: <strong>it shifts conversations from blame to need, which is the only terrain on which genuine resolution becomes possible.</strong> You can argue forever about who did what. It&#8217;s much harder to dismiss someone&#8217;s need once it&#8217;s been clearly named.</p><p>The philosopher J&#252;rgen Habermas wrote about what he called the ideal speech situation: a conversation in which all parties have equal standing, speak honestly, and are open to being persuaded by the better argument rather than by power or status. It&#8217;s utopian, obviously. But it&#8217;s useful as a direction to aim in. The question it generates is practical: <strong>what would need to be true for this conversation to be one where the best idea wins, rather than the loudest voice or the highest rank?</strong></p><div><hr></div><h2>The drama triangle and why nobody moves forward</h2><p>The drama triangle, originally described by Stephen Karpman in the 1960s, is worth knowing because it describes a pattern that plays out in teams with almost mechanical regularity. It involves three roles: the persecutor, the victim, and the rescuer. They seem like fixed identities, but they&#8217;re not. People shift between them constantly, often within a single conversation. The manager who&#8217;s trying to protect their team member (rescuer) becomes the person creating dependency (persecutor). The colleague who feels unfairly treated (victim) becomes the one sabotaging decisions (persecutor). The rescuer feels righteous; the victim feels justified; the persecutor feels either powerful or vindicated. Nobody is moving forward.</p><p>The way out of the drama triangle, according to David Emerald&#8217;s reframing of it, is to shift from reactive to creative: from &#8220;what&#8217;s wrong and who&#8217;s to blame&#8221; to &#8220;what do we want and what can we do.&#8221; This isn&#8217;t about being positive. It&#8217;s about redirecting agency. <strong>Conflict locks people into a past-oriented, causal story. Resolution requires a future-oriented, intentional one.</strong></p><div><hr></div><h2>The leader&#8217;s real job is not to resolve conflict but to contain it</h2><p>The leader&#8217;s role in all of this is often misunderstood. Many leaders believe their job is to resolve conflicts: to step in, assess the situation, and hand down a decision. This sometimes works in the short term. What it teaches the team is that conflict is the leader&#8217;s problem to solve, not their own capacity to develop. Over time, it creates a team that escalates rather than resolves, and a leader who is permanently stuck in the middle.</p><p>A more useful frame, drawn from systemic coaching and family systems theory, is the leader as container. <strong>The job is not to eliminate tension but to create conditions in which the team can sit with tension long enough to understand it</strong>. This requires a specific kind of confidence: the ability to resist the pull towards premature resolution. Ronald Heifetz at Harvard Kennedy School distinguishes between technical problems, which have known solutions that someone with expertise can apply, and adaptive challenges, which require the people involved to change how they think. Most team conflicts are adaptive. They can&#8217;t be solved by the leader implementing the right answer. They require the team to work through something together.</p><p><strong>This is why the most important thing a leader can do in a moment of conflict is often not to solve it, but to slow it down</strong>. Ask questions rather than provide answers. Name what&#8217;s happening without taking sides. Make the implicit explicit. Create enough safety for people to say what they actually think.</p><div><hr></div><h2>Conflict prevention happens in ordinary moments, not just critical ones</h2><p>There&#8217;s a strand of sociology, particularly in the work of Randall Collins on interaction ritual chains, that points to something teams rarely discuss: the emotional energy in a room. When people feel seen, included, and engaged, they generate positive emotional energy that makes collaboration easier. When they feel marginalised, ignored, or dismissed, they generate negative emotional energy that makes everything harder. This isn&#8217;t soft. It&#8217;s how groups actually work. Social identity theory, developed by Henri Tajfel and John Turner, shows how quickly people sort themselves into in-groups and out-groups, and how much of team conflict is actually about belonging rather than the ostensible subject of the disagreement.</p><p>What this means practically is that <strong>a lot of conflict prevention happens not in the difficult moments but in the ordinary one</strong>s. Whether you notice someone&#8217;s contribution in a meeting. Whether you ask about someone&#8217;s view before the decision is made. Whether you check in with the quieter members of the team, not just the loudest. These small acts of inclusion or exclusion accumulate over time and determine whether people feel invested in the team&#8217;s success or quietly estranged from it.</p><div><hr></div><h2>When conflict genuinely cannot be resolved</h2><p>None of this is to say that all conflicts can be resolved by better communication and more empathy. Sometimes values are genuinely incompatible. Sometimes people are acting in bad faith. Sometimes the problem really is that two people cannot work together, and the most honest thing to do is acknowledge that rather than keep pretending otherwise.</p><p>The Stoics were clear about what lies within our control and what doesn&#8217;t. You can change how you communicate, what you invite, what you model, what structures you put in place. You cannot change another person&#8217;s willingness to engage. At some point, continued investment in a conflict that the other party is not interested in resolving becomes a form of self-deception.</p><p>But that point comes much later than most people think. The vast majority of team conflicts that get written off as irresolvable have actually never been properly named, never had the structural conditions examined, never had someone ask sincerely what each person actually needs. They&#8217;ve had plenty of difficult conversations in which two people talked past each other while the real issue stayed underground.</p><div><hr></div><h2>What to actually do</h2><p>If I had to distil this into something actionable, it would be this.</p><p>Start by assuming the conflict is information. Ask what it&#8217;s telling you about values, needs, structures, or dynamics that haven&#8217;t been made visible yet. Resist the pull towards fast resolution. Create conditions where people can speak honestly without penalty. Separate the people from the positions and try to understand the interests underneath. Name what&#8217;s happening, including the things that feel uncomfortable to name. Ask whose voices are missing from the conversation.</p><p>And remember that a team capable of navigating conflict well is not a team that has achieved some kind of permanent peace. <strong>It&#8217;s a team that has developed the muscle to disagree, recover, and keep going together.</strong> That muscle is one of the most valuable things a team can have. It doesn&#8217;t come from avoiding conflict. It comes from doing it better, over and over, until it stops feeling like a threat and starts feeling like work.</p><div><hr></div><h2>References</h2><p>Patrick Lencioni &#8212; <em>The Five Dysfunctions of a Team</em> (2002)</p><p>Amy Edmondson &#8212; <em>The Fearless Organization</em> (2018); &#8220;Psychological Safety and Learning Behavior in Work Teams,&#8221; <em>Administrative Science Quarterly</em> (1999)</p><p>Aristotle &#8212; <em>Nicomachean Ethics</em>, on phronesis (practical wisdom)</p><p>Chris Argyris &#8212; <em>Overcoming Organizational Defenses</em> (1990)</p><p>Marshall Rosenberg &#8212; <em>Nonviolent Communication: A Language of Life</em> (2003)</p><p>J&#252;rgen Habermas &#8212; <em>The Theory of Communicative Action</em> (1981)</p><p>Stephen Karpman &#8212; &#8220;Fairy Tales and Script Drama Analysis,&#8221; <em>Transactional Analysis Bulletin</em> (1968)</p><p>David Emerald &#8212; <em>The Power of TED</em> (2009)</p><p>Ronald Heifetz &#8212; <em>Leadership Without Easy Answers</em> (1994); <em>The Practice of Adaptive Leadership</em> (2009, with Linsky and Grashow)</p><p>Randall Collins &#8212; <em>Interaction Ritual Chains</em> (2004)</p><p>Henri Tajfel and John Turner &#8212; &#8220;An Integrative Theory of Intergroup Conflict,&#8221; in <em>The Social Psychology of Intergroup Relations</em> (1979)</p><p>Epictetus &#8212; <em>Enchiridion</em>, on the dichotomy of control</p>]]></content:encoded></item><item><title><![CDATA[Switching from Scrum to Kanban won’t (necessarily) save you]]></title><description><![CDATA[You're tasting the same jam.]]></description><link>https://a4al6a.substack.com/p/switching-from-scrum-to-kanban-wont</link><guid isPermaLink="false">https://a4al6a.substack.com/p/switching-from-scrum-to-kanban-wont</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Sun, 01 Feb 2026 13:02:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A few days ago, <a href="https://www.linkedin.com/posts/mbowler_the-excellent-book-the-illusionist-brain-activity-7423417322410831872-31JB?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAAAi_0UBvCEtV6uUwi_Wgb6_PVUmPSZDtcQ">Mike Bowler shared a fascinating psychology experiment</a>. Participants taste jam from two jars and pick the one they prefer. The experimenter then secretly switches the jars and asks them to taste their &#8220;chosen&#8221; jam again. Most people don&#8217;t notice. They taste the jam they originally rejected and confidently explain why it&#8217;s the better one.</p><p>This is what happens when some teams switch from Scrum to Kanban.</p><p>A team struggles with Scrum. Sprint planning feels pointless. Retrospectives produce nothing. Velocity is treated as a performance metric. They estimate in story points. Standups are boring status reports. Someone reads a LinkedIn post about how &#8220;sprints are dead,&#8221; and the team decides to make the switch. They rip out the sprints, put up a Kanban board, and carry on.</p><p>Nothing actually changes. The same dysfunctions are still there, running underneath a different label. The planning is still vague. The feedback loops are still weak. The problems to solve are still unclear. The team has picked up the other jar and is now confidently defending why this jam tastes better. It&#8217;s the same jam. They just can&#8217;t tell, because they never understood what they were tasting in the first place.</p><p>This article is about why that happens. It&#8217;s about what agile frameworks actually do, what most people get wrong about Scrum, what most people get wrong about Kanban, and why the conversation about switching between them is almost always asking the wrong question.</p><h2>Frameworks are mirrors, not medicines</h2><p>Here&#8217;s the most important thing I can tell you about any agile framework: it doesn&#8217;t fix your problems. It makes your problems visible.</p><p>A sprint that feels painful isn&#8217;t painful because of the sprint. The sprint is showing you something. Maybe your team can&#8217;t finish anything in two weeks because the work isn&#8217;t broken down properly. Maybe planning feels like fiction because nobody actually understands the requirements. Maybe the demo is embarrassing because the team built something nobody asked for. These are not sprint problems. These are organisational problems, team problems, requirements problems. The sprint is the mirror.</p><p>When you remove the mirror, you don&#8217;t fix the dysfunction. You just stop seeing it.</p><p>This is the role agile frameworks play. Scrum&#8217;s ceremonies create regular, unavoidable moments where reality has to be confronted. Did we build the right thing? Are we improving? Can we actually finish what we started? If the answers are uncomfortable, the instinct is to blame the framework. But the framework is just the thing that forced you to ask.</p><p>Kanban does the same thing, by the way. A Kanban board with a massive pile of work-in-progress is a mirror showing you that your team can&#8217;t say no and can&#8217;t finish things. WIP limits that keep getting overridden are a mirror showing you that your organisation doesn&#8217;t respect capacity constraints. The question is never &#8220;which mirror do I want?&#8221; The question is &#8220;am I willing to look at what the mirror shows me?&#8221;</p><p>Most teams that switch from Scrum to Kanban aren&#8217;t choosing a better mirror. They&#8217;re choosing a less confrontational one. Sprints force you to face reality every two weeks whether you like it or not. Kanban is gentler about it. That&#8217;s not always an advantage.</p><h2>Everything you think you know about Scrum is probably wrong</h2><p>I&#8217;ve coached teams for years, and the number of misconceptions about Scrum is staggering. Most teams that say they&#8217;re &#8220;doing Scrum&#8221; are doing something else entirely, and most teams that say Scrum doesn&#8217;t work have never actually tried it. Let me go through the big ones.</p><p><strong>Sprints are mini-waterfalls.</strong> This is the most common misconception and it&#8217;s completely wrong. A sprint is not &#8220;two weeks to build what we planned.&#8221; A sprint is a timebox for learning. The Sprint Goal gives direction, but the team is expected to discover things along the way, adjust their approach, and deliver the most valuable increments they can. If your sprints feel like waterfall, the problem is how you&#8217;re running them, not the concept itself.</p><p><strong>Story points are part of Scrum.</strong> They&#8217;re not. Read the Scrum Guide. Story points appear nowhere in it. Story points are a widespread practice that became associated with Scrum through cultural osmosis, but Scrum doesn&#8217;t prescribe any estimation technique. If your story point discussions are wasting hours every sprint, that&#8217;s a choice your team made, not something Scrum imposed on you.</p><p><strong>Velocity is a Scrum metric.</strong> It&#8217;s not. It&#8217;s not in the Scrum Guide at all. Velocity became widespread through tools like Jira and through management&#8217;s desire for predictability metrics. In practice, it&#8217;s an antipattern. It incentivises gaming, it conflates output with value, and it gives the illusion of measurement where no real measurement exists. Scrum asks teams to forecast what they can accomplish in a sprint. It doesn&#8217;t say how. If your organisation is using velocity as a performance measure or a commitment mechanism, that&#8217;s a dysfunction your organisation created, not something Scrum asked for.</p><p><strong>You can&#8217;t change direction during a sprint.</strong> Wrong. The Sprint Goal provides focus, but the Scrum Guide explicitly says the scope of the sprint can be clarified and renegotiated with the Product Owner as more is learned. If something urgent comes up on Monday, nothing in Scrum prevents you from addressing it. If your team feels locked in, the problem is your interpretation, not the framework.</p><p><strong>The daily standup is a status report.</strong> It&#8217;s not. The Daily Scrum is a planning event. Its purpose is for the developers to inspect progress toward the Sprint Goal and adapt their plan for the next 24 hours. It&#8217;s not &#8220;tell the Scrum Master what you did yesterday.&#8221; If that&#8217;s what it&#8217;s become, you&#8217;ve turned a collaborative planning session into a reporting ceremony. The framework didn&#8217;t do that. You did.</p><p><strong>The Scrum Master is a project manager.</strong> The Scrum Master has no authority over the team. They don&#8217;t assign work, they don&#8217;t track progress, they don&#8217;t manage resources. They serve the team by helping remove impediments and by coaching the organisation on Scrum. If your Scrum Master is acting as a project manager, you don&#8217;t have a Scrum Master. You have a project manager with a different title.</p><p><strong>Retrospectives are blame sessions or wastes of time.</strong> If your retrospectives produce nothing, the problem is psychological safety, facilitation, or follow-through. The retrospective exists so the team can inspect how the last sprint went and make concrete improvements. If nobody speaks honestly, that&#8217;s a trust problem. If actions are identified but never implemented, that&#8217;s a commitment problem. The retrospective is the mirror showing you those problems exist.</p><p><strong>Scrum requires a specific team structure.</strong> Scrum requires the team to be cross-functional, meaning the team as a whole has all the skills needed to deliver. It does not mean everyone does everything. A team can have specialists. The requirement is that the team doesn&#8217;t depend on external groups to complete their work.</p><p><strong>Scrum is rigid and prescriptive.</strong> Scrum is actually a minimal framework. The Scrum Guide is 13 pages long. It defines three roles, five events, and three artefacts. Everything else, including story points, velocity, burndown charts, task boards, estimation techniques, and technical practices, is something teams and organisations have added on top. Most of the rigidity people complain about isn&#8217;t Scrum. It&#8217;s the layers of process their organisation built around Scrum.</p><p><strong>Scrum doesn&#8217;t allow for technical excellence.</strong> This one is particularly frustrating. Scrum doesn&#8217;t prescribe technical practices because it&#8217;s a project management framework, not an engineering framework. But nothing in Scrum prevents you from doing test-driven development, pair programming, continuous integration, refactoring, or any other XP practice. In fact, the Scrum Guide explicitly states that the Scrum Team is expected to maintain a high-quality Definition of Done. If your Scrum implementation ignores technical quality, you&#8217;ve chosen to do that. </p><p><strong>The Product Backlog is a list of requirements.</strong> The Product Backlog is an ordered list of what might be needed to improve the product. &#8220;Might&#8221; is doing a lot of work in that sentence. It&#8217;s not a specification document. It&#8217;s a living, evolving artefact that reflects current understanding. If your backlog is being treated as a fixed requirements list, you&#8217;ve lost the plot.</p><p><strong>Burndown charts are required.</strong> Not in the Scrum Guide. Use them if they help. Don&#8217;t use them if they don&#8217;t. They&#8217;re a tool, not a mandate.</p><p><strong>Sprints prevent continuous delivery.</strong> This is completely false and I&#8217;ll address it in detail later. The sprint is a planning and learning boundary, not a release boundary. You can deploy as often as you like within a sprint.</p><p>The common thread in all of these misconceptions is that teams confuse the accretions of corporate Scrum with actual Scrum. They rebel against story point negotiations, velocity tracking (an antipattern that Scrum never asked for), burndown charts, and rigid sprint commitments, and they think they&#8217;re rebelling against Scrum. They&#8217;re not. They&#8217;re rebelling against a cargo cult version of Scrum that their organisation created.</p><h2>The six practices of Kanban and why you must respect them</h2><p>Here&#8217;s where things get ironic. Most teams that switch from Scrum to Kanban because Scrum felt &#8220;too prescriptive&#8221; end up doing Kanban in a way that ignores most of what Kanban actually requires.</p><p>The Kanban Method, as defined by David Anderson, has six core practices. They&#8217;re not optional. They&#8217;re not suggestions. If you&#8217;re ignoring them, you&#8217;re not doing Kanban. You&#8217;re just not doing Scrum.</p><p><strong>Visualise the workflow.</strong> This means making the entire flow of work visible, from the moment something is requested to the moment it&#8217;s delivered. Not just &#8220;To Do, In Progress, Done.&#8221; The real workflow, with all its stages, queues, and handoffs. Most teams put up a three-column board and call it Kanban. That&#8217;s not visualising your workflow. That&#8217;s hiding your workflow behind a simplified picture.</p><p><strong>Limit work in progress.</strong> This is the single most important practice in Kanban, and it&#8217;s the one most teams ignore. WIP limits exist because starting many things and finishing few things is the primary cause of slow delivery. When you limit WIP, you force the team to finish things before starting new ones. This creates flow. Without WIP limits, you don&#8217;t have Kanban. You have a to-do list on a wall. And here&#8217;s the critical part: WIP limits must be respected. If your WIP limit is 3 and you have 5 items in progress, you don&#8217;t have a WIP limit. You have a decoration. The moment WIP limits become negotiable, you&#8217;ve lost the core mechanism that makes Kanban work.</p><p><strong>Manage flow.</strong> Kanban requires you to actively monitor and manage the flow of work through the system. This means measuring lead time, cycle time, throughput, and identifying bottlenecks. It means paying attention to where work gets stuck and taking action to unblock it. If you&#8217;re not measuring flow, you have no idea whether your system is improving or degrading. You&#8217;re flying blind.</p><p><strong>Make process policies explicit.</strong> This means writing down the rules. When is an item ready to be pulled into the next stage? What does &#8220;done&#8221; mean? Who can override a WIP limit, and under what conditions? Most teams that adopt Kanban operate on implicit assumptions that nobody has agreed to. This leads to exactly the kind of confusion and conflict they were trying to escape by leaving Scrum.</p><p><strong>Implement feedback loops.</strong> Kanban specifies several feedback loops, including daily standup meetings, service delivery reviews, operations reviews, and risk reviews. Yes, you read that correctly. Kanban has regular meetings. If you switched to Kanban because you wanted fewer meetings, you&#8217;ve misunderstood the method. The feedback loops in Kanban serve the same purpose as the events in Scrum: they force you to inspect reality and adapt.</p><p><strong>Improve collaboratively, evolve experimentally.</strong> Kanban emphasises continuous improvement through small, incremental changes based on evidence. This means you need retrospectives, or something functionally equivalent. You need to regularly examine your process, identify problems, try experiments, and measure results. If you switched from Scrum to Kanban and stopped doing retrospectives, you&#8217;ve removed your improvement mechanism.</p><p>Do you see the irony? Teams abandon Scrum because they find it too structured, then adopt Kanban while ignoring the structures that make Kanban work. The result is a team with no sprints, no WIP limits, no flow metrics, no explicit policies, no feedback loops, and no improvement process. That&#8217;s not Kanban. That&#8217;s anarchy with a board on the wall.</p><h2>Why switching without understanding solves nothing</h2><p>The core fallacy in the &#8220;Scrum to Kanban&#8221; narrative is the assumption that the framework is the problem. It almost never is. The problems are usually one or more of the following:</p><p>The organisation treats agile as a project management methodology rather than a product development approach. Work is pushed onto teams rather than pulled. Success is measured in output rather than outcomes. Management wants predictability but won&#8217;t invest in the conditions that create it. Teams don&#8217;t have the skills, autonomy, or psychological safety to self-organise. Requirements are vague and nobody has the authority or willingness to clarify them. Technical debt has accumulated to the point where every change is slow and risky. Dependencies between teams create bottlenecks that no process framework can solve.</p><p>None of these problems are caused by sprints. And none of them are solved by removing sprints. If you switch from Scrum to Kanban while carrying all of this baggage, you&#8217;ll get the same results with a different board layout.</p><p>The honest conversation isn&#8217;t &#8220;should we use Scrum or Kanban?&#8221; It&#8217;s &#8220;what are the actual problems we need to solve, and which practices help us solve them?&#8221;</p><h2>Scrum and Kanban are not opposites</h2><p>The framing of Scrum versus Kanban is fundamentally wrong. They&#8217;re not competing philosophies. They&#8217;re complementary sets of practices that address different aspects of the same challenge: how to deliver valuable software in the presence of uncertainty.</p><p>Scrum provides a cadence. Regular moments for planning, review, and adaptation. A timebox that creates urgency and focus. Defined roles that clarify responsibility.</p><p>Kanban provides flow management. Visualisation of work. WIP limits that prevent overload. Metrics that show whether your system is healthy.</p><p>There is nothing stopping you from using both. In fact, the most effective teams I&#8217;ve worked with do exactly that. They use sprints as planning and learning cycles while also applying WIP limits within those sprints. They visualise their workflow on a Kanban board while maintaining the Scrum cadence of Sprint Reviews and Retrospectives. They measure flow metrics like cycle time and throughput alongside sprint-based metrics. They treat the sprint as a feedback loop and the Kanban board as a diagnostic tool.</p><p>This integration is sometimes called Scrumban, and while the name is clumsy, the idea is sound. Take the useful parts of both. Timeboxes and WIP limits. Sprint Reviews and flow metrics. Retrospectives and explicit policies. The goal is not to pick a camp. The goal is to assemble a set of practices that helps your team deliver effectively.</p><p>The question is never &#8220;Scrum or Kanban?&#8221; The question is &#8220;what combination of practices makes our specific problems visible and gives us the feedback we need to improve?&#8221;</p><h2>Continuous delivery works with both, and most teams do neither</h2><p>One of the strangest arguments in the &#8220;sprints are dead&#8221; discourse is the idea that sprints prevent continuous delivery. This betrays a fundamental misunderstanding of what a sprint is.</p><p>A sprint is a planning boundary, not a release boundary. The Scrum Guide says the increments must be usable by the end of the sprint. It does not say you can only release at the end of the sprint. You can deploy to production every day, multiple times a day, within a sprint. Many teams do. In fact, the Scrum Guide encourages readers to do that: &#8220;<em>Multiple Increments may be created within a Sprint. The sum of the Increments is presented at the Sprint Review thus supporting empiricism. However, an Increment may be delivered to stakeholders prior to the end of the Sprint. The Sprint Review should never be considered a gate to releasing value.</em>&#8221;</p><p>Continuous delivery is a technical capability. It requires automated testing, continuous integration, a deployment pipeline, feature toggles, and an architecture that supports independent deployment. These are engineering practices. They&#8217;re independent of whether you use Scrum, Kanban, or anything else.</p><p>If your team can&#8217;t do continuous delivery, switching from Scrum to Kanban won&#8217;t change that. You&#8217;ll still have the same manual testing, the same merge conflicts, the same deployment bottlenecks, and the same fear of releasing. Kanban doesn&#8217;t give you a deployment pipeline. It doesn&#8217;t write your tests. It doesn&#8217;t decouple your architecture.</p><p>Conversely, if your team can do continuous delivery, Scrum doesn&#8217;t prevent it. Deploy whenever you want. The sprint planning still gives you a rhythm for deciding what to work on. The Sprint Review still gives you a moment to inspect what was delivered. The retrospective still gives you a moment to improve.</p><p>Continuous delivery is orthogonal to your process framework. It&#8217;s an engineering practice, and it works with any framework you choose, provided you invest in the technical foundations that make it possible.</p><h2>AI changes the speed of writing code, not the nature of risk</h2><p>The latest twist in the &#8220;sprints are dead&#8221; argument is AI. The claim goes something like this: AI generates code so fast that delivery is no longer the bottleneck, so we don&#8217;t need sprints anymore. We need &#8220;context&#8221; and &#8220;discovery&#8221; instead.</p><p>Let me be direct about this. This argument confuses writing code with delivering software, and it confuses speed with risk.</p><p>Delivery has never been primarily about typing. The time spent writing code has always been a small fraction of the total cost of delivering software. Most of the time goes to understanding what needs to be built, designing a solution, testing it, integrating it with existing systems, deploying it, monitoring it, and supporting it in production. AI can accelerate the coding part. It doesn&#8217;t eliminate the rest.</p><p>But more importantly, the fundamental challenge of software development hasn&#8217;t changed. It&#8217;s still risk. Will this feature solve the user&#8217;s actual problem? Will it work correctly under real conditions? Will it integrate with existing systems without breaking anything? Will it scale? Will it be maintainable? Will it be secure? Will it comply with regulations?</p><p>None of these risks are reduced by generating code faster. If anything, they&#8217;re amplified. The faster you can generate code, the faster you can build the wrong thing. The faster you can introduce defects. The faster you can accumulate technical debt. Speed without feedback is not a capability. It&#8217;s a hazard.</p><p>This is exactly why iterative development exists. Not because coding is slow, but because understanding is incomplete. You build a small thing, put it in front of users, learn from the feedback, and adjust. The speed of the coding step is largely irrelevant to this cycle. What matters is the speed of learning.</p><p>AI doesn&#8217;t change the need for feedback loops. AI doesn&#8217;t change the need for WIP limits. AI doesn&#8217;t change the need for regular retrospection. AI doesn&#8217;t change the need for explicit process policies. AI doesn&#8217;t change the need to manage flow. AI doesn&#8217;t change the fact that the biggest risk in software development is building something nobody wants, or building something that doesn&#8217;t work in production, or building something that can&#8217;t be maintained.</p><p>What AI does is make the consequences of poor practices worse. If your feedback loops are weak, you&#8217;ll now generate wrong solutions faster. If your requirements process is broken, you&#8217;ll now build the wrong thing in minutes instead of weeks. If your testing is inadequate, you&#8217;ll now deploy defective code more frequently. AI amplifies whatever practices you already have, good or bad.</p><p>The teams that will benefit most from AI are the teams that already have strong practices: clear requirements processes, fast feedback loops, comprehensive automated testing, continuous integration, and the discipline to validate before they ship. These are the teams that can use AI to accelerate the coding step without introducing additional risk, because the rest of their process catches the errors.</p><p>The teams that will suffer most are the teams that think AI means they can skip the hard parts. The teams that think &#8220;context&#8221; replaces testing. The teams that think discovery can be done entirely upfront instead of continuously. The teams that think generating code fast means delivering value fast.</p><p>Risk management in software development is the same problem it&#8217;s always been. You reduce risk through short feedback cycles, small batch sizes, continuous validation, and close collaboration with the people who will use the software. Sprints do this. Kanban&#8217;s WIP limits and flow management do this. Both frameworks, when properly implemented, are risk management tools.</p><p>AI doesn&#8217;t change what risk is. It just changes how fast you can create it.</p><h2>The real question</h2><p>If you&#8217;re struggling with Scrum, the answer is not to switch to Kanban. If you&#8217;re struggling with Kanban, the answer is not to switch to Scrum. The answer is to look honestly at what the framework is showing you and address the real problems.</p><p>Is your planning painful because nobody understands the requirements? Fix the requirements process. Are your retrospectives useless because people don&#8217;t feel safe? Build psychological safety. Is your delivery slow because of technical debt? Invest in engineering practices. Are your WIP limits constantly overridden because management keeps pushing work in? Have an honest conversation about capacity.</p><p>The framework is the mirror. The dysfunction is yours. Changing the mirror won&#8217;t change what it&#8217;s reflecting.</p><p>Stop debating frameworks. Start fixing the problems the frameworks are showing you.</p>]]></content:encoded></item><item><title><![CDATA[From Story Points to Reliable Planning: A Practical Guide for Teams Ready to Stop Guessing ]]></title><description><![CDATA[Story points have become so ubiquitous in software development that many teams assume they&#8217;re a fundamental part of agile.]]></description><link>https://a4al6a.substack.com/p/from-story-points-to-reliable-planning</link><guid isPermaLink="false">https://a4al6a.substack.com/p/from-story-points-to-reliable-planning</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Tue, 27 Jan 2026 23:17:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Story points have become so ubiquitous in software development that many teams assume they&#8217;re a fundamental part of agile. They&#8217;re not. Story points were invented as a tool to help teams have conversations about work, but somewhere along the way they became a ritual that consumes time, creates false confidence, and rarely delivers the predictability teams actually need.</p><p>If you&#8217;ve ever sat through a planning session where the team spent twenty minutes debating whether something is a 5 or an 8, you&#8217;ve experienced the dysfunction firsthand. That time could have been spent understanding the work, identifying risks, or actually delivering something.</p><p>This article is a practical guide for teams who want to move beyond story points toward planning approaches that are simpler, faster, and grounded in reality rather than collective guessing.</p><h2>Why story points fail</h2><p>The theory behind story points sounds reasonable: estimate relative complexity rather than time, use historical velocity to forecast, and avoid the trap of conflating estimates with commitments. In practice, this rarely works as intended.</p><p>The first problem is that story points mean different things to different people, even within the same team. One developer thinks about complexity. Another thinks about effort. A third is secretly converting to hours in their head and then picking a Fibonacci number. When you average these different mental models together, you don&#8217;t get wisdom of crowds. You get noise.</p><p>The second problem is that velocity, the sum of story points completed per sprint, is treated as if it were a stable measure when it isn&#8217;t. Velocity fluctuates based on who&#8217;s on holiday, how many meetings interrupt the sprint, whether the work was estimated by the same people who did it, and countless other factors. Teams often respond to this instability by adding more process: re-estimation, calibration sessions, reference stories. None of this makes the underlying measure more meaningful.</p><p>The third problem is that story points create perverse incentives. When velocity becomes a performance metric, teams unconsciously inflate their estimates. A 3 becomes a 5. An 8 becomes a 13. Velocity goes up, but throughput stays the same. Everyone pretends not to notice.</p><p>The final problem is opportunity cost. The time spent estimating is time not spent understanding, slicing, or delivering. A team that spends two hours per sprint on estimation and calibration is spending over a hundred hours per year on an activity that doesn&#8217;t move the product forward.</p><h2>What actually predicts delivery</h2><p>If story points don&#8217;t work, what does? The answer is surprisingly simple: count things.</p><p>Throughput, the number of items a team completes per unit of time, is a far more stable and useful measure than velocity. This works because throughput is based on observation rather than prediction. You&#8217;re not asking &#8220;how hard do we think this will be?&#8221; You&#8217;re asking &#8220;how many things did we actually finish?&#8221;</p><p>For throughput to be a reliable predictor, one condition must be met: the items being counted need to be roughly similar in size. This doesn&#8217;t mean identical, just within the same order of magnitude. If most of your work takes one to three days to complete, counting items gives you a useful forecast. If your backlog contains a mix of half-day tasks and three-week epics, counting won&#8217;t tell you much.</p><p>This is where the real skill comes in. The shift from story points to throughput-based planning is really a shift from estimating to slicing. Instead of asking &#8220;how big is this?&#8221; you ask &#8220;how do we make this small enough that size doesn&#8217;t matter?&#8221;</p><h2>The art of slicing work small</h2><p>Slicing is the most valuable planning skill a team can develop, and it&#8217;s almost entirely ignored in mainstream agile training. The ability to take a large, ambiguous piece of work and break it into thin vertical slices that each deliver value is what separates high-performing teams from the rest.</p><p>A good slice has three properties. First, it delivers something a user or stakeholder can see, use, or give feedback on. Second, it can be completed independently, without waiting for other slices. Third, it&#8217;s small enough to finish in a day or two of focused work.</p><p>The third property is the critical one for planning purposes. When everything in your backlog is one to three days of work, the difference between items becomes negligible. A two-day item and a three-day item are close enough that counting them as equivalent introduces less error than trying to estimate them separately.</p><p>Teams often resist slicing this small because it feels unnatural. &#8220;We can&#8217;t deliver anything useful in two days,&#8221; they say. This is almost never true. What&#8217;s actually happening is that the team has become accustomed to thinking in terms of technical tasks rather than user value. A &#8220;user story&#8221; that says &#8220;implement the payment gateway&#8221; isn&#8217;t a story at all. It&#8217;s a technical component. A real slice might be &#8220;a customer can pay for a single item using a saved card&#8221; or even &#8220;a customer sees a payment button that shows a coming soon message.&#8221; Both of these deliver something observable. Both can be done in a day or two. Both give you something to learn from.</p><p>Learning to slice well takes practice, and the best way to practice is to do it together as a team. Every time someone brings a large item to planning, treat it as an opportunity to slice collaboratively. Ask: what&#8217;s the smallest thing we could deliver that would let us learn something? What could we ship that a user would actually notice? What&#8217;s the riskiest part of this, and how could we test that assumption with a thin slice?</p><h2>How to run planning without estimates</h2><p>Once your team has embraced small slices and started tracking throughput, planning becomes remarkably simple.</p><p>Before the session, gather your historical data. How many items has the team completed in each of the last eight to ten sprints? Calculate the average and note the range. If you&#8217;ve been completing between six and ten items per sprint, with an average of eight, that&#8217;s your forecast baseline.</p><p>Start the planning session by reviewing the team&#8217;s throughput data together. This isn&#8217;t about judgement or performance management. It&#8217;s about grounding the conversation in reality. &#8220;Based on our history, we typically complete around eight items per sprint. Sometimes it&#8217;s six, sometimes it&#8217;s ten. Let&#8217;s plan accordingly.&#8221;</p><p>Then work through the backlog in priority order. For each item, ask three questions. Do we understand what done looks like for this? Is this small enough to complete in a couple of days? Are there any blockers, dependencies, or risks we need to address?</p><p>If the answer to the second question is no, stop and slice. This is the most important part of the session. Don&#8217;t let large items into the sprint. Every large item is a forecast risk and a flow impediment.</p><p>Keep pulling items until you&#8217;ve reached your typical throughput number. If your average is eight, stop at eight or perhaps nine. Resist the temptation to overcommit because this sprint &#8220;feels different.&#8221; It doesn&#8217;t. The whole point of using historical data is to protect you from optimism bias.</p><p>That&#8217;s it. No poker. No Fibonacci. No debates about whether complexity and effort should be weighted differently. Just a focused conversation about understanding the work and making sure it&#8217;s small enough to flow.</p><h2>Answering the objections</h2><p>When you propose this approach, you&#8217;ll face pushback. Here&#8217;s how to address the most common objections.</p><p>&#8220;How will we know if we&#8217;ve planned the right amount of work?&#8221; You&#8217;ll know the same way you know now: by comparing what you planned to what you delivered. The difference is that throughput-based planning gives you an honest forecast based on measurement rather than a confident-sounding number based on guessing. If you consistently complete eight items, planning for eight items is a reasonable bet. If you consistently complete somewhere between six and ten, acknowledge that range rather than pretending you can predict exactly.</p><p>&#8220;What about items that are genuinely complex and can&#8217;t be sliced smaller?&#8221; This is almost always a failure of imagination rather than a hard constraint. I&#8217;ve worked with teams building safety-critical systems, complex financial products, and intricate distributed architectures. In every case, we found ways to slice work small. The technique varies depending on context, but the principle holds: there&#8217;s always a thinner slice that still delivers something real. If you truly cannot slice something smaller, that&#8217;s a signal that you don&#8217;t yet understand the work well enough. Do a spike. Build a prototype. Run an experiment. Don&#8217;t commit to delivering something you can&#8217;t decompose.</p><p>&#8220;Stakeholders expect velocity reports and burn-down charts.&#8221; This is a change management challenge, not a technical one. Most stakeholders don&#8217;t actually care about velocity. They care about predictability: when will this be done, and can we count on that date? You can answer those questions more honestly with throughput data and cycle time distributions than with velocity. Have a conversation with your stakeholders about what information they actually need and why. Often they&#8217;re relieved to stop pretending the current reports mean something.</p><p>&#8220;Different team members work at different speeds, so counting items doesn&#8217;t account for who picks up what.&#8221; Neither do story points. You&#8217;re measuring team throughput, not individual productivity. Over time, the mix of who does what averages out. If it doesn&#8217;t, you have a team design problem that no estimation method will solve.</p><p>&#8220;We need estimates for roadmap planning and budgeting.&#8221; Throughput data supports roadmap planning better than story points. If you have fifty items in the backlog and you complete eight per sprint, you can forecast six to seven sprints of work with appropriate confidence intervals. This is more honest than converting everything to story points and dividing by velocity, which gives you a single number that implies false precision.</p><h2>Making the transition</h2><p>If you&#8217;re convinced this approach is worth trying, here&#8217;s how to introduce it without causing chaos.</p><p>Start by gathering data. Even if you&#8217;re still using story points, begin tracking throughput alongside velocity. After a few sprints, you&#8217;ll be able to show the team how throughput compares in terms of stability and predictability.</p><p>Next, invest in slicing skills. Run workshops on vertical slicing. Practice breaking down real backlog items together. Make slicing a core part of your refinement sessions. This is the foundation that makes everything else work.</p><p>Then propose an experiment. Suggest trying throughput-based planning for three sprints. Frame it as a learning exercise rather than a permanent change. This reduces resistance and gives sceptics a face-saving way to engage.</p><p>During the experiment, facilitate planning sessions that focus on understanding and slicing rather than estimating. Keep the conversations grounded in &#8220;is this small enough?&#8221; rather than &#8220;how big is this?&#8221;</p><p>After three sprints, retrospect together. Was planning faster? Did the team deliver roughly what they forecast? Did the conversations during planning feel more useful? In my experience, teams rarely want to go back once they&#8217;ve experienced the simplicity of this approach.</p><h2>The deeper shift</h2><p>Moving away from story points isn&#8217;t just a process change. It&#8217;s a shift in mindset from prediction to measurement, from estimation to understanding, from big batches to small slices.</p><p>This shift has benefits beyond planning. Small slices improve flow, reduce risk, accelerate feedback, and make continuous integration actually possible. Teams that slice well deliver more frequently and learn faster. They spend less time in meetings and more time shipping.</p><p>Story points were never the point. The point was always to deliver valuable software sustainably. If your current approach to planning isn&#8217;t helping you do that, it might be time to try something simpler.</p><div><hr></div><h2>Further Reading</h2><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;9125c3db-9e1d-4219-82b0-c03dcf9274f4&quot;,&quot;caption&quot;:&quot;&#8220;Story points aren&#8217;t about the numbers. They trigger conversations about effort and complexity.&#8221;&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The Story Point Illusion&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:378940401,&quot;name&quot;:&quot;Andrea Laforgia&quot;,&quot;bio&quot;:&quot;Software Engineer, Tech Lead &amp; Architect with 30+ years experience. Expert in design, development, testing, architecture &amp; team collaboration. Lean &amp; XP advocate using TDD, pairing/mobbing, Continuous Integration, and Continuous Delivery.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4d1ea449-01a9-4dac-b00b-30b933b4062c_690x690.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-01-15T18:56:59.391Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!ANa9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://a4al6a.substack.com/p/the-story-point-illusion&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:184682075,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:15,&quot;comment_count&quot;:0,&quot;publication_id&quot;:5917094,&quot;publication_name&quot;:&quot;Andrea&#8217;s Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!JZMQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;cafe6afe-8df3-4a31-94bf-2928e987c06f&quot;,&quot;caption&quot;:&quot;TL;DR Software estimation has a terrible track record. Average cost overruns of 189% and time overruns of 222% aren&#8217;t anomalies; they&#8217;re the norm. But the problem isn&#8217;t that the future is unknowable. The problem is that organisations punish honesty about uncertainty.&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The Honest Estimation Problem: Why Software Forecasting Fails and What Actually Helps&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:378940401,&quot;name&quot;:&quot;Andrea Laforgia&quot;,&quot;bio&quot;:&quot;Software Engineer, Tech Lead &amp; Architect with 30+ years experience. Expert in design, development, testing, architecture &amp; team collaboration. Lean &amp; XP advocate using TDD, pairing/mobbing, Continuous Integration, and Continuous Delivery.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4d1ea449-01a9-4dac-b00b-30b933b4062c_690x690.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-01-20T19:20:05.202Z&quot;,&quot;cover_image&quot;:null,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://a4al6a.substack.com/p/the-honest-estimation-problem-why&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:185128053,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:9,&quot;comment_count&quot;:1,&quot;publication_id&quot;:5917094,&quot;publication_name&quot;:&quot;Andrea&#8217;s Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!JZMQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[What if AI helps us build anti-fragile systems ]]></title><description><![CDATA[On non-determinism, intent, and software that stays alive]]></description><link>https://a4al6a.substack.com/p/what-if-ai-helps-us-build-anti-fragile</link><guid isPermaLink="false">https://a4al6a.substack.com/p/what-if-ai-helps-us-build-anti-fragile</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Sat, 24 Jan 2026 01:03:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The sheer power of AI tools, even with all their non-determinism, feels like a massive opportunity. Not just to build systems with better quality, but to build systems that are actually anti-fragile.<br><br>A lot of the worries we have today about AI agents and reliability seem very influenced by the world we come from. Computing has been dominated by determinism for decades. Same input, same output, otherwise something is wrong. That way of thinking made sense, but it also really limits how we imagine what comes next. We are notoriously bad at this as humans. We almost always picture the future as a slightly tweaked version of the present.<br><br>There are plenty of famous examples. The 640K memory quote often linked to Bill Gates. The idea that nobody would ever need a telephone. Someone important confidently explaining that television would never really take off. Whether the quotes are perfectly accurate or not almost does not matter. The pattern does. We mistake current constraints for universal truths.<br><br>If we are willing to accept that the same behaviour can come from different internal paths, even with the same input, things start to look different. Variability stops being purely a liability. With the right adversarial constraints, it can become a strength. Think of approaches like the <a href="https://nwave.ai/">nWave</a> framework that my friends <a href="https://www.linkedin.com/in/alessandro-di-gioia/">Alessandro</a> and <a href="https://www.linkedin.com/in/michelebrissoni/">Michele</a> have been developing, where challenge and review are deliberately built into the system.<br><br>Set up like that, a system can absorb internal deviations and also cope with the world changing underneath it. That already feels like more than just being robust.<br><br>This is where AI feels most interesting to me. Not as something we use once to generate code or configuration, but as something we inject into the system itself. Almost like the lymph of a living organism. It keeps flowing, reacting to events, while staying within some expected boundaries of behaviour.<br><br>In that world, every system becomes an expert system about its own domain and about the environment it operates in. We set intent and constraints, but the product stays alive. It adapts, reacts to what actually happens, and feeds back to us what no longer makes sense, including where our original assumptions or specs are wrong.<br><br>This might all be a bit wild. It definitely needs more thinking and probably a lot of pushback. But it feels like a direction worth exploring, especially if we stop judging AI purely through a deterministic lens and start designing for change rather than trying to eliminate it.</p>]]></content:encoded></item><item><title><![CDATA[The Honest Estimation Problem: Why Software Forecasting Fails and What Actually Helps]]></title><description><![CDATA[The real enemy isn&#8217;t estimation itself. It&#8217;s false precision and the organisational dysfunction that demands it.]]></description><link>https://a4al6a.substack.com/p/the-honest-estimation-problem-why</link><guid isPermaLink="false">https://a4al6a.substack.com/p/the-honest-estimation-problem-why</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Tue, 20 Jan 2026 19:20:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>TL;DR</h2><p>Software estimation has a terrible track record. Average cost overruns of 189% and time overruns of 222% aren&#8217;t anomalies; they&#8217;re the norm. But the problem isn&#8217;t that the future is unknowable. The problem is that organisations punish honesty about uncertainty.</p><p><strong>The core claim:</strong> Estimation fails not because prediction is impossible, but because most organisations can&#8217;t handle honest answers. They demand single numbers, treat those numbers as commitments, and blame people when reality diverges from the guess. No methodology fixes that.</p><p><strong>What doesn&#8217;t work:</strong> Traditional point estimates pretend certainty where none exists. Story points get gamed the moment they become performance metrics. Even probabilistic forecasting fails if the organisation punishes you for missing the 75th percentile instead of the point estimate.</p><p><strong>What actually helps:</strong> Track what happens (cycle time, throughput). Present ranges with confidence levels. Keep work items small, genuinely small, using deliberate slicing techniques. Treat estimates as decision-support, not promises. Most importantly, build organisational tolerance for uncertainty. If your culture punishes honest answers, no methodology will save you.</p><p><strong>The bottom line:</strong> The question isn&#8217;t &#8220;how do we estimate better?&#8221; The question is &#8220;how do we build organisations mature enough to handle honest answers about uncertainty?&#8221; That&#8217;s harder than switching methodologies, but it&#8217;s the only thing that actually works.</p><div><hr></div><h2>The terminology trap</h2><p><em>In brief: Organisations routinely conflate estimates, targets, and commitments. This vocabulary confusion is where dysfunction begins.</em></p><p>Before examining why estimation fails, we need to untangle a vocabulary problem that poisons most estimation conversations.</p><p>Steve McConnell, in <em>Software Estimation: Demystifying the Black Art</em>, identifies three concepts that organisations routinely conflate:</p><p>An <strong>estimate</strong> is an analytical prediction of how long something will take, based on available information. It&#8217;s a forecast, not a promise. It should include uncertainty ranges and will change as information improves.</p><p>A <strong>target</strong> is a business goal or desired deadline. It represents what the organisation wants to achieve, independent of whether it&#8217;s realistic.</p><p>A <strong>commitment</strong> is a promise to deliver by a specific date. It carries accountability and should only be made when uncertainty has narrowed enough that the risk is acceptable.</p><p>The dysfunction begins when these get confused. A developer saying &#8220;this looks like maybe six weeks&#8221; intends an estimate, a preliminary forecast based on incomplete information. Management hears a commitment. The roadmap records a target. When reality diverges, everyone feels betrayed despite never having agreed on what the original number meant.</p><p>This isn&#8217;t mere pedantry. McConnell argues that consciously separating these concepts is a prerequisite for healthy estimation conversations. An estimate can inform whether a target is realistic. A commitment should only follow once estimates have been refined sufficiently that the uncertainty band becomes acceptable for the business risk involved.</p><p>Most organisational dysfunction around estimation traces back to this confusion. When executives treat preliminary estimates as commitments, or when targets are disguised as &#8220;just estimates,&#8221; the game is rigged before it starts.</p><div><hr></div><h2>The uncomfortable track record</h2><p><em>In brief: The statistics are bad, but often overstated. The real problem isn&#8217;t prediction itself; it&#8217;s what organisations do with predictions.</em></p><p>Let&#8217;s be honest about the numbers, and careful about what they actually mean.</p><p>Bent Flyvbjerg&#8217;s research on over 16,000 projects found that only 0.5% deliver their promised benefits within budget and timeframe. The Standish Group&#8217;s CHAOS Report shows roughly 16% of projects succeed on time, on budget, with all features. About 53% are challenged, meaning over budget, over time, or with fewer features. The remaining 31% are cancelled outright.</p><p>These numbers are often cited to damn estimation itself, but we should be careful. Flyvbjerg&#8217;s research focuses heavily on large infrastructure and megaprojects: airports, railways, Olympic venues. These domains have specific pathologies, including strategic misrepresentation (the polite term for lying to get projects approved), that don&#8217;t translate cleanly to typical software teams. A SaaS product built by a startup is not an airport.</p><p>More importantly, Flyvbjerg&#8217;s 0.5% figure measures whether projects deliver their promised <em>benefits</em>, not whether they hit their estimates. A project can hit its estimates perfectly and still fail to deliver value because the wrong thing was built. Conversely, a project can overrun significantly and still be wildly successful. Amazon Web Services famously launched years late and over budget, but it would be absurd to call it a failure.</p><p>That said, the track record for estimation accuracy in software specifically is still poor. Average cost overruns of 189% and time overruns of 222% aren&#8217;t anomalies; they&#8217;re common. Something is systematically wrong. But importing the worst-case statistics from megaprojects and presenting them as universal software truth overstates the case.</p><p>McConnell adds another sobering observation: projects hardly ever remain stable regarding the assumptions that affect estimates. Staff departs. Priorities shift. Budgets shrink. Markets change. The estimate made at the start described a project that no longer exists by the middle.</p><div><hr></div><h2>Why your brain lies to you</h2><p><em>In brief: Cognitive biases make accurate estimation genuinely hard. The planning fallacy is robust across cultures, personality types, and experience levels. Even knowing about it doesn&#8217;t prevent it.</em></p><p>In 1979, Daniel Kahneman and Amos Tversky identified what they called the planning fallacy: the tendency to underestimate time, costs, and risks while overestimating benefits. This isn&#8217;t a bug in human cognition; it&#8217;s a feature. And it&#8217;s remarkably robust. It appears for small tasks and massive infrastructure projects. It generalises across personality types and cultures. It affects individuals and groups equally. Most troublingly, even knowing about the bias doesn&#8217;t prevent it.</p><p>When you estimate a software project, your brain deploys an arsenal of cognitive shortcuts. Optimism bias makes you believe you&#8217;re less likely to experience problems than others. Anchoring bias means the first number mentioned warps all subsequent thinking. The availability heuristic causes you to weight recent, memorable experiences more heavily than historical data. Self-serving bias leads you to take credit when things go well but blame external factors when they don&#8217;t, preventing you from learning.</p><p>Kahneman distinguished between the &#8220;inside view&#8221; and the &#8220;outside view&#8221;. The inside view focuses on a project&#8217;s specific characteristics: the features, the technology, the team. This feels like the right approach, but it systematically produces overconfident estimates. The outside view instead treats your project as one instance in a class of similar efforts and asks how those similar efforts actually performed. Kahneman called the outside view &#8220;the single most important piece of advice regarding how to increase accuracy in forecasting&#8221;.</p><p>This insight was formalised as Reference Class Forecasting: rather than estimating your project based on its unique characteristics, identify a reference class of similar past projects and use their actual outcomes as your baseline. The technique won Kahneman the Nobel Prize in Economics and has been mandated by the UK Treasury for large public projects. Yet software teams rarely implement it systematically, instead relying on expert judgment that demonstrably produces optimistic bias.</p><p>Ron Jeffries makes a related point: teams typically estimate &#8220;at the moment of maximum ignorance,&#8221; when understanding of what needs to be built is lowest. The estimates that get locked in as commitments are precisely the ones made when the team knew least about the work.</p><div><hr></div><h2>The Cone of Uncertainty: theory meets reality</h2><p><em>In brief: McConnell&#8217;s Cone describes how uncertainty narrows on well-managed projects. The caveat is important: it only works when teams already practice healthy project management.</em></p><p>Steve McConnell popularised the Cone of Uncertainty as a framework for understanding estimation accuracy. The concept is elegant: early in a project, uncertainty spans a wide range (estimates may be off by 4x in either direction). As work progresses and unknowns resolve, the cone narrows. Variability decreases from 4x to 2x to 1.25x as projects move through requirements, design, and implementation phases.</p><p>The Cone offers useful insights. It legitimises early imprecision: demanding precise estimates during initial scoping ignores the fundamental reality that key decisions haven&#8217;t yet been made. It justifies iterative replanning: as information accumulates, estimates can legitimately narrow. Organisations that lock in early estimates and refuse updates are fighting mathematical reality.</p><p>But the critical caveat undermines the Cone as prescription. McConnell himself emphasises that the Cone describes what happens on <em>well-managed</em> projects. He writes: &#8220;It&#8217;s easily possible to do worse.&#8221; Many teams don&#8217;t systematically attack their highest sources of variability early. They postpone difficult problems, creating false certainty that shatters late in the project.</p><p>If the framework only works when teams already practice healthy project management, it describes a symptom of good practice rather than providing a path toward it. Teams struggling with estimation dysfunction can&#8217;t simply &#8220;follow the Cone.&#8221; They must first solve the organisational problems that prevent uncertainty from narrowing naturally.</p><div><hr></div><h2>The story points trap</h2><p><em>In brief: Story points fail not because relative estimation is flawed, but because organisations misuse them. Even the practice&#8217;s inventor now expresses regret.</em></p><p>When Agile emerged, story points and velocity were supposed to solve the estimation problem. Abstract away from time. Focus on relative sizing. Let the data tell you how much you can deliver.</p><p>In theory, elegant. In practice, often dysfunctional.</p><p>The moment velocity becomes a performance metric, teams start gaming it. Story point inflation becomes rampant. Different teams have different reference scales, making cross-team comparisons meaningless, yet organisations constantly try to compare them. When hitting velocity targets matters more than sustainable delivery, teams cut corners. Technical debt accumulates. And despite story points being designed to avoid time estimation, stakeholders want dates, so organisations create conversion factors that destroy whatever value story points might have provided.</p><h3>The inventor&#8217;s second thoughts</h3><p>Perhaps the most striking critique of story points comes from Ron Jeffries himself, one of the original Extreme Programming founders often credited with inventing the practice. In 2012, Jeffries wrote:</p><p>&#8220;A team that is focusing on velocity is not focusing on value. I wish I had never invented velocity, if in fact I did.&#8221;</p><p>This isn&#8217;t nostalgic regret about misuse. Jeffries argues that velocity focuses on &#8220;the short end of the lever,&#8221; optimising the cost side of the equation when what matters in Agile is <em>steering</em> by selecting what to do and what to defer. He observed teams being measured on estimate accuracy, with product owners reduced to &#8220;just following the plan&#8221; rather than genuinely prioritising value.</p><p>The creator&#8217;s disillusionment doesn&#8217;t invalidate story points for all contexts. But it suggests that even under ideal conditions, the practice may misdirect attention. When your estimation system&#8217;s architect concludes it optimises the wrong variable, perhaps the problem runs deeper than organisational dysfunction.</p><h3>The broader pattern</h3><p>This matters because the same critique applies to any metric, including the flow metrics and throughput measures that forecasting advocates prefer. Throughput can be gamed too. Cycle time can be manipulated by how you define &#8220;started&#8221; and &#8220;done&#8221;. If leaders misuse metrics, changing metrics won&#8217;t fix leadership.</p><p>The Agile Manifesto&#8217;s seventh principle states: &#8220;Working software is the primary measure of progress.&#8221; Not story points. Not velocity. Not throughput. Working software. Any metric is a proxy, and any proxy can be gamed.</p><div><hr></div><h2>Forecasting: changing the conversation</h2><p><em>In brief: Forecasting is still estimation. The value isn&#8217;t ontological (it doesn&#8217;t make prediction possible). The value is behavioural (it changes how estimates are discussed).</em></p><p>Around 2011, practitioners like Vasco Duarte, Neil Killick, and Woody Zuill started questioning the estimation orthodoxy under the hashtag #NoEstimates. They drew a distinction between estimating (giving a general idea based on opinion) and forecasting (calculating predictions based on historical data).</p><p>This distinction is useful, but we should be clear about what it actually changes. Forecasting is still estimation. When you run a Monte Carlo simulation and say &#8220;75% chance of completion by October 22nd&#8221;, you&#8217;re still making a prediction about the future. You&#8217;ve wrapped that prediction in probability language, making uncertainty explicit rather than hidden. That&#8217;s genuinely valuable, but it&#8217;s not magic.</p><p>The real value of probabilistic forecasting isn&#8217;t that it makes prediction possible where it wasn&#8217;t before. It&#8217;s that it changes the conversation. It&#8217;s harder to treat a probability distribution as a commitment. It forces discussions about risk tolerance. It makes uncertainty visible and therefore discussable.</p><p>Monte Carlo simulation, invented by John von Neumann and Stanislaw Ulam during World War II, uses repeated random sampling to model uncertain outcomes. For software delivery, you gather historical throughput data (how many items your team completes per week) and cycle time data (how long items take from start to finish). You then run thousands of simulations, randomly sampling from past performance to project possible futures. Instead of a single date, you get a probability distribution: 50% chance by October 15, 75% by October 22, 90% by November 1.</p><p>Daniel Vacanti puts it well: &#8220;If you want to get stuff done by August 31st, sure, we can get stuff done by August 31st, but there&#8217;s about a 40% chance of that happening. Are you okay taking that 60% risk? Now we can have a smarter, more adult economic conversation about what risk we&#8217;re willing to take.&#8221;</p><div><hr></div><h2>What honest estimation looks like</h2><p><em>In brief: A concrete example of the difference between dysfunctional and healthy estimation conversations.</em></p><p>The difference between bad and good estimation isn&#8217;t the technique. It&#8217;s the conversation. Here&#8217;s what the same situation looks like handled badly versus handled well.</p><p><strong>The dysfunctional version:</strong></p><p>Product Manager: &#8220;When will the new payment integration be done?&#8221;</p><p>Tech Lead: &#8220;Hard to say exactly. There&#8217;s a lot of uncertainty around the third-party API.&#8221;</p><p>Product Manager: &#8220;I need a date for the roadmap.&#8221;</p><p>Tech Lead: &#8220;Um... maybe six weeks?&#8221;</p><p>Product Manager: &#8220;Great, I&#8217;ll put it down for March 15th.&#8221;</p><p><em>Four weeks later:</em></p><p>Product Manager: &#8220;We&#8217;re halfway through and you&#8217;re saying it might slip? You committed to March 15th.&#8221;</p><p>Tech Lead: &#8220;I said &#8216;maybe six weeks.&#8217; And we&#8217;ve discovered the API doesn&#8217;t support batch operations, so we need to redesign.&#8221;</p><p>Product Manager: &#8220;This is going to be a difficult conversation with leadership. They&#8217;re expecting March 15th.&#8221;</p><p><strong>The healthy version:</strong></p><p>Product Manager: &#8220;When will the new payment integration be done?&#8221;</p><p>Tech Lead: &#8220;Based on similar integrations we&#8217;ve done, I&#8217;d say 50% chance we&#8217;re done in four weeks, 75% chance by six weeks, 90% by eight weeks. The big unknown is the third-party API. If it&#8217;s well-documented and supports our use cases, we&#8217;re at the fast end. If we hit surprises, we&#8217;re at the slow end.&#8221;</p><p>Product Manager: &#8220;Leadership wants it for the March launch. That&#8217;s six weeks away.&#8221;</p><p>Tech Lead: &#8220;So we&#8217;re looking at roughly 75% confidence for that date. What&#8217;s the cost if we miss it? Is there a fallback plan?&#8221;</p><p>Product Manager: &#8220;The launch can proceed without it, but it&#8217;s a headline feature. Missing it would be disappointing, not catastrophic.&#8221;</p><p>Tech Lead: &#8220;Then let&#8217;s proceed, but I&#8217;ll flag it early if we hit the API problems. If we&#8217;re in trouble by week three, we&#8217;ll know, and you&#8217;ll have time to adjust messaging.&#8221;</p><p><em>Four weeks later:</em></p><p>Tech Lead: &#8220;The API doesn&#8217;t support batch operations. We&#8217;re now tracking toward the six-to-eight week range. I wanted to let you know with two weeks still to go.&#8221;</p><p>Product Manager: &#8220;Thanks for the early warning. I&#8217;ll talk to leadership about plan B for the launch.&#8221;</p><p>The technique (Monte Carlo, throughput tracking, whatever) matters less than what&#8217;s happening in that conversation: ranges instead of points, explicit confidence levels, discussion of risk tolerance, early warning when forecasts change, and no blame when reality differs from prediction.</p><div><hr></div><h2>The assumptions behind forecasting</h2><p><em>In brief: Forecasting requires stable throughput, comparable work items, and sufficient history. When these don&#8217;t hold, forecasts are unreliable, and you should say so.</em></p><p>Monte Carlo forecasting isn&#8217;t magic. It rests on assumptions that often go unstated.</p><p>It assumes reasonably stable throughput, meaning your team&#8217;s capacity tomorrow will resemble its capacity yesterday. It assumes comparable work items, meaning the things you&#8217;re forecasting are similar in nature to the things in your historical data. It assumes sufficient historical data to sample from. And it assumes no major structural changes to team composition, domain, or architecture.</p><p>These conditions often don&#8217;t hold. Greenfield products have no history. Early startups pivot constantly. Platform rewrites involve genuinely novel technical challenges. Regulated environments have external constraints that historical data can&#8217;t capture. Teams undergoing rapid scaling have unstable throughput by definition.</p><p>The situations where estimation is hardest are precisely the situations where you have the least relevant historical data to sample from. A team that&#8217;s been delivering similar features for two years can forecast with reasonable confidence. A team building something genuinely new, in a domain they don&#8217;t know well, with technology they haven&#8217;t used before, has no meaningful reference class.</p><p>For teams without historical data, the honest answer is: your forecasts will be unreliable, and you should say so. Start collecting data immediately, make small commitments, and let your forecasting improve as your history grows. Pretending you can forecast accurately without data isn&#8217;t forecasting; it&#8217;s just the old estimation game with fancier vocabulary.</p><div><hr></div><h2>Cognitive biases don&#8217;t disappear with data</h2><p><em>In brief: Forecasting moves bias into different places (how work is sliced, which history is included) rather than eliminating it. Discipline still required.</em></p><p>The argument that estimation suffers from cognitive biases while forecasting doesn&#8217;t is too clean. The same biases that corrupt our estimates also affect how we collect and interpret historical data.</p><p>Forecasting moves bias into different places rather than eliminating it. Bias affects how work is sliced: teams often break down work in ways that make historical comparisons favourable. Bias affects which history is included: it&#8217;s tempting to exclude that disaster project as an &#8220;outlier&#8221; that doesn&#8217;t represent normal performance. Bias affects how outliers are treated: do you cap them, exclude them, or let them skew your distribution?</p><p>Teams tend to remember successes more vividly than failures. They categorise work items in ways that flatter past performance. The person choosing which historical items to include in the reference class brings all their biases with them.</p><p>This doesn&#8217;t mean data-driven forecasting is worthless. It just means it&#8217;s not the silver bullet it&#8217;s sometimes presented as. Good forecasting requires discipline: consistent categorisation, honest inclusion of failures, resistance to cherry-picking. The methodology helps, but it doesn&#8217;t eliminate the need for intellectual honesty.</p><div><hr></div><h2>The small stories alternative</h2><p><em>In brief: If work items are small enough, estimation becomes almost unnecessary. You count items and measure throughput. The challenge lies in achieving that granularity.</em></p><p>Ron Jeffries argues that if work items are small enough, estimation becomes almost unnecessary. When everything takes roughly a day, you don&#8217;t need sophisticated forecasting. You count items and measure throughput. The challenge lies in achieving that granularity.</p><h3>Technique 1: Single Acceptance Test Method</h3><p>Neil Killick, cited by Jeffries, proposes examining acceptance criteria and implementing them one at a time, starting with the simplest. Rather than estimating &#8220;User can pay for order&#8221; as a single story, break it into separate work items:</p><p>Display payment button (simplest). Accept credit card number. Validate card format. Process test transaction. Handle payment failure gracefully. Store transaction record. Send receipt email.</p><p>Each acceptance test becomes a separate work item. Most will be small, genuinely completable in a day. The few that aren&#8217;t get broken down further.</p><h3>Technique 2: The &#8220;One Dumb Idea&#8221; approach</h3><p>Jeffries describes a psychological technique for unlocking small stories. When teams face seemingly monolithic features, propose a deliberately inadequate but technically possible first step.</p><p>His cable TV example: Instead of building full pay-per-view functionality, propose &#8220;Play one specific movie on a secret channel.&#8221; This exists almost entirely with current infrastructure. No user selection, no payment, no scheduling. Just hard-code a movie playing on channel 999.</p><p>The power lies in team psychology. When someone proposes an obviously insufficient solution, others instinctively respond with &#8220;What we could do instead is...&#8221; Suddenly the conversation shifts from &#8220;this is impossible&#8221; to discussing achievable increments. As Jeffries notes: &#8220;We&#8217;ve gone, in one step, from &#8216;impossible&#8217; to knowing a stupid, but possible, thing to do.&#8221;</p><h3>Technique 3: Vertical slicing with minimal viability</h3><p>Each small story should deliver something end-to-end, however minimal.</p><p>Horizontal slicing (avoid): Build database schema. Create API endpoints. Implement frontend components. Wire everything together.</p><p>Vertical slicing (prefer): User can see one hardcoded product (end-to-end). User can see one product from database. User can see list of products. User can filter products by category.</p><p>The vertical approach produces shippable increments and reveals integration problems immediately rather than concentrating them at the end.</p><h3>Why small stories help estimation</h3><p>Even if you don&#8217;t eliminate estimation entirely, small stories transform the accuracy problem.</p><p>Reduced variability: A 10-day estimate might be off by 5 days. Ten 1-day estimates won&#8217;t all be wrong in the same direction.</p><p>Faster feedback: When items complete daily, you discover problems within days rather than weeks.</p><p>Throughput becomes measurable: With sufficient small items, historical throughput data emerges quickly, enabling forecasting without per-item estimation.</p><p>Cognitive load decreases: Estimating &#8220;this takes about a day&#8221; requires less analysis than forecasting two-week epics.</p><p>This approach requires investment in slicing skills and may initially feel slower than rougher-grained planning. But Jeffries argues the payoff is faster delivery with less estimation overhead, and crucially, less opportunity for estimates to become weaponised commitments.</p><div><hr></div><h2>The NoEstimates philosophy</h2><p><em>In brief: The actual argument isn&#8217;t &#8220;don&#8217;t estimate&#8221; but &#8220;continuously question whether estimation is earning its keep.&#8221;</em></p><p>The NoEstimates movement, associated with Woody Zuill, Vasco Duarte, and Ron Jeffries, often gets reduced to &#8220;don&#8217;t estimate.&#8221; The actual argument is more nuanced: continuously question whether estimation is earning its keep.</p><h3>Estimation as expense</h3><p>Jeffries frames it starkly: &#8220;Estimates are always waste; they are not our product.&#8221; From a lean perspective, any activity that doesn&#8217;t directly produce customer value is expense. Estimation doesn&#8217;t ship features. The question becomes: does this expense generate sufficient return in decision quality to justify its cost?</p><p>On the C3 payroll project (Extreme Programming&#8217;s flagship case study), Jeffries&#8217; team initially used story estimates for planning. They later realised they could have achieved similar outcomes by breaking work into single acceptance tests and counting completions. Mechanical measurement replacing estimation entirely.</p><h3>When estimation provides value</h3><p>Jeffries acknowledges legitimate use cases.</p><p>Sales and contracts: Pricing decisions require some basis. Though he critiques how organisations weaponise estimates in negotiations, he concedes that customers reasonably want cost projections before committing.</p><p>Understanding: Estimation discussions surface differing interpretations of requirements. Team members discover they imagined different solutions. However, Jeffries suggests the <em>conversation</em> provides this value. Written estimates aren&#8217;t strictly necessary.</p><p>Learning: Comparing estimates to actuals reveals systematic biases and process problems. Yet alternative monitoring methods exist; you can track cycle time without estimating individual items.</p><h3>The pragmatic position</h3><p>Jeffries&#8217; conclusion is measured: &#8220;We always <em>could</em> stop estimating, but it&#8217;s not always the right thing to do. It&#8217;s always legitimate to think about it.&#8221;</p><p>This isn&#8217;t dogma. It&#8217;s a heuristic for continuous improvement. Each time estimation seems mandatory, ask: Is there a way to make decisions without this expense? What would we lose? What might we gain? Sometimes the answer favours estimation. Often, teams discover they estimate from habit rather than necessity.</p><div><hr></div><h2>When estimation is unavoidable</h2><p><em>In brief: Fixed-price contracts, regulatory deadlines, capital budgeting, and external coordination all require estimates. The answer is to estimate honestly, not to pretend estimation is unnecessary.</em></p><p>Some contexts don&#8217;t allow the luxury of &#8220;we&#8217;ll deliver what we can when we can&#8221;. In these situations, estimation isn&#8217;t optional, and the question becomes how to do it less badly.</p><p>Fixed-price contracts require estimates. A client asking for a quote needs a number, and &#8220;it depends&#8221; doesn&#8217;t win business. You can build contingency into the price, you can structure contracts with change mechanisms, but you can&#8217;t avoid making a forward-looking commitment.</p><p>Regulatory commitments have hard deadlines. If compliance with a new regulation is required by a specific date, missing it has consequences that probabilistic language doesn&#8217;t soften. You need to know whether you&#8217;re likely to make it, and if not, what to do about it.</p><p>Capital budgeting requires forecasts. Organisations allocate resources annually or quarterly. Someone deciding whether to fund your initiative versus a competing one needs to understand what they&#8217;re getting for their investment. &#8220;Trust us&#8221; isn&#8217;t a capital allocation strategy.</p><p>External stakeholder negotiations depend on estimates. If you&#8217;re coordinating with partners, aligning marketing campaigns, or scheduling dependent work streams, those stakeholders need something to plan against.</p><p>In these situations, the answer isn&#8217;t to pretend estimation is unnecessary. It&#8217;s to estimate honestly: provide ranges rather than points, communicate confidence levels, update forecasts as you learn, and build relationships where changing estimates isn&#8217;t treated as failure.</p><div><hr></div><h2>Estimating less badly</h2><p><em>In brief: Ranges, confidence levels, rolling-wave planning, Bayesian updating, and treating estimates as decision-support rather than promises.</em></p><p>When estimation is required, several practices help reduce the damage.</p><p>Use ranges instead of points. &#8220;Two to four weeks&#8221; is more honest than &#8220;three weeks&#8221; and gives stakeholders useful information about uncertainty. If they need the optimistic end, they know it&#8217;s a stretch. If they need certainty, they can plan for the pessimistic end.</p><p>Express confidence levels explicitly. P50, P75, and P90 estimates communicate that different levels of certainty come with different timelines. A P50 estimate means there&#8217;s a 50% chance of missing it. If that risk is unacceptable, plan for P90.</p><p>Practice rolling-wave planning. Estimate near-term work in detail, further-out work in ranges, and distant work as rough orders of magnitude. Don&#8217;t pretend you know what you&#8217;ll discover.</p><p>Update estimates as you learn. Bayesian updating means revising your forecasts as new information emerges. An estimate made at the start of a project should evolve. Treating the original estimate as a commitment regardless of what you&#8217;ve learned is organisational dysfunction, not estimation failure.</p><p>Treat estimates as decision-support, not promises. The purpose of an estimate is to help someone make a decision: should we fund this, should we commit to this date, should we staff this team. Once the decision is made, the estimate has served its purpose. Holding people to it regardless of changed circumstances misunderstands what estimates are for.</p><p>McConnell uses a helpful metaphor: estimates need not be perfect, just close enough that minor adjustments (equivalent to &#8220;sitting on the suitcase&#8221;) achieve reasonable success. Obsessing over estimation precision often misses the point.</p><div><hr></div><h2>Hybrid approaches</h2><p><em>In brief: Use rough estimates early and forecasts later. Combine discovery and delivery. Layer probabilistic methods on top of expert judgment.</em></p><p>The debate is often framed as estimation versus forecasting, as if you must choose one camp. In practice, hybrid approaches often work best.</p><p>Use rough estimates early, forecasts later. At the inception of a project, you lack historical data for the specific work. Rough expert estimates, honestly communicated as guesses, help with initial go/no-go decisions. As work progresses and you accumulate data, shift to probabilistic forecasting based on actual throughput.</p><p>Combine discovery and delivery tracks. Run a time-boxed discovery phase to reduce uncertainty before committing to estimates. The goal of discovery is to learn enough that your subsequent estimates have a meaningful basis. Don&#8217;t estimate what you haven&#8217;t explored.</p><p>Use scenario-based planning. Instead of a single estimate, develop scenarios: &#8220;If the integration goes smoothly, four weeks. If we hit the authentication complexity we suspect, eight weeks. If we need to rebuild the data layer, three months.&#8221; This surfaces the key risks and lets stakeholders understand what drives the uncertainty.</p><p>Layer probabilistic forecasts on top of rough scoping. Use expert judgment to identify the likely scope, then apply Monte Carlo to the execution. The forecast doesn&#8217;t replace judgment about what needs to be built; it provides rigour around how long building takes.</p><p>McConnell advocates something similar: define requirements upfront with enough detail for story point estimation, then track velocity to calibrate forecasts. This offers a middle path between pure NoEstimates and traditional detailed estimation.</p><div><hr></div><h2>The economics of estimation</h2><p><em>In brief: Cost of delay, risk-adjusted ROI, opportunity cost, and option value. If we want &#8220;adult economic conversations&#8221;, we should actually talk economics.</em></p><p>One thing largely missing from the estimation debate is economic decision theory. If we want &#8220;adult economic conversations&#8221;, we should actually talk economics.</p><p>Cost of delay matters enormously. A feature delivered in January might be worth twice what it&#8217;s worth in June. If you&#8217;re choosing between a certain six-month delivery and a risky four-month delivery, the right choice depends on how value decays over time. Probabilistic forecasting is most useful when connected to explicit cost-of-delay analysis.</p><p>Risk-adjusted return on investment changes decisions. A project with an expected value of &#163;1 million but high variance might be less attractive than one with an expected value of &#163;800,000 and low variance. Portfolio thinking requires understanding not just expected outcomes but distributions of outcomes.</p><p>Opportunity cost is invisible but real. While your team spends six months on Project A, they&#8217;re not working on Projects B, C, and D. The value of better estimation isn&#8217;t just delivering A faster; it&#8217;s making better choices about whether to do A at all.</p><p>Option value exists in uncertainty. Sometimes the right response to uncertainty isn&#8217;t better estimation; it&#8217;s structuring work to preserve options. Small investments that let you learn before committing are often worth more than precise forecasts that lock you in.</p><div><hr></div><h2>Fixed time, variable scope</h2><p><em>In brief: Shape Up&#8217;s approach works well for product-led organisations with autonomy. It doesn&#8217;t work for regulatory deadlines or fixed external commitments.</em></p><p>Basecamp&#8217;s Shape Up methodology offers an interesting flip. Instead of fixing scope and letting time vary, you fix time and let scope vary. You decide that something is worth six weeks of effort, then build the best version you can in six weeks.</p><p>This is liberating in some contexts. Instead of expanding timelines to fit scope (which leads to Parkinson&#8217;s Law), you ruthlessly trim scope to fit timelines. Six weeks turns out to be a sweet spot: long enough to finish something meaningful, short enough to see the end from the beginning.</p><p>But Shape Up is context-specific, not universally applicable. It works well for product-led organisations with strong product management, high team autonomy, and low external deadline pressure. It works poorly for contract-based delivery, regulatory milestones, hardware dependencies, and systems with heavy multi-team coordination.</p><p>If the business requirement is &#8220;we need these specific regulatory features by this compliance deadline&#8221;, you can&#8217;t negotiate scope. If you&#8217;re coordinating with external partners who expect specific functionality, you can&#8217;t just deliver &#8220;whatever fits&#8221;. Shape Up is a valuable tool where it applies, not a universal solution.</p><div><hr></div><h2>The selection bias in success stories</h2><p><em>In brief: Teams that work without traditional estimates have usually earned that trust through years of reliable delivery. Many teams don&#8217;t have that luxury.</em></p><p>Teams that successfully operate without traditional estimates share characteristics that often go unmentioned. They typically have high trust with stakeholders built over years of reliable delivery. They have mature development practices: continuous integration, automated testing, small incremental releases. They have stable funding that doesn&#8217;t require competitive justification.</p><p>In other words, they&#8217;ve earned the right to say &#8220;trust us&#8221;. They can operate with probabilistic forecasts because their stakeholders have seen enough delivery to believe the probabilities are meaningful.</p><p>Many teams operate in environments where that trust hasn&#8217;t been established. Funding is competitive. External dependencies require coordination. Stakeholders have been burned before and want commitments. Telling those teams to &#8220;just stop estimating&#8221; isn&#8217;t practical advice. They need to build trust first, which often means delivering reliably against stated expectations, which requires some form of estimation.</p><div><hr></div><h2>Flow metrics: the foundation</h2><p><em>In brief: Cycle time, throughput, WIP, and work item age. These measure what actually happened rather than what someone guessed. But they can be gamed too.</em></p><p>Whatever approach you take, tracking the right metrics matters. The four essential flow metrics are cycle time (how long work takes from start to finish), throughput (how many items you complete per time period), work in progress (how many items are currently in flight), and work item age (how long an item has been in progress).</p><p>These connect through Little&#8217;s Law: average cycle time equals average work in progress divided by average throughput. Limiting WIP decreases cycle time. Decreased cycle time increases predictability. Increased predictability makes forecasting more accurate.</p><p>Teams that shift from story points and velocity to flow metrics often report reduced cycle times and better predictability. The mechanism is straightforward: flow metrics measure what actually happened, while story points measure what someone guessed would happen.</p><p>But remember: these metrics can be gamed too. The value comes from honest measurement and continuous improvement, not from the metrics themselves. Any metric that becomes a target ceases to be a good metric. The solution is cultural, not methodological.</p><div><hr></div><h2>The real enemy: organisational dysfunction</h2><p><em>In brief: The problem isn&#8217;t estimation. It&#8217;s what organisations do with estimates. False precision is the enemy, not prediction itself.</em></p><p>Here&#8217;s the deeper issue the estimation debate often misses: the problem isn&#8217;t estimation itself. It&#8217;s the organisational dysfunction that surrounds it.</p><p>A team saying &#8220;two to four weeks, depending on what we discover&#8221; is estimating, and that&#8217;s fine. The problem is when that becomes &#8220;we committed to two weeks&#8221; in a status report, which becomes &#8220;why did you miss your commitment?&#8221; in a performance review. The dysfunction isn&#8217;t the estimate; it&#8217;s what the organisation does with it.</p><p>False precision is the enemy, not estimation. When someone asks &#8220;how long will this take?&#8221; and the honest answer is &#8220;probably 2-4 weeks, but it could be longer if we hit complications&#8221;, the organisation needs to be able to hear that. If it can&#8217;t, if it demands a single number and then treats that number as a commitment, the problem is cultural, not methodological.</p><h3>The responsibility question</h3><p>Ron Jeffries raises a pointed question about accountability. Developers cannot reasonably be held responsible for meeting deadlines without corresponding authority. Unless developers can hire or reassign staff, acquire additional resources, decide which features ship versus defer, or adjust scope unilaterally, they cannot control the variables that determine delivery dates.</p><p>Jeffries compares development to a machine with fixed capacity. Pushing harder doesn&#8217;t increase throughput; it risks breakdown. The product owner must select work batches that fit the timeline, not demand the machine work faster.</p><p>This has uncomfortable implications. When projects fail to meet deadlines, the conventional response is blaming developers for poor estimates. Jeffries suggests the failure often lies in management&#8217;s scope decisions, resource allocation, or unrealistic targets, factors developers cannot control.</p><p>What developers <em>can</em> commit to: delivering working, tested features regularly; keeping the codebase shippable at all times; working on whatever sequence management prioritises; surfacing impediments early rather than hiding them. This is narrower accountability than &#8220;hit the date,&#8221; but it&#8217;s accountability developers can actually fulfil without authority over scope and resources.</p><p>Dysfunction arises when organisations hold people accountable for outcomes they cannot control. Estimates become the instrument for manufacturing this false accountability.</p><h3>The cultural barrier</h3><p>Monte Carlo simulations and probabilistic forecasting help with the honesty part. They make it harder to pretend certainty where none exists. But they don&#8217;t solve the cultural part. An organisation that punishes missed estimates won&#8217;t suddenly become healthy because you started expressing estimates as probability distributions. They&#8217;ll just punish you for missing the 75th percentile instead of the point estimate.</p><div><hr></div><h2>What actually helps</h2><p><em>In brief: Track reality, embrace probability, keep things small, focus on flow, estimate honestly when you must, and build a culture that can handle uncertainty.</em></p><p>Based on all the evidence, here&#8217;s what actually improves outcomes.</p><p>Track what actually happens. Historical throughput beats expert guesses. Collect cycle time and throughput data consistently, even if you&#8217;re not sure how you&#8217;ll use it yet.</p><p>Embrace probabilistic thinking. Present ranges with confidence levels rather than single dates. Have risk conversations rather than commitment ceremonies. Acknowledge that you&#8217;re uncertain and explain what drives the uncertainty.</p><p>Keep work items genuinely small. Target 1-day items where possible, certainly no more than 3 days. Use the slicing techniques: single acceptance tests, vertical slicing, the &#8220;one dumb idea&#8221; approach. Smaller items mean tighter distributions and better forecasts.</p><p>Focus on flow. Measure cycle time and throughput. Limit work in progress. The maths is clear: lower WIP means faster cycle times and more predictable delivery.</p><p>Where possible, fix time and vary scope. If you can negotiate what gets built, time-boxing forces real prioritisation and reduces the scope creep that derails traditional projects.</p><p>When estimation is required, do it honestly. Use ranges, express confidence levels, update as you learn, and treat estimates as decision-support rather than commitments. Separate estimates from targets from commitments in your vocabulary and your conversations.</p><p>Build organisational tolerance for uncertainty. This is the hardest part and the most important. If your organisation punishes honest uncertainty, no methodology will save you. Work on the culture alongside the practices.</p><div><hr></div><h2>What about AI?</h2><p><em>In brief: AI agents change the nature of the work being estimated. Historical data becomes less relevant. Variance increases. The estimation problem gets harder before it gets easier.</em></p><p>AI coding agents are already changing how software gets built. Tasks that took a day now take an hour. But not all tasks, and not predictably. This creates a new estimation problem: how do you forecast when the work itself is transforming?</p><p>Historical throughput data assumes some stability in how work gets done. If your team&#8217;s cycle time for &#8220;build a new API endpoint&#8221; was consistently 2-3 days, you could forecast based on that. But now one developer with an AI agent finishes it in two hours, while another developer working on a different endpoint hits edge cases the AI can&#8217;t handle and takes three days anyway. Your historical distribution no longer describes your current capability.</p><p>The variance problem gets worse, not better. AI agents are fast when they work and useless when they don&#8217;t, and predicting which situation you&#8217;ll hit is difficult. A task might complete in minutes if the AI handles it cleanly, or take longer than the pre-AI baseline if you spend hours debugging AI-generated code that almost works. The distribution of outcomes becomes bimodal or worse, which breaks the assumptions behind Monte Carlo forecasting.</p><p>There&#8217;s also a decomposition problem. Traditional estimation assumes humans do the work and you&#8217;re estimating human effort. When AI agents do significant portions of the work, what exactly are you estimating? The human time spent prompting, reviewing, and correcting? The wall-clock time including AI processing? The cognitive load on the human, which might be higher when supervising AI than when doing the work directly? None of our existing frameworks handle this cleanly.</p><p>The honest answer is that we don&#8217;t yet know how to estimate human-AI collaborative work well. Teams adopting AI agents should expect their forecasting accuracy to degrade temporarily. Historical data becomes less useful. New patterns haven&#8217;t stabilised enough to replace the old ones. The best approach is probably radical incrementalism: even smaller batches, even shorter feedback loops, even more willingness to update forecasts as you learn. Treat every AI-assisted task as an experiment until you&#8217;ve built enough new history to see patterns.</p><p>This might ultimately be good news for the estimation debate. If AI makes historical data unreliable and variance unpredictable, organisations will be forced to accept uncertainty whether they like it or not. You can&#8217;t demand false precision when everyone can see the ground shifting. But in the short term, expect estimation to get harder, not easier.</p><div><hr></div><h2>Conclusion: three voices, one uncomfortable truth</h2><p>We&#8217;ve heard three distinct perspectives on software estimation.</p><p><strong>Steve McConnell</strong> argues that estimation is a craft that can be practiced skilfully. Distinguish estimates from targets from commitments. Understand the Cone of Uncertainty. Use historical data and reference classes. Track actuals. The problem isn&#8217;t estimation itself but doing it carelessly.</p><p><strong>Ron Jeffries</strong> questions whether estimation deserves its central role. Every estimate is expense, not product. When work is sliced small enough and delivered continuously, forecasting emerges from counting rather than guessing. At minimum, keep asking: can we make this decision without estimating?</p><p><strong>The organisational dysfunction thesis</strong> identifies cultural problems as the root cause: not estimation technique or philosophy, but what organisations do with estimates, demanding false precision, punishing honest uncertainty, treating preliminary forecasts as ironclad commitments.</p><p>These perspectives aren&#8217;t as contradictory as they first appear.</p><p>McConnell would agree that estimation without distinguishing it from commitment is organisational malpractice. Jeffries would agree that when estimation is necessary, doing it well beats doing it badly. All parties agree that probabilistic language beats point estimates, that smaller work items improve predictability, and that organisational culture determines whether any technique succeeds.</p><p>McConnell offers: &#8220;The primary purpose of software estimation is not to predict a project&#8217;s outcome; it is to determine whether a project&#8217;s targets are realistic enough to allow the project to be controlled to meet them.&#8221;</p><p>Jeffries offers: &#8220;Stop estimating. Start shipping.&#8221;</p><p>Perhaps both are right. When an organisation has the maturity to use estimates as McConnell envisions, for project control rather than prophecy, estimation becomes a valuable tool. When an organisation lacks that maturity, Jeffries&#8217; provocative advice may be the safest path: stop the estimation theatre entirely, focus on small deliverables, and let observable throughput speak for itself.</p><p>The question isn&#8217;t &#8220;Should we estimate?&#8221; but &#8220;Has our organisation earned the right to estimate responsibly?&#8221;</p><p>That requires cultural change harder than switching methodologies, but it&#8217;s the only approach that genuinely works.<br></p><div><hr></div><h2>References</h2><p><strong>Steve McConnell</strong></p><p>McConnell, S. (2006). <em>Software Estimation: Demystifying the Black Art</em>. Microsoft Press.</p><p>McConnell, S. &#8220;Software Engineering Radio Episode 273: Steve McConnell on Software Estimation.&#8221; Software Engineering Radio.</p><p>McConnell, S. &#8220;The Cone of Uncertainty.&#8221; Construx. https://www.construx.com/books/the-cone-of-uncertainty/</p><p><strong>Ron Jeffries</strong></p><p>Jeffries, R. &#8220;The NoEstimates Movement.&#8221; ronjeffries.com. https://ronjeffries.com/xprog/articles/the-noestimates-movement/</p><p>Jeffries, R. &#8220;Estimation is Evil.&#8221; ronjeffries.com. https://ronjeffries.com/articles/019-01ff/estimation-is-evil/</p><p>Jeffries, R. &#8220;Getting Small Stories.&#8221; ronjeffries.com. https://ronjeffries.com/articles/015-10/small-stories/</p><p>Jeffries, R. &#8220;Making the Date.&#8221; ronjeffries.com. https://ronjeffries.com/articles/making-the-date/</p><p>Jeffries, R. &#8220;Story Points Revisited.&#8221; ronjeffries.com. https://ronjeffries.com/articles/019-01ff/story-points/Index.html</p><p>Jeffries, R. (2012). Comment on Scrum Alliance discussion regarding velocity. Referenced in InfoQ: &#8220;Should we stop using Story Points and Velocity?&#8221;</p><p><strong>Daniel Kahneman and Amos Tversky</strong></p><p>Kahneman, D. &amp; Tversky, A. (1979). &#8220;Intuitive Prediction: Biases and Corrective Procedures.&#8221; TIMS Studies in Management Science, 12, 313-327.</p><p>Kahneman, D. (2011). <em>Thinking, Fast and Slow</em>. Farrar, Straus and Giroux.</p><p><strong>Reference Class Forecasting</strong></p><p>Flyvbjerg, B. (2006). &#8220;From Nobel Prize to Project Management: Getting Risks Right.&#8221; Project Management Journal, 37(3), 5-15.</p><p>Wikipedia. &#8220;Reference Class Forecasting.&#8221; https://en.wikipedia.org/wiki/Reference_class_forecasting</p><p><strong>Project Statistics</strong></p><p>Flyvbjerg, B. (2021). &#8220;How Big Things Get Done.&#8221; Oxford Sa&#239;d Business School.</p><p>The Standish Group. &#8220;CHAOS Report.&#8221; https://www.standishgroup.com/</p><p><strong>NoEstimates Movement</strong></p><p>Duarte, V. <em>NoEstimates: How to Measure Project Progress Without Estimating</em>. Leanpub. https://leanpub.com/noestimates</p><p>Zuill, W. &#8220;NoEstimates.&#8221; https://zuill.us/WosenseeBlog/tag/noestimates/</p><p>NoEstimates.org. Links and Resources. https://www.noestimates.org/</p><p><strong>Flow Metrics and Forecasting</strong></p><p>Vacanti, D. <em>When Will It Be Done? Lean-Agile Forecasting to Answer Your Customers&#8217; Most Important Question</em>. Leanpub. https://leanpub.com/whenwillitbedone</p><p>Vacanti, D. <em>Actionable Agile Metrics for Predictability</em>. Leanpub. https://leanpub.com/actionableagilemetrics</p><p>Kersten, M. (2018). <em>Project to Product: How to Survive and Thrive in the Age of Digital Disruption with the Flow Framework</em>. IT Revolution Press.</p><p><strong>DORA Research</strong></p><p>Forsgren, N., Humble, J., &amp; Kim, G. (2018). <em>Accelerate: The Science of Lean Software and DevOps</em>. IT Revolution Press.</p><p>DORA. &#8220;DORA Metrics: The Four Keys.&#8221; https://dora.dev/guides/dora-metrics-four-keys/</p><p><strong>Shape Up</strong></p><p>Singer, R. <em>Shape Up: Stop Running in Circles and Ship Work that Matters</em>. Basecamp. https://basecamp.com/shapeup</p><p><strong>Additional Sources</strong></p><p>Killick, N. &#8220;Slicing Heuristics.&#8221; Referenced in Jeffries&#8217; work on small stories.</p><p>Beck, K. et al. &#8220;Manifesto for Agile Software Development.&#8221; https://agilemanifesto.org/</p>]]></content:encoded></item><item><title><![CDATA[The Story Point Illusion]]></title><description><![CDATA[Why your team's estimates are hiding disagreement instead of resolving it]]></description><link>https://a4al6a.substack.com/p/the-story-point-illusion</link><guid isPermaLink="false">https://a4al6a.substack.com/p/the-story-point-illusion</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Thu, 15 Jan 2026 18:56:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ANa9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ANa9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ANa9!, /__u/a4al6a.substack.com/w_424, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png 424w, /__u/substackcdn.com/image/fetch/$s_!ANa9!, /__u/a4al6a.substack.com/w_848, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png 848w, /__u/substackcdn.com/image/fetch/$s_!ANa9!, /__u/a4al6a.substack.com/w_1272, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ANa9!, /__u/a4al6a.substack.com/w_1456, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_webp, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ANa9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png" width="800" height="603" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:603,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Why Story Pointing Needs to Die. It comes down to an understanding of&#8230; | by  Quinton (Ron) Quartel | The Startup | Medium&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Why Story Pointing Needs to Die. It comes down to an understanding of&#8230; | by  Quinton (Ron) Quartel | The Startup | Medium" title="Why Story Pointing Needs to Die. It comes down to an understanding of&#8230; | by  Quinton (Ron) Quartel | The Startup | Medium" srcset="/__u/substackcdn.com/image/fetch/$s_!ANa9!, /__u/a4al6a.substack.com/w_424, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png 424w, /__u/substackcdn.com/image/fetch/$s_!ANa9!, /__u/a4al6a.substack.com/w_848, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png 848w, /__u/substackcdn.com/image/fetch/$s_!ANa9!, /__u/a4al6a.substack.com/w_1272, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ANa9!, /__u/a4al6a.substack.com/w_1456, /__u/a4al6a.substack.com/c_limit, /__u/a4al6a.substack.com/f_auto, /__u/a4al6a.substack.com/q_auto:good, /__u/a4al6a.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75c4caa1-388e-4cc4-8e94-70f2fcc2dcc2_800x603.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>&#8220;Story points aren&#8217;t about the numbers. They trigger conversations about effort and complexity.&#8221;</p><p>I hear this defence constantly. It sounds reasonable on the surface. Almost wise. But I want you to think carefully about what it&#8217;s actually saying, because underneath the reasonable-sounding words lies a significant admission.</p><p>If the real value of story points is the conversation they trigger, why do you need story points to have that conversation?</p><p>You could just talk about the work. Directly. Without the numerical theatre. Without the planning poker cards. Without the endless debates about whether this piece of work is a 3 or a 5 or maybe an 8.</p><p>The fact that we need to dress up &#8220;having a conversation about work&#8221; in estimation clothing should make us suspicious. It suggests that something else is going on. And that something else is where the danger lies.</p><p>Let me be clear about what I mean when I say story points are dangerous. I don&#8217;t mean they&#8217;re mildly unhelpful or slightly inefficient. I mean they actively harm teams, damage trust, and create dysfunction. They do this in ways that are subtle enough to go unnoticed for years, which makes them more dangerous, not less.</p><h4><strong>Story points don&#8217;t measure anything real</strong></h4><p>When you measure something in metres, everyone agrees what a metre is. When you measure something in kilograms, we have a shared understanding of the unit. Story points don&#8217;t work this way. A &#8220;point&#8221; means something different to every team, and often means something different to every person on the same team.</p><p>Some people think about points as effort. How much work will this take? Some people think about points as complexity. How many moving parts are involved? Some people think about points as risk. How likely is this to go wrong? Some people think about points as uncertainty. How much do we not know?</p><p>These are all different things. Effort and complexity are related but not the same. A task can be simple but effortful, like copying a thousand records by hand. A task can be complex but quick, like making a small change to a critical algorithm where you need to understand the whole system but the actual change is tiny.</p><p>When your team sits down to estimate, each person is measuring a different thing with the same number. You&#8217;re using the same scale to weigh apples, measure distances, and count sheep. Then you&#8217;re surprised when the numbers don&#8217;t mean anything useful.</p><h4><strong>Agreement on a number creates an illusion of shared understanding</strong></h4><p>This problem is more insidious than the first. Imagine your team is estimating a story. After some discussion, everyone holds up a 5. Success, right? You&#8217;ve reached consensus. The process worked.</p><p>But did it?</p><p>Person A is thinking about the database changes required. They&#8217;ve estimated a 5 because they know the schema is messy and migrations are risky.</p><p>Person B is thinking about the front-end work. They&#8217;ve estimated a 5 because there are several screens to update and the design isn&#8217;t finalised.</p><p>Person C is thinking about the integration with the external API. They&#8217;ve estimated a 5 because they&#8217;ve never worked with this API before and the documentation looks sparse.</p><p>Person D is thinking about testing. They&#8217;ve estimated a 5 because this feature touches several existing workflows and regression testing will be extensive.</p><p>Everyone said 5. Everyone had completely different reasons. They&#8217;re not estimating the same work. They&#8217;re not even thinking about the same aspects of the work.</p><p>The number created false consensus. It let everyone believe they were aligned when they weren&#8217;t. And this false consensus is worse than open disagreement, because open disagreement gets resolved. False consensus just waits quietly until it explodes during implementation.</p><p>If Person A starts working on this story, they might ignore the front-end complexity that Person B was worried about. If Person C picks it up, they might not realise the database concerns that Person A had in mind. The &#8220;shared estimate&#8221; shared nothing except a digit.</p><p>This is what I mean when I say story points are dangerous. They create a feeling of alignment without actual alignment. They let teams believe they&#8217;ve communicated when they haven&#8217;t. They substitute a ritual for real understanding.</p><h4><strong>The retreat to &#8220;valuable conversations&#8221;</strong></h4><p>Now let&#8217;s talk about what happens when teams try to defend story points.</p><p>The most common defence is the one I started with: story points trigger valuable conversations. But notice what this defence concedes. It concedes that the points themselves aren&#8217;t valuable. It concedes that accuracy doesn&#8217;t matter. It concedes that the numbers are essentially arbitrary.</p><p>If someone told you that their weighing scale was completely inaccurate but it was still valuable because stepping on it reminded them to think about their health, you&#8217;d suggest they just think about their health directly and throw away the broken scale.</p><p>The &#8220;triggers conversations&#8221; defence is what people say when they can no longer defend the thing on its own merits. It&#8217;s a retreat to secondary benefits. And those secondary benefits can be achieved more directly without the harmful primary effects.</p><p>If you want to understand complexity, ask &#8220;what makes this complicated?&#8221; directly.</p><p>If you want to surface risks, ask &#8220;what could go wrong?&#8221; directly.</p><p>If you want to check whether the team has shared understanding, ask &#8220;how would we approach this?&#8221; directly.</p><p>If you want to know if everyone is thinking about the same work, ask &#8220;what do you think this story involves?&#8221; directly.</p><p>These questions actually achieve the supposed goals of estimation. They surface different perspectives. They reveal hidden assumptions. They create genuine shared understanding rather than the illusion of it.</p><p>And crucially, they don&#8217;t produce a number that will later be misused.</p><h4><strong>Story points get weaponised</strong></h4><p>Because here&#8217;s the third danger of story points: they get weaponised.</p><p>No matter how many times you tell management that story points aren&#8217;t comparable across teams, they will compare them across teams. No matter how many times you explain that velocity isn&#8217;t a productivity measure, it will be used as a productivity measure. No matter how many times you insist that points shouldn&#8217;t be converted to hours, someone will create a conversion formula.</p><p>Story points create numbers. Organisations love numbers. Numbers go into spreadsheets. Spreadsheets go into reports. Reports go to executives. Executives make decisions based on those numbers.</p><p>And now your meaningless, arbitrary, inconsistent, incomparable numbers are driving organisational decisions. Teams get pressured to increase velocity. Developers learn to inflate estimates so they can &#8220;deliver more points.&#8221; Gaming begins. Trust erodes.</p><p>You might say this is a misuse of story points, not a problem with story points themselves. But this misuse is inevitable. It happens in nearly every organisation that adopts story points. If a tool is consistently misused across thousands of different organisations with different cultures and different people, the problem is the tool, not the users.</p><p>A good tool makes the right thing easy and the wrong thing hard. Story points make the wrong thing easy. They produce numbers that look meaningful, invite comparison, and beg to be aggregated. They&#8217;re practically designed to be misused.</p><h4><strong>The myth of improving at estimation</strong></h4><p>Let&#8217;s look at another defence: story points help teams improve at estimation over time.</p><p>This sounds plausible. Practice makes perfect, right? But there&#8217;s a hidden assumption here that deserves examination. The assumption is that there&#8217;s a stable, learnable skill called &#8220;estimation&#8221; that teams can get better at.</p><p>In reality, every piece of work is different. The factors that made your last story take longer than expected are probably not the factors that will affect your next story. Software development is not like manufacturing widgets, where you can measure cycle time and optimise a repeatable process. Each story involves different code, different requirements, different unknowns.</p><p>Teams that track their estimation accuracy over time often find that it doesn&#8217;t improve. Or it improves for a while and then degrades. Or it varies randomly with no clear trend. This isn&#8217;t because the teams are bad at learning. It&#8217;s because there isn&#8217;t a stable underlying skill to learn.</p><p>What teams actually get better at is playing the estimation game. They learn what numbers the organisation wants to hear. They learn to pad estimates to create safety margin. They learn to break work into smaller pieces not because smaller pieces are better but because smaller estimates are less scrutinised.</p><p>The improvement is an illusion. Or worse, the improvement is real but it&#8217;s improvement at gaming the system rather than improvement at understanding work.</p><h4><strong>The false promise of avoiding false precision</strong></h4><p>Here&#8217;s another defence I hear: story points are better than time-based estimates because they avoid the trap of false precision.</p><p>This defence has some truth in it. Time-based estimates do create false precision. Saying &#8220;this will take 3.5 days&#8221; implies a level of accuracy that doesn&#8217;t exist.</p><p>But story points don&#8217;t solve this problem. They just move it. Instead of false precision about time, you get false precision about relative size. Saying &#8220;this is a 5&#8221; implies that you know how this work compares to other work, that you understand the scope well enough to place it on a scale, that the number means something.</p><p>And then the false precision comes back anyway, because most organisations convert velocity into time-based forecasts. If you complete 30 points per sprint and your backlog has 300 points, you forecast 10 sprints. The time-based precision you tried to avoid has snuck back in through the back door.</p><p>You haven&#8217;t eliminated false precision. You&#8217;ve just added a layer of indirection that makes the false precision harder to see and challenge.</p><h4><strong>What story points do to how teams think</strong></h4><p>Now I want to address the deepest problem with story points, which is what they do to the way teams think about work.</p><p>Story points encourage teams to think about work as a quantity to be estimated rather than a problem to be understood.</p><p>When you approach work through the lens of estimation, you&#8217;re asking &#8220;how big is this?&#8221; That&#8217;s a question about the work as a fixed object, something with a predetermined size that you&#8217;re trying to measure.</p><p>But software development doesn&#8217;t work that way. The work isn&#8217;t fixed. The scope isn&#8217;t predetermined. The size emerges from how you approach the problem, what trade-offs you make, what you discover along the way.</p><p>When teams approach work through the lens of understanding, they ask different questions. What are we trying to achieve? What&#8217;s the simplest thing that could work? What do we need to learn? What could we defer?</p><p>These questions lead to better outcomes because they engage with the work as something to be shaped rather than something to be measured. They open up possibilities rather than closing them down into a single number.</p><p>I&#8217;ve watched teams spend an hour debating whether a story is a 5 or an 8. That&#8217;s an hour they could have spent actually understanding the work. Actually talking about the approach. Actually identifying risks and unknowns. Actually building shared mental models.</p><p>The estimation ritual crowds out the valuable activities. It substitutes a proxy for the real thing. And because the proxy feels productive, feels like work, feels like alignment, teams don&#8217;t notice what they&#8217;re missing.</p><h4><strong>What happens when teams stop estimating</strong></h4><p>Let me talk about what happens when teams stop using story points.</p><p>The first thing that happens is fear. Managers worry about losing visibility. They ask how they&#8217;ll know when things will be done. They ask how they&#8217;ll track progress. They ask how they&#8217;ll compare teams.</p><p>These fears are worth examining. What visibility did story points actually provide? If the numbers were arbitrary and inconsistent and gameable, what were managers really seeing? They were seeing a theatrical performance of estimation, not a window into reality.</p><p>When teams stop estimating and start focusing on flow, something interesting happens. They start measuring things that actually matter. How long do items spend in progress? Where do items get stuck? How often do items get blocked? What&#8217;s the cycle time from start to finish?</p><p>These measurements are grounded in reality. They&#8217;re not opinions or guesses. They&#8217;re observations of what actually happened. And they&#8217;re much harder to game because they&#8217;re based on timestamps, not feelings.</p><p>Teams that focus on flow also tend to break work into smaller pieces, not because smaller estimates are safer, but because smaller pieces flow better. They hit problems earlier. They learn faster. They deliver value sooner.</p><p>This is a genuine improvement in the way teams work, not an improvement in the way teams estimate. And it happens precisely because teams stopped spending energy on estimation and redirected that energy toward actually improving their work.</p><h4><strong>What to do instead</strong></h4><p>I&#8217;m not saying all estimation is worthless. Sometimes you need to make decisions that require rough forecasts. Should we start this project? Can we deliver by this deadline? How should we staff this team?</p><p>But these decisions don&#8217;t require story points. They require honest conversations about uncertainty. They require looking at historical data. They require acknowledging what you don&#8217;t know.</p><p>A team that says &#8220;based on how similar work has gone in the past, this might take two to four weeks, but there are significant unknowns around the external integration&#8221; is providing more useful information than a team that says &#8220;we estimated this at 23 points and our velocity is 8 points per sprint.&#8221;</p><p>The first statement is honest about uncertainty. The second creates false precision that will be treated as a commitment.</p><h4><strong>Have the conversations you actually want to have</strong></h4><p>Let me return to where I started. &#8220;Story points trigger conversations about effort and complexity.&#8221;</p><p>If you find yourself defending a practice by pointing to its side effects rather than its primary purpose, it&#8217;s worth asking whether you should replace that practice with something that achieves the side effects directly.</p><p>If you want conversations about effort, have conversations about effort.</p><p>If you want conversations about complexity, have conversations about complexity.</p><p>If you want shared understanding, build shared understanding through discussion, collaboration, and working together on the actual work.</p><p>You don&#8217;t need a numerical ritual to have good conversations. The ritual gets in the way more often than it helps. The numbers create more problems than they solve.</p><p>Stop estimating. Start understanding. Your team will be better for it.</p>]]></content:encoded></item><item><title><![CDATA[Non-Determinism is Not the Problem]]></title><description><![CDATA[The case against AI tools mistakes process variability for unreliability]]></description><link>https://a4al6a.substack.com/p/non-determinism-is-not-the-problem</link><guid isPermaLink="false">https://a4al6a.substack.com/p/non-determinism-is-not-the-problem</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Tue, 13 Jan 2026 23:18:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I keep hearing this argument: AI agents are non-deterministic, LLMs can produce different results when given the same problem, therefore the code they produce cannot be trusted.</p><p>This is a non-sequitur.</p><p>Let&#8217;s start with an observation that should be obvious but apparently isn&#8217;t. Humans are non-deterministic too. Give the same problem to two engineers and you get two different solutions. Give it to the same engineer six months later and you still get a different one. That has never stopped us from shipping software.</p><p>The implicit comparison being made is between LLMs and compilers. Compilers are deterministic: same input, same output, every time. LLMs are not. Therefore LLMs are unreliable. But this comparison misses the point entirely. We don&#8217;t compare developers to compilers. We don&#8217;t expect humans to produce identical output given identical input. <strong>We expect them to produce correct output, verified through review and testing.</strong></p><p>The same standard applies to AI tools.</p><p>Now, I&#8217;m not claiming LLMs fail the same way humans do. They don&#8217;t. LLMs hallucinate with confidence. They don&#8217;t self-correct the way a human rereads their own code and spots the mistake. They can produce plausible-looking nonsense that passes a casual glance. These are real failure modes, and they&#8217;re different from the failure modes we&#8217;re used to managing.</p><p>But this is precisely why we verify. This is why we test. This is why we review outputs rather than shipping them blindly. The failure modes are different, so our verification strategies need to account for them. That&#8217;s an engineering problem with engineering solutions.</p><p><strong>What actually matters is behaviour, not how we got there.</strong> We instruct agents, review the output, and rely on tests as executable statements of expected behaviour. As long as the external behaviour is preserved, the internal path taken to reach it is largely irrelevant. This is exactly why test suites exist and why refactoring is even possible.</p><p>But there&#8217;s a deeper confusion at play here, one about agency and control.</p><p>When people say &#8220;AI is non-deterministic&#8221;, they often mean something more like &#8220;AI does unpredictable things on its own&#8221;. As if we were talking about a colleague who happens to be made of silicon. We&#8217;re not. We&#8217;re talking about software that responds to inputs.</p><p>I&#8217;ll grant that this can feel like an understatement when you&#8217;re watching an agent loop operate over time, calling tools, maintaining context, producing outputs that seem to emerge rather than follow directly from a prompt. There is real complexity here, and reasoning about agent behaviour does add cognitive load. I&#8217;m not dismissing that.</p><p>But complexity is not the same as autonomy. These systems don&#8217;t have intentions, preferences, or agency. They have behaviours that emerge from inputs, context, and training. The complexity makes them harder to reason about, not impossible to control.</p><p><strong>And here&#8217;s what the non-determinism critics often miss: we have extensive control over what these tools do. Modern AI coding tools offer orchestration capabilities, subagents for specific tasks, and guardrails that constrain behaviour. You can define workflows, set boundaries, require confirmations. The tool doesn&#8217;t run wild. You direct it.</strong></p><p>This is no different from how we&#8217;ve always worked with complex systems. We don&#8217;t trust developers to never make mistakes. We build processes around them: social programming, code review, automated testing, continuous integration, staged deployments. The system as a whole produces reliable outcomes even though individual contributors are thoroughly non-deterministic.</p><p>The question isn&#8217;t whether an LLM will give you the same answer twice. The question is whether you have the discipline to verify outputs, the tests to catch regressions, and the workflows to maintain control.</p><p>Here&#8217;s where I need to be direct about something. This argument assumes a certain level of engineering maturity. Tests, reviews, guardrails, orchestration, discipline. Critics might respond that many teams don&#8217;t have this maturity, and that AI makes their problems worse.</p><p>They&#8217;re not wrong.</p><p><strong>AI is an amplifier.</strong> It amplifies whatever practices you already have. If you have strong testing habits, AI helps you move faster with confidence. If you ship code without verification, AI helps you ship bad code faster. This isn&#8217;t a reason to avoid AI tools. It&#8217;s a recognition that using them well requires the same foundations that any reliable software development requires. The bar hasn&#8217;t changed. The speed has.</p><p>This matters because language shapes thinking. If we talk about AI as if it has agency (yes, &#8220;agents&#8221; might be misleading), we start believing we can delegate not just tasks, but responsibility. We start thinking accountability can be outsourced to a machine.</p><p>It can&#8217;t. And this is where the human analogy reaches its limit. Humans aren&#8217;t just another non-deterministic process in the pipeline. We have intent, situational awareness, and moral judgement. AI tools have none of these. That&#8217;s precisely why responsibility remains with us. <strong>We&#8217;re not accountable because we&#8217;re in the loop. We&#8217;re accountable because we&#8217;re the only ones who can be.</strong></p><p>We chose to use the tool. We chose to accept the output. We chose to ship it, publish it, send it. The outcome is ours, good or bad.</p><p>So when I hear the non-determinism argument, it sounds less like a real concern and more like a category error. We&#8217;re comparing the variability of a generative process to the reliability of a deterministic one, and concluding that the former is inherently unsuitable for serious work.</p><p>But we&#8217;ve been doing serious work with non-deterministic processes forever. They&#8217;re called humans.</p><p>AI doesn&#8217;t do anything. We do things with AI. That&#8217;s not pedantry. That&#8217;s the difference between using a tool responsibly and hiding behind one.</p>]]></content:encoded></item><item><title><![CDATA[The Argument You're Winning Is the Wrong One]]></title><description><![CDATA[Why "AI won't replace engineers" defends against an attack nobody is mounting]]></description><link>https://a4al6a.substack.com/p/the-argument-youre-winning-is-the</link><guid isPermaLink="false">https://a4al6a.substack.com/p/the-argument-youre-winning-is-the</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Tue, 06 Jan 2026 12:44:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The debate about artificial intelligence and software engineering has calcified into two predictable camps. On one side, breathless prophets declare that programmers will be obsolete by next Tuesday. On the other, defensive engineers insist that AI will never replace them because software development requires creativity, judgment, and all manner of ineffable human qualities.</p><p>Both positions miss the point entirely. But the second one is particularly insidious because it sounds reasonable while defending against an attack that nobody serious is actually mounting.</p><p>When someone proclaims that &#8220;AI is not going to replace engineers,&#8221; they are constructing and demolishing a straw man. The interesting question was never about wholesale replacement. It never has been, for any technology, in any era. The real questions are far more uncomfortable, which is precisely why we avoid them.</p><p>Let me explain what I mean.</p><p>The replacement fantasy has ancient roots. When John McCarthy, Marvin Minsky, and their colleagues gathered at Dartmouth College in the summer of 1956, they proposed a study based on the conjecture that &#8220;every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it.&#8221; They estimated the project would take about two months with ten researchers.</p><p>This spectacular miscalculation set the template for AI discourse that persists to this day. The field has oscillated between manic overconfidence and depressive retrenchment ever since. The &#8220;AI winters&#8221; of the 1970s and late 1980s followed periods of grandiose promises that could not be kept. Expert systems were going to revolutionise everything until they didn&#8217;t. Neural networks were abandoned as dead ends until they weren&#8217;t.</p><p>Each cycle produces the same rhetorical pattern. Enthusiasts make sweeping claims about imminent transformation. Sceptics point to obvious limitations. The technology fails to deliver on the hype. Everyone declares AI overblown. Then, quietly, the technology improves, finds its niche, and reshapes things in ways nobody predicted.</p><p>The sceptics are always right in the short term and always wrong in the long term. The enthusiasts are always wrong in the short term and often right about directions if not timelines.</p><p>This brings us to the philosophical heart of the matter. The Hungarian-British polymath Michael Polanyi articulated what we now call Polanyi&#8217;s Paradox: &#8220;We can know more than we can tell.&#8221; Much of human expertise consists of tacit knowledge that we cannot easily articulate or formalise. Hubert Dreyfus, the philosopher who spent decades critiquing AI, built his case on similar foundations. He argued that human intelligence is fundamentally embodied, contextual, and resistant to the kind of rule-based formalisation that AI required.</p><p>These arguments were powerful. They were also, as it turned out, arguments against a particular approach to AI rather than against AI itself. The shift from symbolic AI to machine learning, from explicit programming to pattern recognition from data, sidestepped many of Dreyfus&#8217;s objections. He was right that you cannot write down rules for recognising faces or understanding natural language. He did not anticipate that you wouldn&#8217;t need to.</p><p>John Searle&#8217;s famous Chinese Room argument makes a different point. A person who follows rules to manipulate Chinese symbols, without understanding Chinese, does not thereby come to understand Chinese. The room processes the symbols correctly but comprehension is absent. Searle meant this as a refutation of strong AI claims about machine consciousness and understanding.</p><p>But here&#8217;s the uncomfortable truth: much of what we call software engineering does not require understanding in Searle&#8217;s deep sense. It requires symbol manipulation according to patterns. When an AI generates code that compiles, runs correctly, and solves the stated problem, whether it &#8220;understands&#8221; what it&#8217;s doing becomes philosophically interesting but practically irrelevant.</p><p>The defender of human engineers says: but what about the hard parts? What about architecture, requirements analysis, understanding the business domain? What about debugging subtle issues and making judgment calls under uncertainty?</p><p>These are fair points. They are also, almost certainly, temporary ones. The history of automation is the history of tasks that could never be automated until they were.</p><p>Consider the game of Go. For decades, it served as the exemplar of human intuition triumphing over brute computation. Chess might fall to Deep Blue, but Go was different. The branching factor was too large. The positional judgment was too subtle. Human intuition was irreplaceable. In 2016, AlphaGo defeated Lee Sedol, and that particular argument died.</p><p>The pattern repeats. We identify what makes current AI inadequate. We treat those limitations as fundamental. We build our professional identity on being the ones who can do what machines cannot. Then the machines learn to do it anyway.</p><p>This is where the straw man becomes dangerous. When engineers say &#8220;AI won&#8217;t replace us,&#8221; they are defending against total elimination. But total elimination is not what&#8217;s coming. What&#8217;s coming is something more like what happened to bank tellers, travel agents, and countless other professions that still exist but employ far fewer people at different tasks than before.</p><p>The economists have a term for this: labour-saving technology. When a technology allows one worker to produce what previously required five, you don&#8217;t need to replace all the workers. You just need fewer of them. The work still exists. Human judgment is still involved somewhere in the process. But four out of five workers are doing something else, if they&#8217;re lucky, or nothing at all, if they&#8217;re not.</p><p>The augmentation narrative that many engineers embrace is a comforting frame that obscures this dynamic. Yes, AI augments human capability. That&#8217;s precisely the problem. If I can now accomplish with AI assistance what previously required a team, the team is no longer required. The fact that a human remains in the loop is cold comfort to the humans removed from it.</p><p>Karl Marx, whatever one thinks of his prescriptions, was remarkably prescient about machinery and labour. He observed that machines do not simply make work easier. They transform the relationship between labour and capital. They create what he called a &#8220;reserve army&#8221; of unemployed workers whose presence disciplines those still employed. Technology under capitalism serves capital&#8217;s interests, not labour&#8217;s.</p><p>One need not be a Marxist to recognise the dynamic. When AI enthusiasts at technology companies promise that AI will augment developers, they are making a statement about capability. When they predict that this will not affect employment, they are making a statement about economics that does not follow from the first claim and that their own incentives make them unqualified to assess.</p><p>The philosophical tradition offers another useful concept here: the Ship of Theseus. If you replace every plank of a ship, one at a time, is it still the same ship? Apply this to software engineering. If AI takes over code generation, then testing, then documentation, then debugging, then architecture, each time with a human &#8220;in the loop&#8221; providing approval, at what point has the engineer been replaced while still technically being present?</p><p>The answer is that replacement is not a binary event but a gradient. You can be replaced by degrees. You can be diminished incrementally. You can find yourself nominally in control while actually serving as a rubber stamp for decisions made elsewhere. The straw man of total replacement distracts from this more likely and more insidious outcome.</p><p>Martin Heidegger wrote about technology as a mode of &#8220;revealing,&#8221; a way of disclosing the world that also conceals other possibilities. When we frame the question as &#8220;will AI replace engineers,&#8221; we reveal a binary choice and conceal the spectrum of possibilities in between. The framing itself does ideological work, making certain outcomes thinkable and others invisible.</p><p>What should we actually be asking? Not whether AI will replace engineers, but how AI will transform engineering, who will benefit from that transformation, what skills will become more or less valuable, and how we might influence these trajectories rather than simply reacting to them.</p><p>The Luddites, contrary to popular caricature, were not opposed to technology as such. They were skilled textile workers who objected to the use of machinery to circumvent labour standards and degrade working conditions. Their complaint was not that machines were bad but that machines were being used badly. History has vindicated the direction of their concerns while condemning the futility of their methods.</p><p>We might learn from them. The question is not whether to resist AI, which is as futile now as smashing looms was then. The question is whether to participate in shaping how AI transforms our profession or to console ourselves with the fairy tale that it won&#8217;t.</p><p>When someone tells you that AI won&#8217;t replace engineers, ask them what they mean. If they mean that engineers will not vanish overnight, they are trivially correct. If they mean that the number, nature, and compensation of engineering roles will remain unchanged, they are almost certainly wrong. If they mean that they personally will be fine, they may be right or may be engaging in the same wishful thinking that has consoled displaced workers throughout history.</p><p>The honest position is uncertainty. We don&#8217;t know exactly how this will unfold. The capabilities are advancing faster than most predictions. The integration into actual workflows is slower than the hype suggests. The economic incentives are clear but the timeline is not.</p><p>What we can say is that treating &#8220;AI won&#8217;t replace engineers&#8221; as a meaningful contribution to this conversation is an intellectual abdication. It&#8217;s a thought-terminating clich&#233; that protects us from harder questions. It is the equivalent of standing on the shore, watching the tide come in, and reassuring ourselves that water cannot replace sand.</p><p>The sand will still be there when the tide recedes. It will just be in a very different place.</p>]]></content:encoded></item><item><title><![CDATA[Friday Is Not the Problem]]></title><description><![CDATA[How a sensible-sounding rule keeps teams from fixing what is actually broken]]></description><link>https://a4al6a.substack.com/p/friday-is-not-the-problem</link><guid isPermaLink="false">https://a4al6a.substack.com/p/friday-is-not-the-problem</guid><dc:creator><![CDATA[Andrea Laforgia]]></dc:creator><pubDate>Wed, 31 Dec 2025 22:58:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JZMQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b6085b-e5fb-4331-9ec4-f59cc43a2546_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div><hr></div><p>The phrase &#8220;don&#8217;t deploy on Friday&#8221; has become one of those pieces of engineering wisdom that people repeat without much examination. It sounds sensible. It feels responsible. And it is almost entirely wrong.</p><p>I recently read a post defending this rule. It was well-written and made several arguments that, on the surface, seem reasonable. But each of these arguments, when examined closely, reveals a deeper problem that the Friday rule does not solve but merely hides.</p><p>Let me address them one by one.</p><p><strong>&#8220;You rarely have full control over everything involved in a deployment&#8221;</strong></p><p>The argument goes like this: infrastructure, third-party services, configuration differences, traffic patterns, and environment-specific behaviour can all surface issues only after a change is actually applied. These are not always predictable in staging or test environments.</p><p>This is true. It is also completely irrelevant to the question of which day you deploy.</p><p>If your production environment behaves differently from staging in ways that cause incidents, that is a problem on Tuesday just as much as on Friday. The unpredictability does not magically disappear because you have five working days ahead of you. What changes is your ability to throw human hours at the problem.</p><p>The correct response to environmental unpredictability is not to limit your deployment days. It is to invest in observability, automated rollbacks, and production testing strategies like canary deployments and feature flags. These approaches reduce blast radius and enable fast recovery regardless of when something goes wrong.</p><p>If your only mitigation strategy for environmental surprises is &#8220;have people available to fix things manually&#8221;, you have not solved the problem. You have just accepted it.</p><p><strong>&#8220;Data migrations are fundamentally different and risky&#8221;</strong></p><p>The post argues that changes touching real data can be slow, irreversible, dependent on data volume and shape, and risky to roll back. Treating them the same as a small code deploy ignores real operational complexity.</p><p>I completely agree. Data migrations are different and require special handling.</p><p>But this is an argument for treating data migrations differently. It is not an argument for a blanket rule about which days any deployment can happen.</p><p>A well-designed migration strategy includes expand-and-contract patterns, backward-compatible schema changes, and the ability to run migrations independently from code deployments. If your migrations are so risky that they require days of potential cleanup time, the problem is not Friday. The problem is how you are doing migrations.</p><p>Conflating &#8220;don&#8217;t deploy risky migrations on Friday&#8221; with &#8220;don&#8217;t deploy anything on Friday&#8221; is sloppy thinking. It takes a legitimate concern about one specific type of change and inflates it into a universal policy that restricts all changes. This is how reasonable caution turns into irrational fear.</p><p><strong>&#8220;Limited availability of teammates and reduced support coverage&#8221;</strong></p><p>This is perhaps the most honest of the arguments, and it deserves a direct response.</p><p>The concern is that releasing late on a Friday means fewer people are available if something goes wrong. Response times will be slower. The on-call engineer might be alone.</p><p>But notice what this argument assumes: deployments are expected to cause problems that require human intervention. This assumption is so deeply embedded that it passes without comment. It is treated as a law of nature.</p><p>It is not. It is a symptom of an immature deployment process.</p><p>When deployment is a non-event, when changes are small, well-tested, and automatically rolled back if metrics degrade, you do not need a war room on standby. The system handles problems faster than humans could anyway.</p><p>The fear of reduced weekend coverage only makes sense if you expect to need that coverage. And if you expect to need it, that expectation should be treated as a problem to solve, not a constraint to work around.</p><p><strong>&#8220;Mature engineering teams optimise for reliability, not bravado&#8221;</strong></p><p>This framing is clever but misleading. It positions Friday deployments as reckless showmanship and the Friday rule as careful professionalism.</p><p>The reality is the opposite.</p><p>Truly mature teams do not need deployment windows because they have invested in making deployment boring. They deploy dozens of times per day. Each individual deployment is so small and so well-tested that it carries almost no risk. They can deploy on Friday afternoon and go home without anxiety because they have built the systems and practices that make this possible.</p><h3><strong>Avoiding Friday deployments is not a sign of maturity. It is a sign that you have accepted an immature deployment process as permanent.</strong></h3><p>The most reliable teams I have worked with never talked about which days were safe to deploy. The question simply did not arise. Deployment was like committing code or running tests: something you do continuously, without ceremony, because your systems are designed to handle it safely.</p><p><strong>&#8220;Consider blast radius, rollback strategies, observability, and on-call load&#8221;</strong></p><p>Absolutely. Consider all of these things. But consider them as engineering problems to solve, not as reasons to restrict when you can ship.</p><p>Blast radius should be limited by deploying incrementally, using feature flags, and rolling out to small percentages of traffic before going wide. Rollback should be automated and triggered by metric degradation, not by a human noticing a problem and deciding to act. Observability should be good enough that problems are detected in seconds, not hours. On-call load should be reduced by building systems that recover automatically.</p><p>If you have all of these things in place, Friday is just another day. If you do not have them in place, the Friday rule is a bandage over a wound that will keep bleeding every other day of the week.</p><p><strong>&#8220;Good software engineering is about making smart trade-offs&#8221;</strong></p><p>I agree completely. But the smart trade-off is not between deploying on Friday and waiting until Monday.</p><p>The smart trade-off is between maintaining a dysfunctional deployment process and investing in one that actually works. One of these trade-offs solves the problem. The other just delays it until after the weekend.</p><h4><strong>The expanding forbidden zone</strong></h4><p>There is a pattern I have seen repeatedly in teams that adopt the Friday rule. First it is &#8220;don&#8217;t deploy on Friday&#8221;. Then it becomes &#8220;don&#8217;t deploy on Friday afternoon&#8221;. Then &#8220;don&#8217;t deploy after Thursday lunch, just to be safe&#8221;. Then &#8220;don&#8217;t deploy the day before a bank holiday&#8221;. Then &#8220;don&#8217;t deploy during the Christmas period&#8221;, which somehow stretches from early December to mid-January.</p><p>The forbidden zone keeps expanding because the underlying problem remains unaddressed. You have not made deployments safer. You have just reduced the number of days when you are willing to confront how unsafe they are.</p><h4><strong>The rule as technical debt</strong></h4><p>Here is the uncomfortable truth: the &#8220;don&#8217;t deploy on Friday&#8221; rule is a form of technical debt.</p><p>Every time you follow it, you reinforce the idea that deployments are inherently dangerous. You avoid the feedback that would push you to improve. You normalise a dysfunctional process instead of fixing it. The rule does not make you safer. It makes you complacent.</p><p>If your deployments are scary, the solution is not to do them less often. It is to do them more often, in smaller increments, with better automation, faster feedback, and reliable rollback. The fear you feel about Friday deployments should be treated as a signal, not as wisdom to be enshrined in team policy.</p><h4><strong>This is an invitation, not a challenge</strong></h4><p>I want to be clear about what I am not saying. I am not saying &#8220;drop everything and push to prod this Friday&#8221;. Those of us who encourage Friday deployments are not doing it to be provocative or to score points.</p><p>We do it because we have seen what it looks like when teams feel safe enough to do it. We have lived through the pain of not being able to. And we have seen the benefits of moving past that fear.</p><p>I am aware of how many invisible blockers stand in the way. Fear of being blamed if something goes wrong. A belief that &#8220;this is just how we do things&#8221;. Lack of confidence in the release process. Unclear ownership. The absence of good tooling and observability. These are real obstacles, and dismissing them as excuses would be unfair.</p><p>But accepting the Friday rule means accepting these blockers as permanent. It means letting fear drive your decisions and treating the status quo as inevitable.</p><h4><strong>Where the real work begins</strong></h4><p>So no, you do not need to deploy this Friday. But imagine a world where you could, and it was not a big deal.</p><p>Then ask yourself: what would need to be true for that to happen? What stands in the way?</p><p>That question is where the real work begins. And answering it honestly will do far more for your reliability than any rule about which days are safe to ship.</p>]]></content:encoded></item></channel></rss>