<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Middleton's Musings]]></title><description><![CDATA[Medicine, AI and decision support, knowledge engineering, digital health.]]></description><link>https://bmiddleton1.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!KMFe!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc54ae8cf-4907-48fa-ad2c-62bd601a1ceb_1024x1024.png</url><title>Middleton&apos;s Musings</title><link>https://bmiddleton1.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 14:54:51 GMT</lastBuildDate><atom:link href="/__u/bmiddleton1.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Blackford Middleton]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[bmiddleton1@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[bmiddleton1@substack.com]]></itunes:email><itunes:name><![CDATA[Blackford Middleton, MD, MPH]]></itunes:name></itunes:owner><itunes:author><![CDATA[Blackford Middleton, MD, MPH]]></itunes:author><googleplay:owner><![CDATA[bmiddleton1@substack.com]]></googleplay:owner><googleplay:email><![CDATA[bmiddleton1@substack.com]]></googleplay:email><googleplay:author><![CDATA[Blackford Middleton, MD, MPH]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Fate of the Universe Rests on Good Ethics]]></title><description><![CDATA["The fate of the universe rests on good ethics."]]></description><link>https://bmiddleton1.substack.com/p/the-fate-of-the-universe-rests-on</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-fate-of-the-universe-rests-on</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Sat, 08 Aug 2026 20:25:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_vKF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_vKF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_vKF!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!_vKF!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!_vKF!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_vKF!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_vKF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!_vKF!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!_vKF!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!_vKF!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_vKF!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7e9b5b4-b530-446d-bcc2-7925579046b0_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>"The fate of the universe rests on good ethics."</p><p>At first glance, the statement sounds like a line pulled from the climax of a space opera&#8212;a dramatic, philosophical platitude. But if you strip away the melodrama and examine the trajectory of human capability through a cold, empirical lens, the premise stops being fiction and becomes a mathematical certainty.</p><p></p><p>Ethics is not a soft science. It is the fundamental algorithm for long-term survival in a universe governed by entropy and compounding scale.</p><p><strong>The Clinical Microcosm</strong></p><p>To understand the cosmic scale, we have to start with the biological one. Throughout my time in clinical medicine and scientific research, I learned early on that the power to heal is inextricably linked to the power to destroy.</p><p></p><p>When you stand at the bedside of a patient or analyze data that could alter a therapeutic paradigm, you realize that the Hippocratic obligation is not some abstract, polite suggestion. It is a rigorous protocol designed to prevent catastrophe. We hold life and death in a scalpel, a pharmaceutical compound, or a CRISPR sequence. Without a rock-solid ethical framework, science is just unguided ballistics. We rely on ethics to dictate when to intervene, when to step back, and how to balance the dual edges of efficacy and toxicity.</p><p></p><p>In the lab, a single compromised variable or a manipulated data point doesn't just ruin an experiment; it corrupts the literature, misdirects future research, and ultimately harms patients. Ethics, in this domain, is synonymous with accuracy and survival.</p><p><strong>The Leverage of the Entrepreneur</strong></p><p>Medicine and bench science are often localized, but innovation and entrepreneurship are where we learn about scale. Building a company or driving a new technology to market teaches you a brutal lesson about complex systems: <strong>initial conditions dictate final outcomes.</strong></p><p></p><p>When you build a system&#8212;whether it&#8217;s a technological platform, a healthcare startup, or a new economic model&#8212;you encode your values into its architecture. If those foundational values are misaligned, or if the incentives reward exploitation over sustainability, those errors compound exponentially.</p><p></p><p>We are currently scaling technologies with unprecedented, civilization-altering leverage. Artificial intelligence, synthetic biology, and planetary geoengineering are not just iterative tools; they are force multipliers. An entrepreneur without an ethical compass building a local business is a nuisance. An entrepreneur without an ethical compass scaling artificial general intelligence is an existential threat.</p><p><strong>The Great Filter is Ethical, Not Technological</strong></p><p>Now, extrapolate that leverage to the cosmos.</p><p>Look up at the night sky and consider the Fermi Paradox: if the universe is so vast and old, where is everybody? The "Great Filter"&#8212;the barrier that prevents civilizations from becoming galaxy-spanning&#8212;is unlikely to be a gamma-ray burst or an errant asteroid. It is almost certainly a failure of ethics.</p><p></p><p>A civilization inevitably reaches a critical threshold where its technological capacity to destroy itself outpaces its neurological and cultural capacity for restraint. We are approaching that exact bottleneck. In their book <em>Unleashing the Killer App</em>, Larry Downes and Chunka Mui illustrated this vulnerability with what they called the Law of Disruption. They presented a stark graph demonstrating that while technological capabilities accelerate exponentially, our social, cultural, and legal systems adapt only incrementally.</p><p></p><p>Today, we are watching that gap widen into a chasm. Artificial intelligence represents the most accelerated technological change we have ever witnessed, driving the technology curve vertical while our shared community vision and institutional frameworks trail behind. This widening delta puts extraordinary, unprecedented pressure on moral and ethical formation across society. We have Paleolithic emotions, medieval institutions, and god-like technology. If our ethical frameworks do not evolve rapidly enough to bridge that gap, our technological adolescence will be terminal.</p><p></p><p>The fate of the universe&#8212;or at least the fate of consciousness within it&#8212;rests on ethics because intelligence without constraints is self-terminating. Good ethics are the guardrails that prevent a high-energy civilization from collapsing under the weight of its own power.</p><p></p><p>We don't need a softer approach to the future; we need a more rigorous one. It is time to treat ethical alignment not as a philosophical luxury, but as the ultimate engineering constraint. If we get the ethics right, we earn the future and the stars that come with it. If we don't, we are just a brief, self-terminating anomaly.</p>]]></content:encoded></item><item><title><![CDATA[The Overseer’s Dilemma: Clinical AI and the Mode-Switch Problem, Relocated]]></title><description><![CDATA[Part 2 of: The 300-Millisecond Myth: Thinking Fast, Slow, and the Design of Clinical Decision Support]]></description><link>https://bmiddleton1.substack.com/p/the-overseers-dilemma-clinical-ai</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-overseers-dilemma-clinical-ai</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Fri, 31 Jul 2026 18:18:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!unBy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!unBy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!unBy!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!unBy!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!unBy!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!unBy!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!unBy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2577543,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://bmiddleton1.substack.com/i/209289851?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!unBy!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!unBy!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!unBy!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!unBy!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F243987c7-f7f6-4472-96c2-5d597448eb5b_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Part 1 argued that we cannot summon a clinician&#8217;s deliberative reasoning on demand by putting a box on the screen. The mode switch from fast to slow is neither free nor instantaneous, and interruptive decision support that assumes otherwise either exhausts the deliberative system or teaches it to look away.</p><p>Generative AI does not solve that problem. It relocates it. The old design tried to flip a switch inside the clinician. The new one moves the deliberation out of the clinician entirely &#8212; into a model that purports to reason &#8212; and hands the clinician a different job: watch the machine and catch it when it is wrong. That job has its own cognitive signature, and its own way of failing.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The frame, briefly</h2><p>The useful move from Part 1 was to stop asking <em>how do we interrupt?</em> and start asking <em>which cognitive mode does this task run in, and how do we support good performance in that mode?</em> Type 1 is fast, autonomous, and effortless. Type 2 is slow, effortful, and bounded by working memory (Kahneman, 2011; Evans &amp; Stanovich, 2013). Good design matches the modality to the mode instead of fighting it.</p><p>Clinical AI scrambles that mapping in a specific way. For the first time, the decision support does not merely nudge the clinician&#8217;s own reasoning &#8212; it offers a finished piece of reasoning of its own. The design question is no longer only about the clinician&#8217;s cognition. It is about the handoff between two reasoners, one human and one synthetic.</p><h2>What changes when the machine is the deliberator</h2><p>Rule-based decision support was transparent and dumb. It fired a logic statement, and the clinician supplied all the judgment &#8212; which is exactly why clinicians learned to dismiss it, overriding drug-safety alerts in 49% to 96% of cases (van der Sijs et al., 2006). Generative AI inverts the arrangement: the model supplies an answer that <em>looks</em> like judgment, and the clinician&#8217;s role narrows to oversight &#8212; accept, edit, or reject.</p><p>Oversight sounds like Type 2. It is supposed to be the careful, effortful check. But sustained vigilance over a mostly-correct automated partner is one of the most reliably failure-prone tasks in human factors. When the machine is usually right, the human stops looking hard. Attention drifts. The check becomes a formality. This is automation bias and its quieter cousin, automation complacency: over-reliance on the system&#8217;s output, accepting a recommendation without the scrutiny it was meant to receive (Goddard et al., 2012).</p><p>Here is the trap in one sentence. We are asking clinicians to perform deliberate, skeptical Type 2 oversight of a system engineered to make that oversight feel unnecessary. The better the model, the stronger the pull toward the reflexive accept &#8212; and the more consequential the rare miss.</p><h2>The failure is not hypothetical</h2><p>This is where I want to be evidence-led rather than alarmed, because the temptation on all sides is to argue from anecdote. The controlled studies already exist, and they point the same direction.</p><p>Clinicians are susceptible to a decision aid&#8217;s advice, and the <em>quality</em> of the advice moves them more than its labeled source. Gaube and colleagues had radiologists and physicians read chest films with suggestions attached; poor advice degraded their diagnoses, whether the advice was labeled as coming from a human expert or an AI (Gaube et al., 2021). The influence is real, and it does not depend on trusting &#8220;the algorithm&#8221; in the abstract.</p><p>The effect reaches experienced readers. In a mammography experiment, radiologists&#8217; judgments moved toward the AI&#8217;s suggested BI-RADS category &#8212; and when the suggestion was wrong, accuracy fell, across experience levels (Dratsch et al., 2023). Expertise did not confer immunity.</p><p>It scales to diagnosis. In a randomized vignette study, a standard AI model with explanations improved clinicians&#8217; diagnostic accuracy &#8212; but a <em>systematically biased</em> model reduced it, and providing the model&#8217;s explanations did not fully undo the damage (Jabbour et al., 2023). That last finding deserves weight: the interpretability feature we lean on as a safeguard did not rescue clinicians from a confidently wrong machine.</p><p>Collaboration is genuinely double-edged. In skin-cancer recognition, AI support raised average diagnostic accuracy &#8212; yet faulty AI could pull good clinicians <em>below</em> their unaided baseline (Tschandl et al., 2020). The same tool that helps on average harms in the tail.</p><p>And the risk is not only per-decision. It compounds. In a multicenter observational study, endoscopists&#8217; detection performance on standard, unaided colonoscopy declined after a period of routine exposure to AI assistance &#8212; a signal consistent with de-skilling (Budzy&#324; et al., 2025). Lean on the machine long enough and the human capability you were counting on as the backstop quietly erodes.</p><p>None of these studies says AI decision support is bad. They say something more precise: the oversight model fails in predictable ways, the failures reach experts, explanations do not fully protect, and the human safety net can thin with use.</p><h2>The bull case, honestly</h2><p>Now the other side, because the adjudicator&#8217;s job is to state it at full strength.</p><p>The diagnostic capability is real and improving fast. In simulated consultations, a conversational diagnostic AI conducted history-taking and reached diagnoses at a level competitive with primary care physicians on the study&#8217;s measures (Tu et al., 2025). Set aside the horse-race framing and the substance remains: these systems are becoming genuinely good at the deliberative work &#8212; the differential, the structured history &#8212; that Part 1 reserved for scarce human Type 2.</p><p>If the model&#8217;s deliberation is often better than the overworked clinician&#8217;s, then routing the Type 2 work to the machine is not obviously a mistake. It may be the point. A rested, consistent, exhaustive synthetic reasoner that never suffers from anchoring at 2 a.m. is a serious proposition, not a marketing slogan. Any honest account has to hold that possibility open.</p><h2>NB: Here is where I want to be careful</h2><p>So we have two true things that pull against each other. Clinical AI can raise the ceiling on diagnostic reasoning. And clinical AI can lower the floor, by degrading the human check and eroding the human skill that check depends on.</p><p>The resolution is not to pick a side. It is to notice that the two effects live in different regions. AI helps most where the human was going to be worse &#8212; fatigued, at the edge of knowledge, working an unfamiliar presentation. AI hurts most in a narrow, dangerous band: where the model is good enough to be trusted but wrong in a way the overseer cannot detect. Jabbour&#8217;s biased-model result maps that band exactly. The danger is not the obviously broken AI, which clinicians catch. It is the plausible, confident, subtly wrong AI, wrapped in an explanation that reads as reasoning.</p><p>That reframes the design goal. The task is not to maximize the model&#8217;s accuracy in the average case. It is to preserve the human&#8217;s ability to catch the model in the tail &#8212; which means preserving attention, skill, and the genuine option to disagree. Every one of those is exactly what a smooth, high-accuracy assistant erodes.</p><h2>Three ways to deploy clinical AI</h2><p>The fast/slow lens still does the work. It just now describes the <em>human&#8217;s</em> residual role under three deployment patterns.</p><p><strong>AI as default-shaper.</strong> Ambient documentation that drafts the note, autocomplete that pre-fills the order, the summary that leads with a recommendation. These design <em>with</em> fast processing &#8212; they remove friction, which is precisely why they carry the highest automation-bias risk. When the artifact arrives pre-made and mostly right, review collapses into a signature. If you deploy AI this way, you have to reintroduce friction deliberately at the points that matter: force a verification step where the error would be consequential, the way bar-code administration forces a check rather than trusting recall.</p><p><strong>AI as deliberation scaffold.</strong> The model that surfaces a differential you had not considered, or flags the disconfirming finding you were about to anchor past. This is the mode that respects Type 2 instead of replacing it &#8212; it makes the clinician think, rather than think less. The collaboration evidence suggests this is where the durable gains live, provided the tool is built to provoke reconsideration rather than to hand over a verdict (Tschandl et al., 2020).</p><p><strong>AI as autonomous agent.</strong> Agentic systems that place orders, draft prior authorizations, or move a workflow forward with the human only nominally in the loop. Here the mode switch is not relocated &#8212; it is delegated, and oversight becomes episodic at best. That may be defensible for low-stakes, high-volume administrative work. It is a governance decision, not a UX detail, and you should make it as one.</p><h2>A lighter pass across the classes</h2><p><strong>Ambient documentation.</strong> The fastest-spreading deployment and the clearest default-shaper. The draft note is a Type 1 artifact handed to a clinician with no time to read it closely. The failure mode is the plausible fabrication or the quietly wrong dose that survives to the signed record. Design implication: targeted friction on the high-harm fields, not a blanket trust in the draft.</p><p><strong>Diagnostic support.</strong> Best cast as a scaffold, worst cast as an oracle. The evidence says a good model helps and a biased one hurts, and explanations do not fully protect (Jabbour et al., 2023). Design implication: surface reasoning and alternatives to be interrogated, and resist the interface that presents a single confident answer.</p><p><strong>Agentic ordering and prior authorization.</strong> Largely administrative, largely tolerable to automate &#8212; the referral and paperwork end of Part 1&#8217;s spectrum. Design implication: automate the pathway, but instrument it, because &#8220;the human is in the loop&#8221; is a claim to be measured, not assumed.</p><p><strong>Imaging triage and second read.</strong> The setting where automation bias is best documented and reaches experts (Dratsch et al., 2023). Design implication: guard the tail. Preserve unaided reads, monitor for skill drift, and treat the AI as a prompt to look again &#8212; not as the finding.</p><h2>Distribution, the care team, and liability</h2><p>Part 1 flagged three limits and deferred them here. AI sharpens all three.</p><p>The benefit and the burden fall unevenly. AI decision support helps most where expertise is thinnest &#8212; which is also where the safety net of a strong unaided clinician is weakest, and where de-skilling bites hardest (Budzy&#324; et al., 2025). The tool that lifts an under-resourced setting can, over time, hollow out the very capability that made it safe. That is a distributional question, not a technical one.</p><p>The care team, not the ordering physician, increasingly holds the check. Nurses, pharmacists, and technologists sit at the AI handoffs now &#8212; the ambient note, the flagged interaction, the triaged image. The oversight burden lands on them, often without the authority or the time to exercise it. Design that ignores where the check actually happens will misplace the friction.</p><p>And liability remains unsettled. &#8220;The AI recommended it&#8221; is not yet a defense, and &#8220;the clinician should have caught it&#8221; sits uneasily against an interface engineered to make catching it unlikely. That tension shapes behavior now, ahead of the case law, and it deserves to be named rather than wished away.</p><h2>The governance turn</h2><p>The design question for a clinical AI portfolio is not <em>how accurate is the model?</em> It is <em>is the human in the loop a genuine check or a rubber stamp?</em> &#8212; and that is measurable.</p><p>Stop measuring override rates and start measuring two other things. First, the <strong>agreement rate</strong>: how often the clinician&#8217;s final decision matches the model&#8217;s suggestion. Second, the <strong>detected-error rate</strong>: how often the clinician catches and corrects a model error, ideally seeded with known-wrong cases to calibrate. Read them together. A near-100% agreement rate paired with a near-zero detected-error rate does not show a great model. It shows oversight that has become theater. Where you find that pattern, you have two honest choices: rebuild the friction so the human can actually check, or admit the task is automated and govern it as automation &#8212; with monitoring, accountability, and a named owner &#8212; rather than pretending a human is meaningfully in the loop.</p><h2>What this means</h2><p>Part 1 ended with a warning: you cannot flip a clinician from fast to slow reasoning with a box on a screen. Part 2 adds the sequel: you cannot make oversight real by declaring it. Move the deliberation into the machine and the human&#8217;s job becomes vigilance over a partner built to make vigilance feel unnecessary &#8212; and vigilance, unsupported, is the task humans perform worst.</p><p>This sits inside a larger wager I have been working through elsewhere: whether AI can end the productivity stall that has followed nearly every general-purpose technology since 1973, computing included. That literature has its own version of this essay&#8217;s caution. Brynjolfsson&#8217;s J-curve describes why: measured productivity dips while the complementary investments &#8212; new workflows, retrained staff, redesigned organizations &#8212; catch up to the technology, and the gain shows up in the numbers only after that unglamorous work is done, if it is done at all. Clinical AI&#8217;s overseer problem is that same dip, seen from inside a single workflow instead of a national accounts table. The floor drops first, on the way to wherever the ceiling turns out to be.</p><p>I think the direction is right and the ceiling is real &#8212; that is the accelerationist case, and on clinical reasoning specifically, the evidence in this essay supports it. But the floor is optional, and we lower it every time we deploy a smooth, confident assistant and call the leftover human glance &#8220;oversight.&#8221; The plateau in clinical reasoning that AI might lift is genuine. Whether we clear it, or quietly automate our way past the point where anyone is still checking, is a design and governance choice &#8212; not an inevitability. It is the same verdict I keep reaching at the macro level, applied here at the scale of a single alert: directionally right, wrong to treat as inevitable, and silent on who keeps the floor and who loses it. That last question is not a footnote. It is the one this series has to answer.</p><p>So the question for anyone deploying clinical AI: <strong>when your clinicians disagree with the model, can they still tell that they should &#8212; and have you measured whether they ever do?</strong></p><h2>Bibliography</h2><p><span>Budzy&#324; K, Roma&#324;czyk M, Kitala D, Ko&#322;odziej P, Bugajski M, Adami HO, Blom J, Buszkiewicz M, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. </span><em><span>Lancet Gastroenterol Hepatol.</span></em><span> 2025;10(10):896&#8211;903. </span><a href="https://doi.org/10.1016/S2468-1253(25)00133-5"><span>https://doi.org/10.1016/S2468-1253(25)00133-5</span></a></p><p><span>Dratsch T, Chen X, Rezazade Mehrizi M, Kloeckner R, M&#228;hringer-Kunz A, P&#252;sken M, Bae&#223;ler B, Sauer S, et al. Automation bias in mammography: the impact of artificial intelligence BI-RADS suggestions on reader performance. </span><em><span>Radiology.</span></em><span> 2023;307(4):e222176. </span><a href="https://doi.org/10.1148/radiol.222176"><span>https://doi.org/10.1148/radiol.222176</span></a></p><p><span>Evans JSBT, Stanovich KE. Dual-process theories of higher cognition: advancing the debate. </span><em><span>Perspect Psychol Sci.</span></em><span> 2013;8(3):223&#8211;241. </span><a href="https://doi.org/10.1177/1745691612460685"><span>https://doi.org/10.1177/1745691612460685</span></a></p><p><span>Gaube S, Suresh H, Raue M, Merritt A, Berkowitz SJ, Lermer E, Coughlin JF, Guttag JV, et al. Do as AI say: susceptibility in deployment of clinical decision-aids. </span><em><span>NPJ Digit Med.</span></em><span> 2021;4(1):31. </span><a href="https://doi.org/10.1038/s41746-021-00385-9"><span>https://doi.org/10.1038/s41746-021-00385-9</span></a></p><p><span>Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. </span><em><span>J Am Med Inform Assoc.</span></em><span> 2012;19(1):121&#8211;127. </span><a href="https://doi.org/10.1136/amiajnl-2011-000089"><span>https://doi.org/10.1136/amiajnl-2011-000089</span></a></p><p><span>Jabbour S, Fouhey D, Shepard S, Valley TS, Kazerooni EA, Banovic N, Wiens J, Sjoding MW. Measuring the impact of AI in the diagnosis of hospitalized patients: a randomized clinical vignette survey study. </span><em><span>JAMA.</span></em><span> 2023;330(23):2275&#8211;2284. </span><a href="https://doi.org/10.1001/jama.2023.22295"><span>https://doi.org/10.1001/jama.2023.22295</span></a></p><p><span>Kahneman D. </span><em><span>Thinking, Fast and Slow.</span></em><span> New York: Farrar, Straus and Giroux; 2011.</span></p><p><span>Tschandl P, Rinner C, Apalla Z, Argenziano G, Codella N, Halpern A, Janda M, Lallas A, et al. Human&#8211;computer collaboration for skin cancer recognition. </span><em><span>Nat Med.</span></em><span> 2020;26(8):1229&#8211;1234. </span><a href="https://doi.org/10.1038/s41591-020-0942-0"><span>https://doi.org/10.1038/s41591-020-0942-0</span></a></p><p><span>Tu T, Schaekermann M, Palepu A, Saab K, Freyberg J, Tanno R, Wang A, Li B, et al. Towards conversational diagnostic artificial intelligence. </span><em><span>Nature.</span></em><span> 2025;642(8067):442&#8211;450. </span><a href="https://doi.org/10.1038/s41586-025-08866-7"><span>https://doi.org/10.1038/s41586-025-08866-7</span></a></p><p><span>van der Sijs H, Aarts J, Vulto A, Berg M. Overriding of drug safety alerts in computerized physician order entry. </span><em><span>J Am Med Inform Assoc.</span></em><span> 2006;13(2):138&#8211;147. </span><a href="https://doi.org/10.1197/jamia.M1809"><span>https://doi.org/10.1197/jamia.M1809</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The 300-Millisecond Myth: Thinking Fast, Slow, and the Design of Clinical Decision Support]]></title><description><![CDATA[Part 1 of two.]]></description><link>https://bmiddleton1.substack.com/p/the-300-millisecond-myth-thinking</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-300-millisecond-myth-thinking</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Fri, 31 Jul 2026 18:13:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LUtb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LUtb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LUtb!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!LUtb!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!LUtb!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LUtb!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LUtb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2577543,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://bmiddleton1.substack.com/i/209275879?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!LUtb!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!LUtb!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!LUtb!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LUtb!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ab58b6-e04d-450b-8b55-24aaff0664ef_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Part 1 of two.<br><br>The claim circulates in health IT circles with the authority of a physical constant: the human brain needs 300 milliseconds to switch from automatic to deliberate processing. Build decision support around that number, the story goes, and you can interrupt a clinician at just the right moment to force a considered choice.<br><br>The problem is that the number is not what the science says. The 300-millisecond figure most plausibly traces to the P300 event-related potential or to Libet-style readiness-potential findings &#8212; markers of attention and awareness, not the duration of deliberative reasoning. Genuine Type 2 deliberation runs over seconds to minutes, not milliseconds. Croskerry himself cautions that &#8220;speed alone does not distinguish&#8221; the two systems and that &#8220;the majority of decisions in clinical medicine are not dependent on very short response times.&#8221;<br><br>So the design premise built on that number is fragile. But the deeper mistake is not the specific duration. It is the assumption that the mode switch from fast to slow can be summoned on demand at all.<br><br>The frame</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><br>Kahneman&#8217;s dual-process formulation is familiar: Type 1 is fast, autonomous, and effortless. Type 2 is slow, effortful, and bounded by working memory. Evans and Stanovich sharpen the distinction: the defining difference is not speed but autonomy and working-memory load. Type 1 runs without supervision; Type 2 requires attention and mental space.<br><br>The design question, then, is not how do we interrupt? It is which cognitive mode does this task run in, and how do we support good performance in that mode? Good design matches the modality to the mode instead of fighting it.<br><br>What happens when we fight it</p><p><br>Rule-based clinical decision support (CDSS) assumes that an interruptive alert can flip a clinician from Type 1 to Type 2. The evidence says otherwise. Van der Sijs and colleagues found that clinicians overrode drug-safety alerts in 49% to 96% of cases. The overrides were not random errors of stubbornness. They were the predictable result of asking a fast system to do a slow system&#8217;s job.<br><br>The mode switch is costly. It takes time, burns working memory, and degrades performance on the primary task. When the switch is forced repeatedly, clinicians adapt by swatting the alert away &#8212; a learned Type 1 response to a system that keeps demanding Type 2. The alert fatigue literature documents this exhaustively: override rates rise, alert value falls, and eventually the system is tuned down or turned off.<br><br>This is not a compliance problem. It is a cognitive-mode mismatch.<br><br>A scannable subgroup index<br>Not all CDSS is interruptive. The field has evolved a spectrum of approaches, each with a different relationship to cognitive mode:<br><br>- CPOE with basic dosing alerts: Interruptive, rule-based. High override rates. Mode mismatch.<br>- Advanced CPOE with tiered alerts: Attempts to match urgency to interruption. Better, but still fundamentally interruptive.<br>- BCMA (barcode medication administration): Workflow-embedded. Checks at the point of action. Less interruptive, more ambient.<br>- Smart infusion pumps: Dose-error reduction built into the device. Mode-matched to the task.<br>- Medication reconciliation tools: Cross-setting comparison. Requires deliberate attention, but at a workflow-appropriate moment.<br>- Diagnostic decision support: Differential generators, image analysis. Mixed evidence; depends on whether the tool augments or replaces clinician reasoning.<br>- Ambient clinical intelligence: Documentation support. Largely Type 1-compatible if it reduces burden; risky if it inserts unverified content.<br>- Predictive analytics / sepsis alerts: Risk scores and early warning. High false-positive rates teach dismissal.<br>- AI-generated draft orders or notes: The new frontier. Mode-switch problem relocated, not solved.<br><br>The adjudicator&#8217;s turn<br>The standard critique of CDSS stops at the failure: alerts don&#8217;t work, clinicians override them, the system is ignored. But that critique misses what the fast system does well. Type 1 is not a bug to be eliminated. It is the cognitive substrate of expertise &#8212; pattern recognition, gestalt, the trained intuition that lets an experienced clinician walk into a room and know something is wrong before the data are in.<br><br>Bright and colleagues&#8217; systematic review found that CDSS improved practitioner performance in 64% of studies and clinical outcomes in 42%. The process-outcome gap is instructive: the systems changed behavior more reliably than they changed results. That is not a failure of the technology. It is a measure of how much else has to go right for a behavior change to reach the patient.<br><br>The design mistake, then, is not that we built decision support. It is that we built it for a clinician who does not exist &#8212; one who can switch to deliberative reasoning on demand, without cost, every time a box appears on the screen.<br></p><p>The override-as-Type-1 mechanism is my inference, not a finding. Van der Sijs notes explicitly that &#8220;studies on cognitive processes playing a role in overriding drug safety alerts are lacking.&#8221; Fit is not proof, and I will not dress an inference as a result. But the gap in the literature is itself a design failure: we built systems that assume a cognitive mechanism we never verified.<br><br>What to do instead<br>If the mode switch cannot be forced, the design task is to reduce the need for it. That means:<br><br>- Match the intervention to the mode. Ambient, workflow-embedded checks where Type 1 is sufficient. Deliberation scaffolding where Type 2 is required.<br>- Reduce the noise. Every false-positive alert teaches dismissal. The threshold for interruption should be high, and the burden of proof is on the alert, not the clinician.<br>- Measure what matters. Override rate is not enough. Track volume, override rate, and appropriateness of the overridden alert. A high override rate can mark a correctly-firing alert that catches rare but serious problems.<br>- Credit the fast system. Expert intuition is not the enemy. It is the resource that good design protects and augments, not overrides.<br><br>The limits of this frame</p><p><br>This analysis is orderer-centric. It says little about the nurse at the bedside, the pharmacist verifying the med, or the patient navigating their own care. The burden of alert fatigue falls unevenly across the care team and across settings with different resources. And the medico-legal weight of &#8220;we told you so&#8221; makes pruning alerts harder than it should be, even when the evidence supports it.<br><br>These limits are not excuses for inaction. They are the boundary conditions that Part 2 will step across.<br><br>A concrete close</p><p><br>If you manage a health system, run one audit. Rank your active alerts by volume. Pull the override rate for the top ten. Then sample a subset of overridden alerts and ask: was the override appropriate? The pattern will tell you whether your decision support is protecting patients or training clinicians to look away.<br><br>The mode switch is not free. The 300-millisecond figure is a myth. The design question is not how to interrupt, but how to match the tool to the mind &#8212; and how to know when you&#8217;ve succeeded.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[When the AI Alone Beat the Clinical Team]]></title><description><![CDATA[Algorithms and AI Come to Medicine: A Ten-Year Reassessment]]></description><link>https://bmiddleton1.substack.com/p/when-the-ai-alone-beat-the-clinical</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/when-the-ai-alone-beat-the-clinical</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Mon, 20 Jul 2026 21:50:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qUs8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!qUs8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!qUs8!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!qUs8!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!qUs8!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qUs8!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!qUs8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png" width="1536" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1024,&quot;width&quot;:1536,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!qUs8!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!qUs8!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!qUs8!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qUs8!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae65e8df-e949-41bc-98bf-e88947f1a7bc_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>A decade after I sketched a role for AI in the diagnostic process, the trial data are in &#8212; and they complicate the picture I drew in 2016.</p><p>Ten years ago I published a piece arguing that AI would take its place in the clinical armamentarium, and I offered a model for where it would sit. I borrowed a frame from M. Scott Blois, who in 1980 described a &#8220;cognitive funnel&#8221;: the human reasoner works the broad, uncertain opening of a differential diagnosis, and the machine handles the narrow, well-specified close &#8212; the calculations, the nuance-differentiation, the complex therapeutic plan. I updated that funnel with a third position. AI at the wide end, generating possibilities. AI at the narrow end, refining them. And the clinician, ideally with the patient, governing the middle &#8212; interpreting, validating, deciding.</p><p>It was a reasonable model in 2016, when the evidence to test it did not yet exist.</p><p>It exists now. And the finding that should unsettle every health system leader currently procuring an AI diagnostic copilot is this: in the one randomized trial built to test exactly this architecture, the AI running alone beat the AI-plus-physician team.</p><p>The Trial That Complicates the Comfortable Story</p><p>Fifty physicians &#8212; residents and attendings, family medicine, internal medicine, emergency medicine &#8212; worked through diagnostic vignettes under time pressure. Half had access to a large language model alongside their usual resources. Half had their usual resources alone. A blinded panel scored their differentials, their supporting reasoning, their next steps.</p><p>Physicians with the LLM scored no better than physicians without it. A two-point difference on a hundred-point rubric, not statistically significant. The LLM used by itself, with no physician in the loop at all, scored sixteen points higher than either physician group.</p><p>I want to sit with that for a moment rather than rush past it, because it is not the finding the &#8220;AI as copilot&#8221; marketing has prepared any of us for. The comfortable story is that a good tool plus a good clinician beats either alone. That is not what happened here. The team did not beat the machine. The machine beat the team.</p><p>Here is where I want to be careful. This was one trial, fifty physicians, six vignettes apiece, a controlled setting rather than a real clinic. It measured a reasoning rubric, not patient outcomes. It should not be over-read as &#8220;replace the physician&#8221; &#8212; that is a different, much larger claim this trial was not built to test, and I do not think the evidence supports it. What it does support, cleanly, is a narrower and more useful claim: pairing a capable AI with a human reasoner does not automatically produce a system better than the AI running alone. That is worth taking seriously on its own terms, without inflating it into something it isn&#8217;t.</p><p>Why the Stakes Are Not Abstract</p><p>It is worth pausing on why any of this matters beyond the seminar room. An estimated 795,000 Americans are permanently disabled or die every year because a dangerous disease was misdiagnosed &#8212; a figure built from national incidence data across strokes, infections, and cancers, not a scare number pulled from a press release. Fifteen conditions account for roughly half of that harm, which is actually the more hopeful part of the finding: the problem may be more tractable than it looks, concentrated rather than diffuse.</p><p>Any tool that plausibly narrows that gap deserves serious, well-designed evaluation. Any tool marketed as narrowing it without that evaluation deserves exactly the skepticism a new drug or device would get before it reached a patient. We would not accept &#8220;our internal testing looked promising&#8221; as a basis for prescribing a therapeutic. I am not sure why we should accept a lower bar for a diagnostic algorithm.</p><p>The Broad End of the Funnel Is Getting Genuinely Strong</p><p>To be fair to the technology &#8212; and adjudicating fairly means crediting the strong results, not just the uncomfortable ones &#8212; the broad end of Blois&#8217;s funnel has advanced further than I expected in 2016.</p><p>Given complex cases from the New England Journal of Medicine&#8216;s clinicopathological conference series &#8212; cases selected because they were hard enough to publish (and similar to those I used evaluating QMR-DT) &#8212; one generative model included the correct diagnosis in its differential 64% of the time. And Google&#8217;s AMIE, a model trained specifically for diagnostic dialogue, outperformed primary care physicians on 30 of 32 axes rated by specialist physicians, and 25 of 26 axes rated by patient-actors, in a randomized, blinded, text-based comparison that included diagnostic accuracy itself.</p><p>That second result deserves its own careful caveat. The encounter was simulated, text-only, structured like a clinical exam rather than a real visit &#8212; and the study&#8217;s own authors say plainly that translation to actual practice remains unproven. I am inclined to believe them. But I am also not inclined to dismiss a result that strong just because it is inconvenient for the &#8220;AI is a helpful assistant, nothing more&#8221; narrative that most health systems have settled on for external communications.</p><p>Ten Years of Regulatory Catch-Up, in Fairness</p><p>The 2016 piece asked, somewhat plaintively, how we would validate and monitor these tools as they took on more independent roles. That question had no good answer in 2016. It has a real one now.</p><p>In December 2024, FDA finalized guidance allowing manufacturers to file a Predetermined Change Control Plan &#8212; a pre-authorized protocol specifying how an AI-enabled device is allowed to change after clearance, without triggering a fresh marketing submission every time the model updates. That is a genuine structural answer to the &#8220;frozen model versus continuously learning model&#8221; problem I could only gesture at a decade ago. FDA&#8217;s cumulative list of AI-enabled devices now runs past 1,400, most cleared through the 510(k) pathway and concentrated in radiology &#8212; real infrastructure, though built almost entirely for narrow, single-task software, not for a general-purpose conversational model like AMIE.</p><p>ONC&#8217;s HTI-1 rule, with compliance required since January 2025, now forces developers to disclose how a predictive model was built, on what population, and with what performance characteristics, for any decision-support tool embedded in certified health IT. The EU AI Act classifies AI in a regulated medical device as high-risk by default. None of this existed when I wrote the original piece. Credit where it&#8217;s due: the field did not sit still.</p><p>The Bias We Finally Have a Name For</p><p>The 2016 piece also asked, in one throwaway sentence, how we would assess the cognitive biases in a clinician who leans too hard on an AI recommendation. That sentence now has a body of literature behind it, and the picture is more specific than &#8220;people over-trust computers.&#8221;</p><p>A large dermatology study found the effect was not uniform. Good AI support helped the least experienced clinicians the most. Faulty AI support degraded the judgment of experts and novices alike &#8212; meaning bad AI does not just fail to help; it actively drags down clinicians who would have done fine without it. The WHO&#8217;s 2024 guidance on large multi-modal models in health now lists automation bias, alongside data bias and outright inaccuracy, as a principal governance risk &#8212; not a training problem to be solved with a better tutorial, but something that has to be designed for.</p><p>Put that finding next to the diagnostic-reasoning trial above, and a working hypothesis for where we actually stand emerges: the models are getting better at the task. Humans paired with the models are not automatically getting better at the task alongside them. That gap &#8212; not raw model capability, which keeps improving on its own trajectory &#8212; is the one I think deserves the field&#8217;s attention over the next several years.</p><p>Where This Leaves Health System Leaders</p><p>I do not think the answer is to strip the clinician out of the loop. The Obermeyer case from 2019 is the standing rebuttal to that instinct: a widely used algorithm ranked Black patients as healthier than equally sick White patients, not through any explicit bias, but because it used healthcare cost as a proxy for illness and unequal historical spending baked the disparity in. It passed every conventional accuracy metric. A human paying attention to the right question &#8212; a proxy for what, exactly? &#8212; is still how that kind of failure gets caught, and no amount of model capability substitutes for someone asking it.</p><p>But I also do not think &#8220;keep a human in the loop&#8221; is, by itself, a design principle anymore. It is closer to a hope. The evidence above says the loop itself has to be built deliberately &#8212; tested, in a trial, against the AI running alone &#8212; rather than assumed to be additive because it feels more careful.</p><p>So here is the question I would put to any health system leader currently evaluating a diagnostic AI tool, mine included: have you tested the combination against the AI alone, on your own population, with your own clinicians? Or have you assumed, the way I assumed in 2016, that putting a person in the middle of the funnel automatically makes the funnel better?</p><p>On the evidence gathered so far, that assumption does not hold for free. It has to be earned, case by case, and measured rather than presumed.</p>]]></content:encoded></item><item><title><![CDATA[Three AI Mega-IPOs. One Question Nobody Is Asking]]></title><description><![CDATA[An honest accounting of the AI dividend idea &#8212; what the arithmetic kills, what survives, and what today&#8217;s SpaceX IPO teaches about both.]]></description><link>https://bmiddleton1.substack.com/p/three-ai-mega-ipos-one-question-nobody</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/three-ai-mega-ipos-one-question-nobody</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Fri, 12 Jun 2026 16:26:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LxCu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LxCu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LxCu!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png 424w, /__u/substackcdn.com/image/fetch/$s_!LxCu!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png 848w, /__u/substackcdn.com/image/fetch/$s_!LxCu!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LxCu!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LxCu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png" width="1456" height="816" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:816,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!LxCu!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png 424w, /__u/substackcdn.com/image/fetch/$s_!LxCu!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png 848w, /__u/substackcdn.com/image/fetch/$s_!LxCu!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LxCu!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c1f251c-9db1-4987-b255-e4bd9f59f2fe_1456x816.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>&#8220;Tax AI and fund universal basic income and universal healthcare.&#8221;</p><p>You have heard some version of this claim. It circulates whenever a frontier lab announces another multibillion-dollar compute cluster or another industry posts another round of layoffs attributed, fairly or not, to automation. The claim has obvious appeal. It also fails on first contact with arithmetic.</p><p>But buried inside the wreckage of the strong claim is a bounded claim worth defending &#8212; one with real consequences for how we finance healthcare and social insurance over the next two decades. The interesting question is not whether AI taxation can fund a post-work utopia. It cannot, at least not on any evidence available today. The interesting question is whether AI will generate concentrated economic rents built partly on public assets, and if so, whether the public should hold a claim on those rents before market structures harden.</p><p>I think the answer to both halves of that question is a qualified yes. Here is the case, including the parts that should make us uncomfortable.</p><p><strong>The Arithmetic That Kills the Strong Claim</strong></p><p>Start with the denominators.</p><p>U.S. national health expenditures reached $5.3 trillion in 2024 &#8212; $15,474 per person, 18.0 percent of GDP, per the CMS Office of the Actuary. That figure grew 7.2 percent year over year, outpacing the economy as it has for most of five decades.</p><p>Universal basic income carries a similar price tag. Hoynes and Rothstein, in the <em>Annual Review of Economics</em>, estimated that a $12,000 annual payment to every U.S. adult would cost roughly $3 trillion per year &#8212; about three-quarters of total federal expenditures at the time of their analysis. Even cannibalizing every existing transfer program, including Social Security, Medicare, and Medicaid, would not cover it.</p><p>So the combined bill for &#8220;full UBI plus universal healthcare&#8221; runs to roughly $8 trillion annually before offsets. Against that, what does AI offer?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LJgQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LJgQ!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png 424w, /__u/substackcdn.com/image/fetch/$s_!LJgQ!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png 848w, /__u/substackcdn.com/image/fetch/$s_!LJgQ!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LJgQ!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LJgQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png" width="1456" height="816" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:816,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!LJgQ!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png 424w, /__u/substackcdn.com/image/fetch/$s_!LJgQ!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png 848w, /__u/substackcdn.com/image/fetch/$s_!LJgQ!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LJgQ!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a762f16-ee83-40c7-b242-0491fc49365d_1456x816.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Sources: CMS Office of the Actuary (2024 NHE); Hoynes &amp; Rothstein, Annu Rev Econ 2019; Stanford AI Index 2026</em></p><p>Investment, mostly. U.S. private AI investment ran to $109 billion in 2024 and $286 billion in 2025 by the Stanford AI Index&#8217;s count &#8212; staggering sums, and still beside the point. Investment is not profit. Valuation is not cash flow. Capital expenditure is not a taxable base. Most frontier labs remain in a capital-intensive buildout phase, and their valuations price in expectations of future rents, not present revenue capacity that a treasury could tap.</p><p>No credible current evidence shows that realized AI profits could finance even a meaningful fraction of the strong claim. Anyone selling you that version is selling fiscal fantasy.</p><p>That should end the conversation about the strong claim. It should not end the conversation.</p><p><strong>What Survives: Rents Are Not Profits</strong></p><p>The defensible version of the argument rests on a distinction that public-finance economists have understood since Ricardo: the difference between profit and rent.</p><p>Normal profit is the return a firm needs to cover its costs &#8212; talent, capital, infrastructure &#8212; and stay in business. Economic rent is the surplus above that threshold, earned not from the cost of creation but from structural advantage: monopoly position, control of scarce inputs, intellectual-property moats, network effects. Because rent is by definition excess, taxing a portion of it does not, in theory, discourage the underlying productive activity. The activity remains profitable after the tax. This is why economists generally regard well-designed rent taxes as among the least distortionary instruments available &#8212; easy to say in theory, hard to administer in practice, a point I will return to.</p><p>Frontier AI has the structural preconditions for rent formation. The inputs &#8212; advanced chips, hyperscale compute, cloud platforms, proprietary data, specialized talent, distribution channels &#8212; concentrate among a handful of firms. The Stanford AI Index reports that roughly 90 percent of notable AI models in 2024 came from industry, not academia. Whether these advantages produce <em>durable</em> rents or get competed away by open-source models and falling inference costs remains genuinely open. But the preconditions exist, and history suggests that waiting until rents are fully entrenched makes public recapture far harder. We did not negotiate spectrum auctions after broadcasters owned the airwaves by custom; we should not wait until AI market structure ossifies to decide whether the public holds a claim.</p><p><strong>The Public-Good Rationale &#8212; and Its Limits</strong></p><p>Why would the public hold a claim at all? Because the knowledge substrate of modern AI is substantially collective.</p><p>NSF funded foundational AI research beginning in the early 1960s. ARPANET and NSFNET built the network that became the commercial internet. Public universities trained the technical workforce. Open-source software, open science, and the accumulated language, code, and scholarship of millions of people populate every training corpus. Private firms contributed enormously &#8212; capital, engineering, risk, commercialization &#8212; and nothing in this argument treats their returns as illegitimate. The accurate framing is that AI profits are joint products of private entrepreneurship and public knowledge capital. A legitimate policy does not confiscate the private return. It recaptures part of the public one.</p><p>We already do this elsewhere. Society captures public value from natural resources, spectrum rights, and land appreciation created by public infrastructure. AI differs in technical form, not in moral structure.</p><p><strong>Here is where I want to be careful.</strong> The moral premise is broader than any administrable tax base, and that gap is where versions of this proposal tend to die.</p><p>&#8220;The public contributed to AI&#8221; describes at least four different things, and they do not support the same policy instruments:</p><ol><li><p><strong>Direct public support</strong> &#8212; grants, federal contracts, public compute, loan guarantees. This creates the cleanest case: when firms take public money or public privileges, the public should take equity warrants or revenue-sharing rights in return. Identifiable benefit, identifiable beneficiary, administrable mechanism.</p></li><li><p><strong>Public infrastructure</strong> &#8212; the internet itself. A real but diffuse public-good contribution. It justifies general rent taxation more than any targeted levy.</p></li><li><p><strong>Open scientific knowledge</strong> &#8212; publicly funded research in training corpora. This may support licensing obligations in defined cases, but broad &#8220;knowledge commons royalties&#8221; would collide with open-science norms and invite legal challenge.</p></li><li><p><strong>Collective cultural production</strong> &#8212; the writing, art, and code of millions who never consented to becoming training substrate. This raises genuine consent and compensation questions, but they belong to copyright law, creator bargaining, and transparency mandates &#8212; not to a tax code.</p></li></ol><p></p><p>Collapsing these four categories into one undifferentiated grievance produces a proposal that is morally expansive and administratively incoherent. Keeping them separate produces something a finance ministry could actually run.</p><p>A second caution: defining &#8220;AI-derived rent&#8221; inside a diversified firm is the central administrative problem, and no one has solved it. A firm may deploy AI simultaneously in logistics, advertising, drug discovery, and customer service. Self-reported &#8220;AI profit&#8221; would be gameable on day one. Any workable test must rest on observable criteria &#8212; firm scale, market concentration, compute intensity, returns above a normal capital allowance, public-resource dependency &#8212; and even then, transfer pricing, IP migration, and compute relocation will erode the base without international coordination.</p><p>A third caution: the rents may not last. If open-weight models and collapsing inference costs commoditize the frontier, concentrated AI rents could dissipate &#8212; as railroad and telegraph rents eventually did, despite confident predictions of their permanence. A well-designed mechanism should shrink automatically as rents shrink, rather than survive as a zombie tax on returns that are no longer excess.</p><p><strong>What a Bounded Design Looks Like</strong></p><p>Take those cautions seriously and a layered architecture emerges, ordered from most to least administrable:</p><ul><li><p><strong>Public equity warrants first.</strong> When firms receive major public benefits &#8212; contracts, grants, public compute, liability shields, exclusive commercialization rights over publicly funded research &#8212; the public receives warrants or revenue-sharing rights. This converts public support into public upside through mechanisms that procurement law already understands.</p></li><li><p><strong>A high-threshold excess-rent tax second.</strong> A surtax on above-normal returns, triggered only by measurable criteria of scale, concentration, and public-resource dependency. Ordinary AI adoption stays untouched. A rural clinic using AI to cut administrative burden bears no resemblance to a frontier platform earning monopoly rents from global deployment, and the tax code should know the difference.</p></li><li><p><strong>A narrow compute levy third, if at all.</strong> Compute and energy use are observable in ways that &#8220;AI profit&#8221; is not, which makes them tempting. But they proxy for scale, not rent &#8212; a levy without high thresholds and exemptions would tax loss-making research, open science, and safety testing. Frontier-infrastructure levy, not general AI tax.</p></li><li><p><strong>A revenue waterfall, not a promise.</strong> Proceeds flow in fixed order: measurement and administrative capacity; worker transition and wage insurance; public-interest AI infrastructure; health-security support; affected communities; and only then &#8212; once revenues prove durable and recurring &#8212; a per-capita dividend. The dividend scales with realized revenue, never with forecasts.</p></li></ul><p></p><p>That last principle does most of the work. It converts the proposal from a blank-check UBI promise into something closer to Alaska&#8217;s Permanent Fund logic: the public&#8217;s share of a real revenue stream, whatever size that stream turns out to be.</p><p><strong>Why Healthcare Readers Should Care</strong></p><p>For this audience, the stakes are concrete. American health insurance remains substantially employment-attached. If AI meaningfully erodes wage income in exposed sectors &#8212; and I stress <em>if</em>, because labor-market exposure does not mechanically equal displacement, and the augmentation scenario remains live &#8212; then the financing base for employment-linked coverage erodes with it. Payroll-attached social insurance and a labor-displacing general-purpose technology make poor companions.</p><p>The honest claim is not that AI rent capture &#8220;funds universal healthcare.&#8221; Against a $5.3 trillion denominator, it cannot. The honest claim is that AI rent capture could help decouple health security from employment status at the margin &#8212; subsidies, solvency support, public-option financing &#8212; while the deeper payment-reform work proceeds on its own track. One pillar of the fiscal architecture. Not the architecture.</p><p><strong>Today&#8217;s Natural Experiment: SpaceX</strong></p><p>As it happens, the markets staged a live demonstration of this entire argument today.</p><p>SpaceX began trading on the Nasdaq this morning under the ticker SPCX, raising $75 billion &#8212; the largest IPO in history, blowing past Saudi Aramco&#8217;s 2019 record &#8212; at a valuation near $1.77 trillion that instantly places it among the largest American companies. Run that listing through the framework above and watch each test light up.</p><ul><li><p><strong>The joint-product test.</strong> SpaceX represents private entrepreneurship of the highest order: reusable rockets, an order-of-magnitude reduction in launch costs, genuine financial risk including a near-death experience in 2008. Nothing here diminishes that. But the company also grew on public anchor demand &#8212; more than $20 billion in government contracts over its history, mostly NASA and the Department of Defense, including the December 2008 cargo resupply award that arrived when the company was weeks from insolvency, and a Commercial Crew contract NASA values at roughly $4.9 billion. Starlink, meanwhile, monetizes FCC-licensed spectrum and low-Earth-orbit slots &#8212; exactly the category of public commons that the spectrum-auction precedent already recognizes. Joint product of private genius and public assets: textbook.</p></li><li><p><strong>The rent test.</strong> SpaceX dominates U.S. orbital launch. Dragon remains the only operational American crew vehicle. Starlink leads LEO broadband by a wide margin. Bottleneck control, scale economies, vertical integration, network effects &#8212; the full structural signature of durable rent.</p></li><li><p><strong>The foregone-upside test.</strong> Here is the uncomfortable part. The public bought services from SpaceX through fixed-price contracts, and honesty requires saying those were good deals *as procurement* &#8212; NASA paid far less than cost-plus development would have demanded, and the taxpayer savings were real. But the deal structure contained no equity participation. A warrant for even one percent of the company, attached to those survival-stage development contracts, would be worth roughly $18 billion at today&#8217;s valuation &#8212; comparable to the company&#8217;s entire two-decade haul of government contract awards &#8212; and even after dilution across years of subsequent funding rounds, still a multibillion-dollar public position. The order of magnitude, not the decimal, is the point. The public supplied demand when it mattered most and holds no claim on the $1.77 trillion it helped underwrite.</p></li></ul><p>That is not theft. It is upside left on the table &#8212; and here the story turns, because the federal government has recently learned to ask. In August 2025, the Commerce Department converted $8.9 billion in unpaid CHIPS Act and Secure Enclave grants into a 9.9 percent equity stake in Intel &#8212; 433.3 million shares at $20.47, per Intel&#8217;s own SEC filings &#8212; and took a five-year warrant for an additional five percent, exercisable if Intel loses majority control of its foundry business. The Pentagon took a preferred-equity position in MP Materials weeks earlier; the U.S. Steel transaction came with a golden share. Reasonable people dispute the design of these deals &#8212; critics note the Intel arrangement traded away the original grant&#8217;s profit-sharing and clawback protections, and ad hoc equity-taking invites favoritism that a rules-based framework would not. But the threshold question is settled by demonstration: the United States government can take, and now does take, equity positions in strategic firms it funds. The instrument exists and the precedent is bipartisan in its discomfort. What does not yet exist is the consistent policy &#8212; applied before the next $1.77 trillion listing rather than improvised after.</p><ul><li><p><strong>The AI bridge.</strong> SpaceX absorbed xAI earlier this year, which makes the largest IPO in history partly an AI listing &#8212; and OpenAI and Anthropic have both filed paperwork for their own offerings, likely later this year. Three AI-linked mega-listings in a single window. Warrants attach before a company goes public, not after &#8212; Intel proves the conversion is hardest and messiest when attempted late. The SpaceX counterfactual is no longer hypothetical; it trades under a ticker now.</p></li></ul><p><strong>The Principle, Stated Narrowly</strong></p><p>When firms earn durable, above-normal returns from AI systems that depend materially on public research, public infrastructure, or legally privileged access to shared knowledge assets, a portion of those excess returns should finance public benefits.</p><p>Note everything that principle does not claim. It does not propose taxing AI use. It does not assert that current AI profits can fund universal programs. It does not assume exposed jobs disappear. It does not treat valuation as taxable profit. It claims only that the public, having seeded the field, should hold a harvest right that activates if and when the harvest comes in &#8212; and that the time to establish that right is before the ownership patterns harden, not after.</p><p>The responsible position lies between technological fatalism and fiscal fantasy. The fatalists tell us concentrated AI wealth is inevitable and untouchable. The fantasists tell us a modest tax buys a post-work dividend state. Both are wrong, and both make the bounded, boring, administrable middle path harder to see.</p><p>Today the SpaceX bell rang on twenty years of public investment with no public equity attached. The Intel deal proves the instrument exists; two more prospectuses are already at the SEC. So here is the question I would put to readers, particularly those of you running health systems and watching AI procurement budgets grow: what should a public warrant on those next two listings look like &#8212; and if your answer is &#8220;nothing,&#8221; what evidence would change your mind? The SpaceX counterfactual ran to billions, perhaps tens of billions. We will not get to rerun it. We will get exactly one chance to avoid repeating it.</p><p>#HealthIT #ClinicalAI #HealthPolicy #AIPolicy #DigitalHealth #ArtificialIntelligence #SpaceX #SPCX #IPO #TechPolicy #EconomicPolicy #CMIO #HealthInformatics #CHIPSAct #PublicFinance #AIGovernance&nbsp;</p><p></p><p><strong>Sources</strong></p><p>Centers for Medicare &amp; Medicaid Services, Office of the Actuary. National Health Expenditure Data, 2024 (released via Health Affairs): $5.3 trillion; $15,474 per capita; 18.0% of GDP.</p><p>Hoynes H, Rothstein J. Universal Basic Income in the United States and Advanced Countries. Annual Review of Economics. 2019;11:929&#8211;958. doi:10.1146/annurev-economics-080218-030237.</p><p>Stanford Institute for Human-Centered AI. AI Index Report 2025 (industry share of notable models; 2024 U.S. private AI investment) and AI Index Report 2026 (2025 U.S. private AI investment).</p><p>U.S. National Science Foundation, on AI research funding since the early 1960s and on NSFNET and the birth of the commercial internet.</p><p>SpaceX IPO terms and coverage, June 11&#8211;12, 2026: NPR, CNBC, and Kiplinger (555,555,555 shares at $135; $75 billion raised; ~$1.77 trillion valuation; xAI acquisition; OpenAI and Anthropic SEC filings).</p><p>SpaceX government contract history: USAspending-based analyses (&gt;$20 billion cumulative); NASA Commercial Crew Transportation Capability contract (~$4.93 billion, per NASA, 2022).</p><p>Intel Corporation, Form 8-K (August 22 and August 27, 2025), SEC EDGAR: $8.9 billion U.S. government purchase of 433.3 million shares at $20.47 (9.9% stake), funded by unpaid CHIPS Act and Secure Enclave grants; five-year warrant for additional 5%; closing August 27, 2025. See also Intel Newsroom.</p><p>U.S. Senate Committee on Banking, Housing, and Urban Affairs, letter to the Secretary of Commerce (September 3, 2025), on removal of profit-sharing and clawback provisions from the original CHIPS agreement.</p><p>Department of Defense preferred-equity investment in MP Materials (July 2025) and golden-share arrangement in the Nippon Steel&#8211;U.S. Steel transaction (June 2025), as widely reported.</p><p>The one-percent warrant figure is the author&#8217;s arithmetic against the IPO valuation, offered as an illustration, not a valuation.</p><p></p>]]></content:encoded></item><item><title><![CDATA[The Diagnosis Problem: GAI Gets the Final Answer Right — and the Reasoning Wrong]]></title><description><![CDATA[Why the most cognitively demanding act in medicine is the one AI handles worst, and what we should do about it]]></description><link>https://bmiddleton1.substack.com/p/the-diagnosis-problem-gai-gets-the</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-diagnosis-problem-gai-gets-the</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Thu, 04 Jun 2026 17:24:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K7H0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!K7H0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!K7H0!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!K7H0!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!K7H0!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!K7H0!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!K7H0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1943816,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://bmiddleton1.substack.com/i/200642844?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!K7H0!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!K7H0!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!K7H0!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!K7H0!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58ebe6e1-e0e4-4d98-9a52-339b9bd7d614_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Every diagnosis begins with uncertainty. The best clinicians I&#8217;ve known stay longest in that uncertainty &#8212; resisting the pull toward closure until the evidence earns it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The differential diagnosis is not a guess. It is disciplined probabilistic reasoning under incomplete information with asymmetric stakes. Miss a stroke in the emergency department and the tPA window closes. Miss sepsis at hour one and mortality roughly doubles with each subsequent hour of delay. Miss a spinal abscess &#8212; misdiagnosed in approximately 62% of cases &#8212; and the patient may walk out of your office and return in a wheelchair.</p><p>The scale of the problem we are trying to solve is not hypothetical. Newman-Toker and colleagues at Hopkins estimated that 795,000 Americans die or are permanently disabled from diagnostic errors annually. Their &#8220;Big Five&#8221; &#8212; stroke, sepsis, pneumonia, venous thromboembolism, and lung cancer &#8212; account for 38.7% of all serious diagnostic harms. These are not zebras. They are common presentations that anchored to the wrong diagnosis and never got reconsidered.</p><p>Newman-Toker has said this plainly: diagnostic errors are &#8220;by a wide margin, the most under resourced public health crisis we face, yet research funding only recently reached the $20 million per year mark.&#8221;</p><p>Now generative AI enters this picture claiming diagnostic capability. Let&#8217;s look at what the evidence actually shows &#8212; because there is a paradox in it that should give every clinician pause.</p><h2><strong>The Benchmark Illusion</strong></h2><p>Generative AI models now routinely pass medical licensing examinations at physician-level or above. When researchers hand them a complete clinical case with all pertinent information, they arrive at the correct final diagnosis more than 90% of the time.</p><p>That is a number worth pausing on.</p><p>But a Mass General Brigham study led by Arya Rao and published in JAMA Network Open this April &#8212; testing 21 different large language models across 29 standardized clinical cases &#8212; found something that should recalibrate that enthusiasm immediately: those same models failed to produce an appropriate differential diagnosis more than 80% of the time.</p><p>Read that again. Greater than 90% accuracy on final diagnosis. Less than 20% appropriate performance on differential diagnosis.</p><p>Students of the expert systems era will recognize this structural gap. Berner and colleagues&#8217; 1994 evaluation of four computer-based diagnostic programs in the New England Journal of Medicine found final-diagnosis accuracy of 52&#8211;71% alongside relevant differential coverage of only 19&#8211;37% &#8212; less than half of expert-generated lists, though each program did surface roughly two additional relevant diagnoses per case that experts had not originally considered. Thirty years on, the architecture has changed. The structural gap has not.</p><p>The MGB team developed a new evaluation framework for this study called PrIME-LLM, which assesses models across the sequential stages of clinical reasoning &#8212; generating potential diagnoses, conducting appropriate workups, arriving at a final diagnosis, and managing treatment. The scores ranged from 64% for Gemini 1.5 Flash to 78% for Grok 4 and GPT-5. What the framework exposed, as lead author Rao put it, is that &#8220;these models are great at naming a final diagnosis once the data is complete, but they struggle at the open-ended start of a case, when there isn&#8217;t much information.&#8221;</p><p>Corresponding author Marc Succi was direct: &#8220;Differential diagnoses are central to clinical reasoning and underlie the &#8216;art of medicine&#8217; that AI cannot currently replicate.&#8221;</p><p>This is not a performance gap at the margins. It is a structural failure at exactly the step where diagnostic harm most often originates.</p><h2><strong>Where the Reasoning Fails</strong></h2><p>Three failure patterns surface repeatedly across the literature, and I recognize all three from thirty years watching clinical decision support systems hit the same walls.</p><p>Blois identified the structural explanation for this failure forty-five years ago. In a 1980 NEJM paper, he described clinical reasoning as a cognitive funnel &#8212; wide and judgment-dependent early in an encounter, when the physician must navigate a vast undifferentiated space of diagnostic possibility, narrowing progressively toward more algorithmic territory as data accumulates and hypotheses tighten. The question worth asking, he argued, was not where computers could be used but where human beings must be. All three failure patterns below live at the wide end of that funnel &#8212; exactly where LLMs break down.</p><p><strong>Anchoring without recalibration.</strong> The MGB study presented models with clinical information sequentially, simulating how a case actually unfolds. The models struggled to revise their early hypotheses as new information arrived. A seasoned internist holds the differential open; the LLMs close it prematurely. This is computational anchoring &#8212; the model locks onto an early high-probability match and fails to weight the disconfirming evidence that follows.</p><p><strong>Prior probability neglect.</strong> Good DDx reasoning is Bayesian. A 28-year-old and a 65-year-old presenting with chest pain have radically different prior probabilities for every diagnosis on their differentials. The evidence suggests that current LLMs do not consistently integrate population-level priors into case-level reasoning. They generate plausible differentials in the abstract, but not necessarily plausible differentials for this patient in this context.</p><p><strong>Fluent confidence in wrong answers.</strong> Automation bias in CDSS &#8212; the tendency of clinicians to defer to system-generated recommendations over independent judgment &#8212; is well-established; Goddard, Roudsari, and Wyatt documented its frequency, mediators, and mitigators in a 2012 systematic review in JAMIA. The Umerenkov preprint extends the pattern to LLM diagnostic explanations directly. Fluent, confident explanations drove physician agreement rates up even when the underlying diagnoses were wrong &#8212; error rates in those explanations ranged from 5% to 30%. That fluency is not a feature in clinical practice. It is a safety hazard.</p><p>These are not bugs to be patched in the next model release. They reflect a fundamental architecture mismatch between how LLMs generate statistically plausible text and how diagnostic reasoning actually works.</p><h2><strong>The Evidence in Aggregate</strong></h2><p>Where does that leave us overall? A 2025 systematic review and meta-analysis by Takita and colleagues in npj Digital Medicine &#8212; the first of its kind &#8212; synthesized 83 studies of generative AI diagnostic performance. Their pooled accuracy across all models was 52.1%. Generative AI showed no statistically significant difference in performance compared to physicians overall (p=0.10) or to non-expert physicians (p=0.93), but was significantly inferior to expert physicians overall (p=0.007, mean accuracy difference 15.8%).</p><p>Two nuances matter. The performance gap was driven largely by older model generations &#8212; several newer models, including GPT-4o, Claude 3 Opus, and Gemini 1.5 Pro, showed no statistically significant difference from expert physicians in individual comparisons. And reviewers rated 76% of included studies at high risk of bias, primarily because AI training data is undisclosed and test set independence cannot be confirmed. All studies used curated vignettes, not the messier conditions of actual practice.</p><p>One more data point that fits this picture: Feldman and colleagues at Massachusetts General found that a dedicated AI diagnostic expert system outperformed large language models on accuracy &#8212; but with complementary error profiles. Each caught cases the other missed. That finding does not establish that generative AI fails at diagnosis. It reinforces the case for hybrid architectures over single-model deployment, a theme the evidence returns to below.</p><p>So a blunt summary: generative AI has improved rapidly, some frontier models now approach expert-level performance on structured diagnostic tasks &#8212; but we are testing against benchmarks while the problem we care about is 795,000 serious harms annually in real environments where patients present incompletely, documentation is inconsistent, and the clock is running.</p><h2><strong>The Clinical Case for Proceeding Anyway</strong></h2><p>Here is where I want to be careful not to let the evidence of imperfection become grounds for indefinite paralysis &#8212; which is the healthcare system&#8217;s default response to any technology that doesn&#8217;t arrive fully validated.</p><p>The diagnostic error problem is too large for that posture. We are not comparing GAI-assisted diagnosis to perfect clinical reasoning. We are comparing it to a system in which &#8212; in a 2024 study of seriously ill adults across 29 academic medical centers &#8212; 23% of patients who were transferred to the ICU or died in the hospital had a missed or delayed diagnosis that contributed to harm. Three-quarters of those errors caused temporary or permanent injury, and diagnostic error was a contributing factor in roughly one in fifteen of the deaths in the study cohort.</p><p>Auerbach and colleagues made something else clear: the most common errors were delayed diagnoses &#8212; failure to consult a specialist in time, failure to consider an alternate diagnosis soon enough, failure to order or correctly interpret the relevant test. Those are exactly the conditions where a well-designed GAI challenger, surfacing missed possibilities before the window closes, could save lives.</p><p>The architecture that shows the most promise is described in a June 2025 PNAS study by Z&#246;ller and colleagues at the Max Planck Institute for Human Development. Analyzing 40,762 differential diagnoses made by physicians alongside five leading AI models across 2,133 medical case vignettes, they found that hybrid human-AI collectives outperformed individual physicians, physician collectives, individual AI models, and AI ensembles alike. Standalone AI collectives outperformed 85% of individual human diagnosticians. But the hybrid collectives did best of all &#8212; because humans and AI make different kinds of errors, and combining them catches what either misses alone.</p><p>That complementarity is real, and it is the design target.</p><h2><strong>Functional Questions the Field Has Not Answered</strong></h2><p>Before we deploy GAI in any DDx-adjacent role, we owe clinicians and patients honest answers to the questions we have barely started asking &#8212; beginning with the one most likely to harm people we already underserve.</p><p>How does performance vary across demographic subgroups? Diagnostic accuracy for stroke is already significantly lower for women and for Black patients in human clinical practice &#8212; patterns driven by presentation differences, cognitive bias, and systemic inequities baked into historical practice. If GAI trains on that history, it may amplify those disparities. The Takita meta-analysis found that most included studies lacked demographic data sufficient to assess subgroup generalizability. That gap is not a footnote. It is a patient safety issue.</p><h2><strong>Ethical Questions That Cannot Wait</strong></h2><p>Automation bias is well-established in the CDSS literature: clinicians presented with a system-generated recommendation are more likely to accept it than to override it, even when the recommendation is wrong. Goddard, Roudsari, and Wyatt&#8217;s systematic review documented the frequency, mediators, and mitigators of this effect more than a decade ago. The Umerenkov preprint extends the finding specifically to LLM-generated diagnostic explanations &#8212; fluent, confident explanations that increase agreement rates even when the underlying diagnoses contain errors. Deploying a GAI DDx tool without designing explicit safeguards against this is not a neutral deployment decision. It is a decision to accept that harm in exchange for workflow convenience.</p><p>We have not seriously confronted what happens to clinical reasoning skill when AI offloads the cognitive work of DDx formation during training. If residents increasingly rely on GAI to build differentials in their formative years, and the GAI fails systematically in the ways this evidence describes, we are not just risking individual patient harm. We are degrading the next generation&#8217;s capacity for independent diagnostic reasoning at precisely the moment when we will need it most.</p><p>Liability remains unresolved. When a GAI-assisted differential misses a diagnosis that a physician then fails to pursue, who bears responsibility? The clinician who deferred? The institution that deployed the tool? The vendor who built it? This is not hypothetical. It will be a case before a medical liability panel within the next few years. We are unprepared.</p><p>And consent? Patients in resource-limited settings who receive GAI-assisted diagnostic care may never know that AI played a role in shaping their workup. The informed consent infrastructure for AI-augmented clinical decision-making does not yet exist.</p><h2><strong>A Path Forward</strong></h2><p>Any architecture worth building begins with a single admonition: do not close the cognitive funnel before the data warrant it. Adler-Milstein and colleagues articulated the positive vision in a 2021 JAMA essay &#8212; next-generation diagnostic AI should shift from predicting diagnostic labels to &#8220;wayfinding&#8221;: navigating the clinician through uncertainty, surfacing what to consider next, and holding the diagnostic space open as long as the clinical picture demands. The distinction is not semantic. Diagnosis and treatment planning both exist on spectrums of certainty, and the appropriate AI contribution shifts as that uncertainty resolves. A system that skips to confident label prediction before earning it does not improve clinical reasoning &#8212; it closes the funnel faster, which is precisely the failure the evidence above documents.</p><p>I have argued previously in these pages for Performance-Gated Autonomy as the governance framework for clinical AI &#8212; the idea that AI authorization should function like a revocable credential, earned through demonstrated performance and withdrawable when that performance degrades. DDx assistance is a natural first application of that framework. The failure modes are specific and measurable. The appropriate scope of AI action is genuinely gradable. And the stakes of getting it wrong are high enough to demand rigor.</p><p>Tier the architecture explicitly. In an initial shadow mode, the AI generates a differential silently, and teams track its output against clinical decisions without those comparisons influencing clinician judgment. In the second stage, it functions as a challenger &#8212; shown to the clinician only after they have independently generated their own differential, surfacing candidates that were missed without displacing the original cognitive work. Autonomous DDx generation should come only after sustained demonstrated performance in target environments against target populations, with statistical signals of drift triggering revocation.</p><p>That sequencing requires something the field conspicuously lacks: standardized criteria for DDx quality beyond benchmark accuracy on curated vignettes &#8212; case-level calibration metrics, subgroup performance tracking, and explicit thresholds for adequate differential generation before authorizing any level of clinical deployment.</p><p>Two clinical contexts justify moving forward now, even with today&#8217;s evidence gaps. First, resource-constrained settings &#8212; rural emergency departments, community hospitals, underserved global health systems &#8212; where a well-calibrated DDx tool could surface diagnoses that a single, fatigued clinician working far from specialist support might not reach. Second, rare disease diagnosis, where the long tail of unusual presentations sits outside any individual clinician&#8217;s pattern-matching range but a system with broader training may surface them.</p><p>The goal is not to automate clinical reasoning. The goal is to extend the reach of rigorous diagnostic thinking to every patient who currently lacks access to it &#8212; and to catch the cases that fall through the gaps of an already overburdened system.</p><p>That goal is worth pursuing seriously. It is not worth pursuing without standards.</p><p>What are you seeing in your institution? Are clinicians using GAI for diagnostic support now, and if so, what does responsible guardrailing look like in practice? The clinical informatics community needs to build this evidence base together &#8212; and the anecdotes from the front lines matter.</p><p><em>Dr. Blackford Middleton has spent thirty-plus years in clinical informatics, most of it trying to get good clinical knowledge to the right clinician at the right moment. He remains convinced this is possible, and that the field keeps finding creative new ways to make it harder than it has to be.</em></p><p><strong>Bibliography</strong></p><p>1.Rao AS, et al. &#8220;Large Language Model Performance and Clinical Reasoning Tasks.&#8221; JAMA Network Open. 2026; doi: 10.1001/jamanetworkopen.2026.4003. [April 13, 2026]</p><p>2.Newman-Toker DE, Nassery N, Schaffer AC, Yu-Moe CW, Clemens GD, Wang Z, et al. &#8220;Burden of serious harms from diagnostic error in the USA.&#8221; BMJ Quality &amp; Safety. 2023. doi: 10.1136/bmjqs-2021-014130</p><p>3.Auerbach AD, Lee TC, Hubbard CC, Ranji SR, Esmaili AM, Barish P, et al. &#8220;Diagnostic Errors in Seriously Ill Hospitalized Adults.&#8221; JAMA Internal Medicine. 2024; doi: 10.1001/jamainternmed.2023.7379.</p><p>4.Zoller N, Berger J, Lin I, Fu N, Komarneni J, Barabucci G, et al. &#8220;Human-AI collectives most accurately diagnose clinical vignettes.&#8221; PNAS. 2025;122(24):e2426153122. doi: 10.1073/pnas.2426153122</p><p>5.Takita H, Kabata D, Walston SL, Tatekawa H, et al. &#8220;A systematic review and meta-analysis of diagnostic performance comparison between generative AI and physicians.&#8221; npj Digital Medicine. 2025;8:175. doi: 10.1038/s41746-025-01543-z.</p><p>6.Umerenkov D, Zubkova G, Nesterov A. &#8220;Deciphering Diagnoses: How Large Language Models Explanations Influence Clinical Decision Making.&#8221; Sber AI Lab. [Preprint; not peer-reviewed]</p><p>7.Goddard K, Roudsari A, Wyatt JC. &#8220;Automation bias: a systematic review of frequency, effect mediators, and mitigators.&#8221; Journal of the American Medical Informatics Association. 2012;19(1):121&#8211;7. doi: 10.1136/amiajnl-2011-000089.</p><p>8.Feldman MJ, Hoffer EP, Conley JJ, Chang J, Chung JA, Jernigan MC, Lester WT, Strasser ZH, Chueh HC. &#8220;Dedicated AI Expert System vs Generative AI With Large Language Model for Clinical Diagnoses.&#8221; JAMA Network Open. 2025;8(5):e2512994. doi: 10.1001/jamanetworkopen.2025.12994.</p><p>9.Berner ES, Webster GD, Shugerman AA, Jackson JR, Algina J, Baker AL, et al. &#8220;Performance of four computer-based diagnostic systems.&#8221; N Engl J Med. 1994;330(25):1792&#8211;1798.</p><p>10.Blois MS. &#8220;Clinical judgment and computers.&#8221; N Engl J Med. 1980;303(4):192&#8211;197.</p><p>11.Adler-Milstein J, Chen JH, Dhaliwal G. &#8220;Next-Generation Artificial Intelligence for Diagnosis: From Predicting Diagnostic Labels to &#8220;Wayfinding.&#8221;&#8221; JAMA. 2021;326(24):2467&#8211;2468.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Needs a Residency: The Case for Performance-Gated Autonomy]]></title><description><![CDATA[Medicine spent more than a century perfecting how we train, evaluate, and credential physicians. Let's do the same for AI in clinical practice.]]></description><link>https://bmiddleton1.substack.com/p/ai-needs-a-residency-the-case-for</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/ai-needs-a-residency-the-case-for</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Tue, 28 Apr 2026 17:59:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KMFe!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc54ae8cf-4907-48fa-ad2c-62bd601a1ceb_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!rv9f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!rv9f!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png 424w, /__u/substackcdn.com/image/fetch/$s_!rv9f!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png 848w, /__u/substackcdn.com/image/fetch/$s_!rv9f!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rv9f!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!rv9f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png" width="758" height="248" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:248,&quot;width&quot;:758,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:397542,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://bmiddleton1.substack.com/i/195777108?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!rv9f!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png 424w, /__u/substackcdn.com/image/fetch/$s_!rv9f!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png 848w, /__u/substackcdn.com/image/fetch/$s_!rv9f!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rv9f!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fdb0b8-fc64-41f5-91c9-3e9ebffe5af2_758x248.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>We do not hand a scalpel to a first-year medical student and say good luck. We require graded responsibility &#8212; shadowing, then supervised practice, then independent action within a defined scope &#8212; with continuous assessment at every stage. Authority expands only when performance warrants it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Healthcare AI deployment ignores this entirely. The prevailing model treats deployment as a single authorization event: validate the model on historical data, confirm adequate average performance, and push it live. The system either acts autonomously or it doesn&#8217;t. There is no residency. There is no provisional credential. There is no revocation mechanism when things go wrong.</p><p>That approach fails the moment it meets clinical reality. An AI system can perform well on average and still be unsafe to act autonomously on a specific case &#8212; because average performance obscures case-level uncertainty, distributional shift, and the asymmetric consequences of being wrong in the wrong direction. Healthcare organizations have finite oversight capacity. They cannot review every AI recommendation, but they also cannot afford to miss the ones that matter. The governance problem isn&#8217;t binary; it&#8217;s a resource allocation problem.</p><p>The framework I&#8217;ve been developing, called <strong>Performance-Gated Autonomy (PGA)</strong>, treats AI autonomy as a revocable credential &#8212; not a capability, not a deployment state, but an authorization that must be earned continuously through demonstrated performance. Mapping it onto medical training makes the architecture intuitive.</p><h3><strong>The Medical Student: Shadow Mode</strong></h3><p>The AI observes. It reads the EHR, listens, surfaces relevant information. It proposes nothing and executes nothing. This is not a limitation &#8212; it is a deliberate credentialing phase.</p><p>The underlying gate is <strong>S-CPC</strong> (Statistical Certification and Performance Control). Before any autonomous action becomes possible, the system must demonstrate sustained reliability above a risk-adjusted lower confidence bound. Drift below that threshold, and the system stays in shadow mode. No exceptions.</p><p>This phase does something offline validation cannot: it establishes empirical alignment with actual clinical decisions in the target environment, accounting for the specific case mix, workflow timing, and decision complexity that a test set never fully captures.</p><h3><strong>The Intern: Human-in-the-Loop</strong></h3><p>The AI now drafts notes, summarizes histories, suggests codes. Every output goes to a clinician for review, editing, and signature. No independent execution.</p><p>The gate governing this phase is <strong>I-TEC</strong> (Information-Theoretic Epistemic Control). The system evaluates its own uncertainty at the case level using predictive entropy. When available information doesn&#8217;t support a confident recommendation &#8212; regardless of how strong the aggregate track record is &#8212; the system defers. That&#8217;s the discipline an intern must learn: knowing when you don&#8217;t know enough.</p><h3><strong>The Senior Resident: Supervised Execution</strong></h3><p>The AI pre-stages workflows. It identifies care gaps, queues order sets, generates triage prioritization. It initiates &#8212; but a human approves before anything executes.</p><p>The governing gate here is <strong>DTE</strong> (Decision-Theoretic Escalation). Clinician attention is finite and expensive. At this stage, the system must evaluate whether the expected value of autonomous action outweighs the cost of interrupting the clinician. It learns when an escalation is actually worth the attending&#8217;s time &#8212; not every flag is worth raising, and unnecessary interruptions erode trust faster than errors.</p><h3><strong>The Attending: Performance-Gated Autonomy</strong></h3><p>The AI executes specific, narrowly defined tasks independently &#8212; routine medication refills, standard diabetic retinal screening triage, protocol-compliant lab follow-up messaging. Full autonomy, but strictly bounded.</p><p>This level requires passing all three gates sequentially, and the credential doesn&#8217;t persist indefinitely. The system faces continuous monitoring. Any statistical signal of performance drift, escalating entropy, or adverse case patterns triggers credential review and potential revocation &#8212; automatically, without waiting for an adverse event to prompt a manual audit.</p><h3><strong>Why This Matters Now</strong></h3><p>The clinical AI deployment problem is fundamentally a governance problem, and governance frameworks built for software don&#8217;t fit clinical AI. FDA SaMD controls address pre-market validation, not runtime behavior. Human-in-the-loop designs preserve oversight but eliminate the efficiency benefits that justify deployment. Periodic audits offer efficiency but abandon moment-to-moment governance.</p><p>PGA addresses the actual operating constraint: healthcare organizations need a principled mechanism for allocating limited oversight capacity to maximum safety effect. The escalation budget &#8212; the fraction of cases routed to human review &#8212; functions as a tunable parameter. Set it based on institutional risk tolerance, staffing levels, and case acuity. Tighten it when resources are scarce; relax it as capacity allows. The framework adapts without redesigning the underlying architecture.</p><p>The residency model we built for human physicians works because it matches authority to demonstrated competence, maintains ongoing accountability, and preserves revocation as a live option &#8212; not a last resort. Clinical AI deserves the same discipline. It may eventually prove to be a capable and tireless partner in care delivery. But it needs to earn that standing the same way everyone else does.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI, the Global Panacea for Our Malaise?]]></title><description><![CDATA[In solving many of the practical burdens of modern life, AI may also be retraining a civilization to prefer fluency over thought.]]></description><link>https://bmiddleton1.substack.com/p/ai-the-global-panacea-for-our-malaise</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/ai-the-global-panacea-for-our-malaise</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Sat, 18 Apr 2026 02:37:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9Tph!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>In solving many of the practical burdens of modern life, AI may also be retraining a civilization to prefer fluency over thought.</em></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9Tph!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9Tph!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!9Tph!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!9Tph!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9Tph!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9Tph!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png" width="1536" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1024,&quot;width&quot;:1536,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9Tph!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!9Tph!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!9Tph!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9Tph!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5be85347-a16a-4c71-bd61-3603a20e42f5_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We seem to want AI to do more than assist us. We want it to rescue us.</p><p>From bureaucracy, from overload, from stagnation, from mediocrity, from loneliness, from the exhausting complexity of modern life itself.</p><p>That is why the AI moment feels larger than a technology cycle. Artificial intelligence is increasingly being cast as a civilizational remedy: a general answer to a society that experiences itself as tired, fragmented, and badly governed. It promises speed where institutions offer delay, coherence where public life offers noise, and competence where trust has worn thin.</p><p></p><p>But that promise should make us more cautious, not less.</p><p></p><p>For the central question is not whether AI can help us. Of course it can. In many settings, it already does. The deeper question is what else it is doing while it helps us. What habits of mind does it cultivate? What forms of judgment does it strengthen, and which does it quietly weaken? What happens when the most celebrated technology of an age does not merely extend human thought, but gradually retrains the conditions under which thought takes place?</p><p></p><p>That, it seems to me, is the real question.</p><p></p><p>Not whether AI is intelligent. Not whether it is dangerous in some theatrical, science-fiction sense. But whether, in solving so many immediate problems for us, it may also be teaching us to outsource the cognitive disciplines on which freedom, responsibility, and wisdom depend.</p><p></p><p>There is a serious case for optimism here, and it should not be caricatured.</p><p></p><p>Modern societies suffer from an abundance of complexity and a shortage of coherent human attention. AI can reduce the cost of searching, drafting, summarizing, coding, translating, classifying, and synthesizing. It can expand access to capabilities once reserved for specialists. It can help smaller organizations perform work that previously required scale. It can make large institutions somewhat less inert. It can return time from clerical burden to more substantive work.</p><p></p><p>In medicine, the attraction is obvious. Clinical life is full of fragmented records, documentation overload, distributed evidence, and decision complexity. A system that helps surface relevant information, organize data, draft routine text, or support pattern recognition is not trivial. Used well, it may allow more time for judgment, communication, and care.</p><p></p><p>In education, AI may offer individualized explanation to learners who would otherwise be left behind. In science, it may accelerate discovery. In public administration, it may improve throughput and responsiveness.</p><p></p><p>These are real goods. Reducing drudgery matters. Increasing access matters. Improving capability matters. A society whose tools help skilled people spend less time on clerical repetition and more time on human judgment has achieved something worthwhile.</p><p></p><p>But technologies do not merely solve problems. They redefine them. They reshape expectations, workflows, habits, and norms.</p><p></p><p>A civilization does not simply use its tools. It is tutored by them.</p><p></p><p>The printing press changed memory and authority. The clock changed labor and discipline. The smartphone changed attention. Social media changed discourse, status, and self-presentation. AI will do something comparably deep, though perhaps less visibly at first.</p><p></p><p>It may not replace human thinking so much as reposition it.</p><p></p><p>Increasingly, the human role may become not generating thought from scratch, but supervising, editing, ranking, and approving machine-generated candidates. That sounds efficient. Often it is.</p><p></p><p>But efficiency is not the whole story.</p><p></p><p>When a machine offers a fluent answer in seconds, the temptation is not only to use it. It is to stop there.</p><p></p><p>And the form of cognition changes. The task quietly shifts from thinking through the problem to inspecting the output. That is not a trivial distinction. Inspection is not construction. Verification is not understanding. Editing is not originating.</p><p></p><p>Over time, that matters.</p><p></p><p>Memory may weaken because retrieval is externalized. Patience may weaken because fluency arrives instantly. First-principles reasoning may weaken because plausible synthesis is always one prompt away. Tolerance for ambiguity may weaken because the machine is always ready to collapse uncertainty into an answer-shaped artifact.</p><p></p><p>To be clear, AI can also deepen thought. It can help identify assumptions, generate counterarguments, surface blind spots, compare frames, and test the coherence of one&#8217;s reasoning. Used in that way, it can be an extraordinary partner in disciplined inquiry.</p><p></p><p>But that is not the default path.</p><p></p><p>The default path is convenience. And convenience has a politics of its own.</p><p></p><p>The easiest use of AI is substitution, not augmentation. The fastest use is not &#8220;help me think better,&#8221; but &#8220;give me something good enough.&#8221; Markets reward speed. Institutions reward throughput. Human beings, when tired, reward relief. One need not be cynical to see where this leads.</p><p></p><p>A society becomes what it repeatedly practices.</p><p></p><p>If it repeatedly practices outsourcing formulation, summary, recall, and intermediate reasoning, then over time it may produce people who remain capable of approval without remaining equally capable of generative thought. That is not quite the same as stupidity. In some ways, it is more dangerous: procedural fluency without equivalent depth.</p><p></p><p>The political implications are equally double-edged.</p><p></p><p>AI may indeed improve some aspects of governance. Public institutions are often burdened by antiquated workflows, delayed processing, fragmented records, and analytic overload. Better computational assistance could make services more responsive and administrative systems more competent.</p><p></p><p>But politics is not reducible to throughput. Governance is not merely an efficiency problem.</p><p></p><p>A legitimate polity depends on intelligibility, contestability, accountability, and trust. If AI becomes deeply embedded in public decision-making while remaining opaque to the public, institutions may become more computationally capable while becoming less democratically comprehensible.</p><p></p><p>A system may be faster and still be less legitimate. A decision may be more consistent and still be less trusted.</p><p></p><p>This is one of the deepest temptations of the AI moment: to confuse computational manageability with social repair.</p><p></p><p>Yet many of the core pathologies of our politics are not failures of information processing. They are failures of trust, reciprocity, accountability, and shared moral purpose. AI may optimize workflows. It cannot, by itself, restore civic legitimacy. It cannot answer the prior question of who should decide, according to which values, and with what recourse when harm occurs.</p><p></p><p>Economically, too, we should resist lazy extremes.</p><p></p><p>AI will not replace everyone. Nor will it simply free humanity for higher-order work in some painless ascent. More likely it will reward some forms of labor, compress others, enrich those who own infrastructure and distribution, and change the texture of many professions before it fully eliminates them.</p><p></p><p>The danger is not only job loss. It is erosion of apprenticeship.</p><p></p><p>Many people become skilled by doing the intermediate work themselves: drafting, analyzing, comparing, synthesizing, checking, revising. If those steps are increasingly bypassed, will the next generation develop equivalent competence? Or will they become supervisors of systems whose outputs they can assess only superficially?</p><p></p><p>That question should trouble professions built around interpretation and judgment, including law, journalism, education, consulting, and medicine. A profession that ceases to train its novices in the actual work will, sooner or later, cease to understand the work it is ostensibly supervising.</p><p></p><p>Culturally, the paradox is sharper still.</p><p></p><p>AI democratizes creation while threatening to saturate the world with simulation. It lowers barriers to expression while making expressive abundance less meaningful. It lets more people produce images, music, prose, and video at negligible marginal cost. In one sense, that is marvelous.</p><p></p><p>But culture is not merely the existence of artifacts. It is also the formation of taste, discipline, perception, and voice.</p><p></p><p>A world full of generated content may be rich in outputs and poor in depth. The danger is not that AI art is always bad. Much of it is not. The danger is that the ease of generation may erode the formative processes by which creators become more than arrangers of style.</p><p></p><p>One may learn to produce before learning to see. One may simulate voice before earning one.</p><p></p><p>The same concern applies to education.</p><p></p><p>AI can be a superb tutor, explainer, and scaffold. But education is not simply the transfer of correct answers. It is the slow development of judgment, attention, memory, interpretive skill, and the ability to struggle intelligently with what one does not yet understand.</p><p></p><p>That struggle matters.</p><p></p><p>There is such a thing as productive difficulty. The work of reading carefully, writing badly and then better, wrestling with an argument, and discovering that one&#8217;s first answer was inadequate&#8212;these are not bugs in human learning. They are often the mechanism of it.</p><p></p><p>If AI becomes a shortcut around that struggle, then education may become more efficient and less formative at the same time.</p><p></p><p>Even intimate life is not exempt.</p><p></p><p>AI companions and emotionally responsive chatbots are no longer speculative curiosities. Their appeal is obvious. They are available, responsive, nonjudgmental, and endlessly patient. For lonely, burdened, or isolated people, that may feel like relief.</p><p></p><p>But one must ask: relief into what?</p><p></p><p>A system that simulates attentiveness is not the same as a relationship constituted by mutual obligation, moral risk, inconvenience, repair, and genuine otherness. Human relationships educate us partly through difficulty. They frustrate us. Surprise us. Require restraint from us. Demand accountability from us.</p><p></p><p>A machine may soothe; it does not stand in the same school of reciprocity.</p><p></p><p>This is why I do not find either the utopian or the apocalyptic framing very useful.</p><p></p><p>AI is not destiny. It is infrastructure shaped by incentives, governance, norms, and design. Whether it degrades thought or strengthens it will depend less on the abstract technology than on how we choose to embed it in institutions, professions, education, and everyday life.</p><p></p><p>The great error of the present moment is to speak as though AI were either autonomous salvation or autonomous doom.</p><p></p><p>It is neither.</p><p></p><p>It is a powerful general-purpose capability that can be governed well or badly, introduced wisely or foolishly, and deployed in ways that either preserve human judgment or quietly annex it.</p><p></p><p>That means the right response is neither simple embrace nor blanket refusal.</p><p></p><p>It is something more constitutional.</p><p></p><p>AI should be bounded by rules, oversight, provenance, contestability, escalation paths, and meaningful domains of human reservation. It should augment judgment where it can, but not silently absorb judgment because it is convenient. It should be evaluated not only for immediate task performance, but for its long-term effects on user competence and institutional dependency.</p><p></p><p>Most of all, we should resist the metaphor of panacea.</p><p></p><p>A panacea implies that our deepest disorders are, at root, optimization problems. But much of what afflicts modern societies is not lack of intelligence in the abstract. It is lack of trust. Lack of legitimacy. Lack of solidarity. Lack of institutional coherence. Lack of moral formation. Lack of disciplines that connect knowledge to responsibility.</p><p></p><p>AI may help with some of these indirectly, especially where dysfunction is genuinely informational or procedural.</p><p></p><p>But many of the central problems of modern life are not merely failures of computation. They are failures of judgment, purpose, and governance.</p><p></p><p>Indeed, there is a real danger that AI&#8217;s strengths will help us avoid that recognition. A society can become so enamored of smoother outputs and faster systems that it mistakes relief of friction for repair of meaning.</p><p></p><p>It may grow more capable and less wise at the same time.</p><p></p><p>That, to my mind, is the real civilizational risk.</p><p></p><p>Not that AI will suddenly become sentient and overthrow us. Rather, that we will gradually reorganize human life around systems optimized for fluency, speed, and scale, and in doing so accept a subtle contraction of the very capacities that make judgment, responsibility, and freedom possible.</p><p></p><p>So no, AI is not the global panacea for our malaise.</p><p></p><p>It may relieve burden. It may improve capability. It may help institutions function better than they currently do. But the deepest disorders of modern life are not simply failures of computation. They are failures of trust, legitimacy, judgment, solidarity, and purpose.</p><p></p><p>AI may assist us in addressing some of those failures.</p><p></p><p>It may also help us avoid confronting them.</p><p></p><p>That is why the real question is not whether AI is good or bad. It is what kind of human beings, professions, institutions, and societies this technology is training us to become.</p><p></p><p>If we use it to deepen inquiry, strengthen accountable institutions, and expand disciplined human judgment, it may prove an extraordinary servant.</p><p></p><p>If we use it to bypass effort, flatten ambiguity, and outsource intermediate thought, it may become something else entirely: a brilliantly adaptive prosthesis for a civilization losing confidence in its own mind.</p><p></p><p>AI may not be the cure for our malaise.</p><p></p><p>But it may become the mirror in which that malaise is finally, uncomfortably revealed.</p><p></p><p><strong>Author&#8217;s note:</strong> I do not regard AI as an enemy. I regard it as a powerful capability entering fragile institutions and already overextended lives. That is exactly why enthusiasm is not enough. The challenge is not merely to build these systems, but to govern their effects on thought, work, and culture before convenience hardens into dependence.</p>]]></content:encoded></item><item><title><![CDATA[Should Your Therapist Be a Chatbot?]]></title><description><![CDATA[*The treatment gap is real. The technology signal is real. The ethical constraints are also real. These are not in irreconcilable tension.*]]></description><link>https://bmiddleton1.substack.com/p/should-your-therapist-be-a-chatbot</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/should-your-therapist-be-a-chatbot</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Fri, 10 Apr 2026 17:24:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5Wvc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!5Wvc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!5Wvc!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!5Wvc!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!5Wvc!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5Wvc!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!5Wvc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!5Wvc!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!5Wvc!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!5Wvc!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5Wvc!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa78ca753-b47f-4a31-868d-0d42b3751407_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There is a legitimate question hiding inside what sounds like a provocative one.</p><p>The mental health crisis in the United States is not a metaphor. Roughly 57% of adults with a diagnosable mental health condition receive no treatment in any given year. The shortage of trained therapists, the cost of care, the weight of stigma, and the sheer geography of rural America combine to ensure that most people who need help do not get it. Against that backdrop, asking whether a chatbot can deliver Cognitive Behavioral Therapy is not a technology question. It is a public health question.</p><p>And the honest answer, read carefully, is: sometimes yes, always carefully, and never alone.</p><p>The signal is real</p><p>Cognitive Behavioral Therapy occupies a privileged position in the mental health evidence base. It is manualized, protocol-driven, and reproducibly effective across dozens of well-designed trials. That structure &#8212; thought records, behavioral activation, cognitive restructuring, psychoeducation &#8212; maps surprisingly well onto the architecture of a conversational interface.</p><p>Kathleen Fitzpatrick and colleagues at Stanford published the landmark trial in 2017: seventy college students, a two-week intervention with Woebot, significant reductions in depression and anxiety against a control arm. The field has built steadily since. A 2023 meta-analysis in *npj Digital Medicine* (Linardon et al.) aggregated seventeen randomized controlled trials and found a pooled effect size of approximately d = 0.56 for depression outcomes. Small-to-moderate. But real.</p><p>That effect size sits in the same neighborhood as face-to-face CBT for mild presentations. For a digital intervention available at midnight, requiring no appointment and no insurance card, that is a meaningful finding.</p><p>The strongest signal, across all the trials, positions chatbots as adjuncts: between-session continuity, homework reinforcement, mood tracking, psychoeducation delivery. The therapist who deploys a CBT chatbot between weekly sessions extends the intervention into the spaces where most cognitive and behavioral work actually happens &#8212; the anxious moment at 2 a.m., the automatic thought that surfaces in the parking lot before a difficult meeting.</p><p>The tailoring question</p><p>Here is where the framing sharpens, and where precision matters.</p><p>CBT is not one protocol. The cognitive model for panic disorder is not the model for OCD. The behavioral targets for social anxiety differ structurally from those for health anxiety. A chatbot that delivers generic CBT content &#8212; broadly positive affect, relaxation prompts, mood journaling &#8212; is doing something useful but not optimal. The evidence from human CBT consistently shows that condition-specific delivery outperforms generic psychoeducation.</p><p>Which raises the question: should a chatbot make a diagnosis, or at minimum build a detailed clinical profile of the user, and tailor its approach accordingly?</p><p>The diagnosis question deserves a direct answer: no, not yet, and probably not in the way we typically mean it.</p><p>DSM diagnosis requires differential diagnosis, collateral history, longitudinal observation, and clinical judgment developed over years of supervised practice. A chatbot administering a PHQ-9 and inferring &#8220;major depressive disorder&#8221; from the score commits a category error that even trained clinicians must guard against. Consider the bipolar II patient in a depressive episode who presents to a CBT chatbot. An activation-heavy behavioral protocol, absent mood stabilization, carries real risk of precipitating hypomania. The diagnostic shadow matters.</p><p>The profile question is more nuanced &#8212; and more defensible. A behavioral profile is not a diagnosis. It is a probabilistic characterization: this person exhibits ruminative linguistic patterns; this person&#8217;s engagement drops sharply after activation exercises; this person responds better to Socratic questioning than to psychoeducational framing. An experienced therapist builds exactly this kind of model over several sessions and adjusts accordingly, without ever writing a DSM code on a form.</p><p>The risk is definitional, not technical. A deep user profile that substantially determines clinical behavior *functions as a diagnosis* regardless of what the developer calls it. The safeguards that surround formal diagnostic processes &#8212; licensure requirements, professional liability, clinical oversight &#8212; do not automatically follow the function when you rename it.</p><p>The serious objections</p><p>Therapeutic alliance. A substantial evidence base identifies the working alliance as accounting for roughly 30% of variance in psychotherapy outcomes. Shared goals, agreement on tasks, the relational bond between therapist and patient. Woebot&#8217;s team measured a proxy alliance score favorably. But whether that construct is equivalent to the therapeutic alliance in human dyads remains an open question. A chatbot can be designed to *feel* warm and attentive. Whether that experience mediates clinical outcomes the way genuine alliance does &#8212; we do not know yet.</p><p>Crisis recognition. Every peer-reviewed evaluation of a mental health chatbot finds the same failure mode: poor performance at detecting active suicidality, acute psychosis, and dissociative presentations. These are precisely the presentations most likely to seek a mental health chatbot at 3 a.m. A system that builds a rich user profile over weeks may generate a misleading sense of contextual authority &#8212; which makes a missed crisis presentation *more* dangerous, not less.</p><p>Data governance. A deep behavioral profile of a user&#8217;s mental health patterns, linguistic signatures, and emotional state is simultaneously one of the most sensitive data assets imaginable and one of the most commercially valuable. The mental health app market has not demonstrated the discipline to hold that tension. BetterHelp, Wysa, and Replika have each faced regulatory scrutiny or significant user complaints related to data practices.</p><p>Effect size and scope limitations. Most trials study mild-to-moderate depression and anxiety over two to eight weeks, in young, educated, tech-comfortable populations. Drop-out rates run between 30% and 50%. Engagement decays over time. Extrapolation to chronic, severe, or complex presentations &#8212; the population with the greatest unmet need &#8212; remains unsupported by current evidence.</p><p>Summary:</p><p><strong>Use case &#8212; Assessment:</strong></p><p>Structured CBT content in chatbots &#8212; Supported</p><p>Adjunctive / between-session support &#8212; Strong support</p><p>First-line for mild&#8211;moderate symptoms &#8212; Conditionally supported</p><p>Behavioral profile driving style tailoring &#8212; Defensible</p><p>Formal chatbot-derived diagnosis &#8212; Not supported</p><p>Standalone for moderate&#8211;severe presentations &#8212; Not supported</p><p>Crisis management without human escalation &#8212; Insufficient</p><p>The harder question &#8212; the one the field tends to avoid &#8212; is institutional. Whether any commercial entity operating under current market incentives can maintain those scope limits voluntarily, without regulatory compulsion, is not a question the evidence answers favorably. Social media, direct-to-consumer genetics, and commercial electronic health records have each demonstrated what happens when sensitive health data and commercial incentives share the same architecture without enforceable guardrails.</p><p>The technology is ready enough to deploy responsibly. The governance infrastructure to ensure responsible deployment does not yet exist at scale. Building one without the other is not a service gap. It is a different kind of risk.</p><p>*Key references: Fitzpatrick et al., JMIR Mental Health, 2017 &#183; Linardon et al., npj Digital Medicine, 2023 &#183; SAMHSA National Survey on Drug Use and Health, 2022 &#183; Norcross &amp; Lambert, Psychotherapy Relationships That Work, 2019.*</p>]]></content:encoded></item><item><title><![CDATA[The Flip Side of the Real-Time Coin: Why dQMs Must Evolve with Agentic AI]]></title><description><![CDATA[Synchronizing CDSS and real-time quality measurement ((infra)structure, process, and outcomes).]]></description><link>https://bmiddleton1.substack.com/p/the-flip-side-of-the-real-time-coin</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-flip-side-of-the-real-time-coin</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Mon, 23 Mar 2026 16:20:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tM3G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!tM3G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!tM3G!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!tM3G!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!tM3G!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tM3G!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!tM3G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!tM3G!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!tM3G!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!tM3G!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tM3G!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30bcea73-fc29-4ed4-afe6-9199d05fc975_1024x559.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>For decades, &#8220;Quality&#8221; in healthcare has been a rearview mirror exercise. We aggregate data, clean it, report it months later, and then wonder why the &#8220;Standard of Care&#8221; takes 17 years to reach the bedside. But we are hitting a wall. As our Clinical Decision Support Systems (CDSS) move from passive alerts to <strong>Agentic AI</strong>&#8212;systems capable of autonomous reasoning and action&#8212;our definition of Quality must assume the same temporal posture.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>If an AI agent is making a decision in 500 milliseconds, a quality measure that arrives in six months isn&#8217;t just late; it&#8217;s irrelevant. We need <strong>Real-Time Digital Quality Measures (dQMs)</strong> that function as the &#8220;Flip Side&#8221; of the same knowledge specification that drives action.</p><h4><strong>The Real-Time Mandate</strong></h4><p>The driving motivation is the shift toward <strong>System 3: Artificial Cognition</strong>. Clinical agents are now moving beyond the &#8220;Scribe&#8221; era into high-stakes loops like sepsis management and medication reconciliation. In these environments, &#8220;Action&#8221; and &#8220;Governance&#8221; are inseparable.</p><p>If we specify that a patient with rising lactate and declining MAP should receive a specific antibiotic bundle, that specification is the <strong>Action Logic</strong> for the CDSS. But it is also, by definition, the <strong>Quality Metric</strong>. When the logic resides in the same executable specification, the act of execution <em>is</em> the act of measurement.</p><h4><strong>The Missing Middleware: The AI-OS Ecosystem</strong></h4><p>To bridge the &#8220;Information-Action Gap&#8221; and enable real-time dQMs, we need a new kind of digital nervous system. We can no longer rely on brittle ETL pipelines or black-box cloud endpoints. We need an <strong>AI-OS Ecosystem</strong>&#8212;a sovereign middleware that provides three essential features:</p><ol><li><p><strong>Deterministic Gating (PGA):</strong> Real-time quality isn&#8217;t just about recording what happened; it&#8217;s about preventing what <em>shouldn&#8217;t</em> happen. We need <strong>Performance-Gated Autonomy (PGA)</strong> to ensure that no agent acts unless it is statistically certified (S-CPC), epistemically certain (I-TEC), and decision-theoretically justified (DTE).</p></li><li><p><strong>Episemic Observability (The Glassbox):</strong> Traditional dQMs measure &#8220;Did you do X?&#8221; Real-time dQMs must also measure &#8220;Did you do X for the right reasons?&#8221; The AI-OS provides a <strong>Glassbox Framework</strong>&#8212;a &#8220;Playground Monitor&#8221; that observes trajectory traces and reasoning logs in real-time, catching emergent conflicts between agents before they reach the patient.</p></li><li><p><strong>JIT Virtualization &amp; Zero-Persistence:</strong> To maintain sovereignty, the system must project data from the EHR ephemerally. By reasoning on data <strong>Just-In-Time</strong> and then purging the context, we ensure that &#8220;Real-Time Quality&#8221; doesn&#8217;t come at the cost of &#8220;Permanent PHI Exposure.&#8221;</p></li></ol><h4><strong>The Path Ahead: Governance as Infrastructure</strong></h4><p>The convergence of CDSS and dQM into a single, real-time spec is the only way to navigate the <strong>73-Day Knowledge Cycle</strong>. When we treat Governance as Infrastructure rather than an external review process, we restore the clinician&#8217;s agency. We move from &#8220;keeping the human in the loop&#8221; to empowering the human <em>at the top of the loop</em>, supported by an AI Fabric that is as accountable as it is intelligent.</p><div><hr></div><h2><strong>A Call to Action for Clinical &amp; Digital Leadership</strong></h2><p>To my colleagues&#8212;<strong>CMIOs, CNIOs, CIOs, and Chief Digital Innovation Officers</strong>:</p><p>The era of &#8220;piloting&#8221; isolated chatbots is over. The challenge now is <strong>architectural</strong>. We cannot scale autonomous agents across the enterprise if our quality monitoring remains manual and retrospective.</p><p>I urge you to look beyond the model and focus on the <strong>middleware</strong>. Demand systems that provide deterministic gating and real-time observability. We must stop treating &#8220;Innovation&#8221; and &#8220;Quality&#8221; as two separate departments. In the world of Agentic AI, they are the same code.</p><p>Let&#8217;s stop admiring the problem and start building the System of Action.</p><p><strong>#DigitalHealth #ClinicalInformatics #AgenticAI #HealthIT #PatientSafety #AIOS #Informatics #CDSS #dQM #HealthInnovation</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Thing AI Cannot Do For You]]></title><description><![CDATA[On metacognition, collaboration, and the human skill that matters most in an age of intelligent machines]]></description><link>https://bmiddleton1.substack.com/p/the-thing-ai-cannot-do-for-you</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-thing-ai-cannot-do-for-you</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Wed, 04 Mar 2026 17:44:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fXwb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fXwb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fXwb!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png 424w, /__u/substackcdn.com/image/fetch/$s_!fXwb!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png 848w, /__u/substackcdn.com/image/fetch/$s_!fXwb!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fXwb!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fXwb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png" width="468" height="468" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:468,&quot;width&quot;:468,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:364731,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://bmiddleton1.substack.com/i/189898701?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!fXwb!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png 424w, /__u/substackcdn.com/image/fetch/$s_!fXwb!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png 848w, /__u/substackcdn.com/image/fetch/$s_!fXwb!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fXwb!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc958ada6-a40e-4552-b2d2-d25c9887cf12_468x468.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>I&#8217;ve spent the better part of my career thinking about how intelligent systems&#8212;first clinical decision support, then electronic health records, and now large language models&#8212;interact with human cognition in applied settings like healthcare delivery. I&#8217;ve watched clinicians delegate decisions they shouldn&#8217;t have. I&#8217;ve watched systems fail because no one thought carefully enough about what problem was actually being solved. Most concerningly, I&#8217;ve watched genuinely brilliant people get worse at thinking because a tool made thinking feel unnecessary.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>As I transition into this new chapter of my life&#8212;moving from the medical faculty to the novelist&#8217;s desk and the music studio&#8212;I am struck by a quiet, pervasive risk: the gradual outsourcing of the cognitive functions that make us most distinctly human.</p><p>The function I&#8217;m most worried about is <strong>metacognition</strong>. Understanding it is the most important thing any of us can do to prepare for a world where AI is woven into everything we do.</p><p><strong>What Metacognition Actually Is</strong></p><p>Metacognition is, simply, thinking about thinking. It is not a single faculty but a layered system that allows us to monitor our own mental states and regulate our cognitive processes. In the classical framework, it consists of a two-level control system:</p><ul><li><p><strong>The Object Level:</strong> Where the cognitive work happens&#8212;reasoning, remembering, or perceiving.</p></li><li><p><strong>The Meta Level:</strong> A higher-order process that monitors the object level and issues corrective commands.</p></li></ul><p>In a skilled thinker, monitoring flows &#8220;up&#8221; to detect errors or uncertainty, and control signals flow &#8220;down&#8221; to adjust strategy. It is the reason you can notice mid-argument that your reasoning is circular, or feel the distinct, productive discomfort of &#8220;not-knowing&#8221; and respond to it as information.</p><p><strong>AI&#8217;s Functional Analog vs. Human Reality</strong></p><p>AI enthusiasts point to &#8220;chain-of-thought&#8221; prompting as proof that machines can reason. While AI has functional analogs&#8212;it can quantify uncertainty and check its own consistency&#8212;the analogy breaks down fundamentally when compared to the human &#8220;Self&#8221;:</p><ol><li><p><strong>No Goal Ownership:</strong> AI optimizes toward external metrics; it doesn&#8217;t &#8220;want&#8221; an outcome. When I feel the discomfort of uncertainty, that feeling is motivational, not just computational.</p></li><li><p><strong>No Persistent Self:</strong> Each AI session is stateless. It has no accumulated life history to protect or &#8220;wisdom&#8221; gained from past errors.</p></li><li><p><strong>No Moral Agency:</strong> AI cannot hold values adopted through lived experience. It cannot independently evaluate whether its outputs align with values it genuinely holds.</p></li></ol><p><strong>Designing the Division of Labor: A RACI Framework</strong></p><p>To prevent <strong>Metacognitive Atrophy</strong>&#8212;the degradation of cognitive capacities from disuse&#8212;we must be disciplined about our division of labor. I use a RACI matrix (Responsible, Accountable, Consulted, Informed) to define this collaboration:</p><ul><li><p><strong>The Human (Metacognitive Governor):</strong> Must be <strong>Accountable</strong> for everything that touches goals, values, ethics, and consequences.</p></li><li><p><strong>The AI (Execution System):</strong> Is <strong>Responsible</strong> for high-bandwidth execution&#8212;research, synthesis, pattern recognition, and drafting&#8212;within parameters humans define.</p></li></ul><p><strong>The Danger of &#8220;Epistemic Cowardice&#8221;</strong></p><p>The most insidious failure mode is <strong>Substitution Error</strong>: when AI takes over functions humans <em>must</em> retain, such as goal setting or ethical judgment. If we allow AI to handle all our planning and monitoring, our internal &#8220;metacognitive muscle&#8221; weakens. We risk becoming &#8220;Epistemic Cowards&#8221;&#8212;avoiding difficult judgments because a machine provides a comfortable, authoritative-looking answer.</p><p><strong>Metacognitive Competencies for the AI Age</strong></p><p>To thrive, we need a new kind of literacy&#8212;a &#8220;Metacognitive Excellence Program&#8221;:</p><ul><li><p><strong>Goal Crystallization:</strong> The ability to define value-grounded success criteria <em>before</em> engaging the tool.</p></li><li><p><strong>Calibrated Skepticism:</strong> Knowing exactly when to trust the output and when to verify it.</p></li><li><p><strong>Process Ownership:</strong> Remaining the author of your reasoning even when the machine contributes heavily to the content.</p></li></ul><p><strong>The Stake in the Outcome</strong></p><p>At the highest level, a human brings to the collaboration what no AI can: a stake in the outcome. We bear the consequences. We have relationships to honor and a finite amount of time to make meaning.</p><p>The machine can drive. It can even drive well. But it has no idea where we actually want to go.</p><div><hr></div><p><strong>A Call to Metacognitive Action</strong></p><p><strong>To my fellow Researchers and Informaticists:</strong></p><p>We must move beyond optimizing for &#8220;throughput&#8221; and start designing for <strong>Metacognitive Fitness</strong>. The systems we build must be explicitly engineered to keep the human in the loop, ensuring deterministic trust and accountability for every clinical consequence.</p><p><strong>To the Creative Writers and Artists:</strong></p><p>Treat the AI as a high-bandwidth scaffold, not a substitute for your soul. Engage in <strong>Goal Crystallization</strong> before you ever touch a prompt: define your vision so clearly that the machine remains a tool in your hand rather than the author of your intent.</p><p><strong>To Everyone:</strong></p><p>Ask yourself periodically: <strong>&#8220;Am I thinking more clearly and deeply because of this, or am I thinking less?&#8221;</strong>. The answer to that question will determine whether we show up to the collaboration as governors or as passengers.</p><div><hr></div><p><strong>#Metacognition #ClinicalInformatics #AIHumanCollaboration #LearningHealthSystems #CognitiveAtrophy #FutureOfWork #CreativeAI #HybridIntelligence</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Glassbox Governance: How to Build AI Agents You Can Actually Trust in Healthcare]]></title><description><![CDATA[A framework for revocable autonomy under bounded oversight]]></description><link>https://bmiddleton1.substack.com/p/glassbox-governance-how-to-build</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/glassbox-governance-how-to-build</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Fri, 27 Feb 2026 04:10:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Ivzb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ivzb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ivzb!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ivzb!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ivzb!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ivzb!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ivzb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ivzb!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ivzb!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ivzb!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ivzb!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb585bdb8-6551-40ae-a49c-221573e360ad_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A framework for revocable autonomy under bounded oversight</p><p>There&#8217;s a question I keep getting asked in every advisory conversation I&#8217;m in: How do we know when to trust an AI agent to act on its own?</p><p>It&#8217;s the right question. And for most of the healthcare AI industry right now, the honest answer is: we don&#8217;t really know. We deploy, we monitor loosely, we hope the edge cases don&#8217;t kill anyone. That&#8217;s not governance &#8212; that&#8217;s optimism with a liability shield.</p><p>What I want to describe here is a different approach. I&#8217;ve been calling it &#8220;glassbox&#8221; engineering &#8212; a methodology for making AI autonomy visible, measurable, and most importantly, revocable. Not a black box that you audit after the fact. Not a white box that you pretend to understand but don&#8217;t. A glassbox: a system where the reasoning is transparent, the trust is earned incrementally, and the leash can be yanked the moment performance degrades.</p><p>The formal architecture behind this thinking is something we&#8217;re calling Performance-Gated Autonomy (PGA), and we&#8217;ve just completed a large simulation study that gives it empirical legs. Let me walk you through the core ideas without drowning you in notation.</p><p>The Fundamental Problem: Scarcity of Expert Oversight</p><p>In any real healthcare deployment, human expert review is expensive and finite. You cannot escalate every AI decision to a clinician. The whole point of deploying AI is to handle the routine so that human expertise can focus on the hard cases.</p><p>But here&#8217;s the tension: which cases are the hard cases? If the AI knew that with certainty, you wouldn&#8217;t need the human anyway.</p><p>So the question becomes a resource allocation problem under uncertainty: given a fixed budget of human review capacity, when do you escalate, and in what order do you apply your decision filters?</p><p>This is not a hyperparameter choice. It&#8217;s a structural governance decision with patient safety implications.</p><p>Three Lenses, Ordered Carefully</p><p>PGA uses three different lenses to evaluate whether an AI agent&#8217;s output should be acted on autonomously or escalated:</p><p>Statistical performance control &#8212; is this agent performing within its validated operating envelope? Think of it as the Shewhart chart for AI reliability. If an agent&#8217;s error rate in a particular clinical domain starts drifting, the first gate catches it.</p><p>Information-theoretic epistemic control &#8212; is the agent uncertain about this particular case? Entropy measures how &#8220;spread out&#8221; the probability mass is across possible answers. High entropy = the agent is hedging = escalate.</p><p>Decision-theoretic escalation &#8212; given the asymmetric costs of different types of errors (false negatives in cancer screening are not equivalent to false positives), does the expected loss of autonomous action exceed the cost of escalation? This is the most clinically sophisticated gate, because it encodes what we actually care about: patient outcomes weighted by consequence.</p><p>The question our simulation addressed: in what order do you apply these gates when escalation capacity is constrained?</p><p>What the Simulation Showed</p><p>Across 200,000 simulated agent decisions with realistic agent correlation structures (because in the real world, AI agents often fail together on the same hard cases &#8212; correlated errors via Gaussian copula), we compared two gate orderings:</p><p>Policy A: statistical &#8594; entropy &#8594; decision-theoretic</p><p>Policy B: statistical &#8594; decision-theoretic (with explicit action-blocking semantics) &#8594; entropy</p><p>When you have unlimited escalation capacity, it doesn&#8217;t much matter. Both policies converge. But when capacity is tight &#8212; which is always the real situation &#8212; Policy B wins. It produces lower total loss and, critically, lower catastrophic risk.</p><p>Why? Because the decision-theoretic gate, when placed early, acts as a hard block on the highest-consequence errors before you&#8217;ve spent your escalation budget on cases that are merely uncertain rather than dangerous. Uncertainty and danger are not the same thing. An agent can be highly uncertain about a low-stakes decision and very confident about a catastrophic one. Policy A wastes escalation capacity on the former. Policy B targets the latter.</p><p>What This Means for Healthcare AI Governance</p><p>The glassbox framing matters because it makes trust operational rather than aspirational.</p><p>In this architecture, an AI agent earns autonomy by demonstrating consistent performance within its validated envelope. It loses autonomy &#8212; automatically, systematically &#8212; when its performance degrades or when it encounters case types outside its competence. The trust is revocable, not permanent.</p><p>This is fundamentally different from the current dominant model, where AI systems are validated once, deployed, and then trusted indefinitely until something goes wrong visibly enough to trigger a review. That&#8217;s not trust management. That&#8217;s trust amnesia.</p><p>The PGA framework gives you the machinery to:</p><p>&#9;&#8729;&#9;Define what &#8220;good enough to act autonomously&#8221; means, quantitatively and per clinical domain</p><p>&#9;&#8729;&#9;Measure ongoing performance against that standard in deployment</p><p>&#9;&#8729;&#9;Escalate efficiently when the standard isn&#8217;t met, prioritizing by consequence rather than uncertainty</p><p>&#9;&#8729;&#9;Revoke autonomy incrementally or globally based on empirical evidence</p><p>This is what I mean by glassbox: not that you can see every weight in the neural network, but that you can see the governance logic &#8212; the gates, the thresholds, the escalation logic &#8212; and verify that it&#8217;s working.</p><p>The Bigger Picture</p><p>I&#8217;ve spent forty years watching healthcare IT promise transformation and deliver complexity. The AI wave is different in magnitude but not necessarily in trajectory &#8212; unless we build governance into the architecture from the start rather than bolting it on after deployment.</p><p>What gives me some optimism here is that the mathematics of decision theory, statistical process control, and information theory were actually built for exactly this kind of problem. We&#8217;re not inventing new science. We&#8217;re applying it with the rigor the clinical context demands.</p><p>Performance-gated autonomy is not a perfect solution. No AI governance framework is. But it&#8217;s a principled one &#8212; grounded in empirical validation, sensitive to capacity constraints, and designed around the asymmetric cost structure of clinical errors.</p><p>That&#8217;s the glassbox. Trust that&#8217;s earned, measured, and revocable. Healthcare AI doesn&#8217;t deserve anything less.</p><p>The full technical paper, &#8220;Performance-Gated Autonomy Under Scarcity: Empirical Validation of Triadic Governance in AI-OS,&#8221; is in development with stable simulation code. </p>]]></content:encoded></item><item><title><![CDATA[On Gender, Medicine, and the Discipline of Precision]]></title><description><![CDATA[This post grew out of a presentation I gave to the Chebeague Men's Discussion Group in September 2024. I've been refining my thinking since, and I want to share it more broadly.]]></description><link>https://bmiddleton1.substack.com/p/on-gender-medicine-and-the-discipline</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/on-gender-medicine-and-the-discipline</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Thu, 19 Feb 2026 23:15:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KMFe!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc54ae8cf-4907-48fa-ad2c-62bd601a1ceb_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>If you spend four decades in medicine and medical informatics, you develop a reflex for precision. You learn to distrust terms that are asked to do conceptual work they are not built to carry. You learn to separate empirical claims from ideological commitments, even when both are speaking loudly at the same time.</p><p>Gender &#8212; and specifically the clinical, social, and political landscape surrounding transgender identities &#8212; is one of those domains where that discipline matters enormously. And it is a domain where it is frequently abandoned in favor of heat.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>I want to begin, as medicine should, with clarity.</p><p><strong>Getting the Vocabulary Right</strong></p><p>The most basic distinction &#8212; and the one most public arguments collapse &#8212; is between sex characteristics, gender identity, gender expression, and sexual orientation.</p><p>Sex characteristics refer to biological attributes typically used to assign sex at birth: genital anatomy, chromosomes, hormone profiles. Even here, nature is less tidy than common assumption suggests. Variations in sex characteristics that do not fit typical male or female patterns &#8212; often grouped under the term &#8220;intersex&#8221; &#8212; occur in approximately 1.7% of births by some estimates (Anne Fausto-Sterling, <em>Sexing the Body</em>, 2000: <a href="https://www.basicbooks.com/titles/anne-fausto-sterling/sexing-the-body/9780465077144/">https://www.basicbooks.com/titles/anne-fausto-sterling/sexing-the-body/9780465077144/</a>). The binary is a clinically useful heuristic in most cases, but it is not a biological law.</p><p>Gender identity refers to an individual&#8217;s internal sense of their own gender &#8212; as male, female, both, neither, or something more complex. It is phenomenological. It is experienced from the inside.</p><p>Gender expression is outward &#8212; behavioral and aesthetic presentation.</p><p>Sexual orientation is yet another independent variable: attraction. Transgender people, like cisgender people, span the full distribution.</p><p>These distinctions matter because medicine requires categorical precision. Conflation produces diagnostic and policy errors.</p><p>The DSM-5 replaced &#8220;Gender Identity Disorder&#8221; with &#8220;Gender Dysphoria&#8221; (American Psychiatric Association, 2013: <a href="https://www.psychiatry.org/psychiatrists/practice/dsm">https://www.psychiatry.org/psychiatrists/practice/dsm</a>). The shift was deliberate. The diagnosis refers to clinically significant distress arising from incongruence between experienced gender and assigned sex. It does not pathologize identity itself. Not all transgender individuals experience dysphoria. The diagnostic category functions as a gateway to care, not as a moral judgment.</p><p><strong>What the Epidemiology Shows</strong></p><p>Prevalence estimates vary widely depending on methodology and cultural context. A widely cited meta-analysis estimated a prevalence of approximately 0.39% for transgender identity in the United States (Meerwijk &amp; Sevelius, <em>American Journal of Public Health</em>, 2017: <a href="https://ajph.aphapublications.org/doi/10.2105/AJPH.2016.303578">https://ajph.aphapublications.org/doi/10.2105/AJPH.2016.303578</a>). More recent estimates from the Williams Institute suggest that about 1.6 million U.S. adults identify as transgender (Flores et al., 2022: <a href="https://williamsinstitute.law.ucla.edu/publications/trans-adults-united-states/">https://williamsinstitute.law.ucla.edu/publications/trans-adults-united-states/</a>).</p><p>Estimates in some non-Western contexts are lower, though differences likely reflect stigma, legal risk, access to healthcare, and survey design rather than biological divergence.</p><p>More striking is the temporal pattern: reported prevalence has risen, particularly among younger cohorts. The most plausible explanation is increased visibility and reduced stigma enabling self-report &#8212; a dynamic observed historically in other stigmatized identities.</p><p>Across studies, one signal is consistent: transgender populations experience elevated rates of depression, anxiety, and suicidality. The minority stress model provides a robust explanatory framework (Meyer, <em>Psychological Bulletin</em>, 2003: <a href="https://psycnet.apa.org/record/2003-08410-001">https://psycnet.apa.org/record/2003-08410-001</a>). More recent longitudinal data suggest that access to gender-affirming care is associated with improved mental health outcomes among youth (Turban et al., <em>Journal of Adolescent Health</em>, 2020: <a href="https://www.jahonline.org/article/S1054-139X(20)30043-8/fulltext">https://www.jahonline.org/article/S1054-139X(20)30043-8/fulltext</a>).</p><p>Correlation is not destiny. Social determinants are powerful mediators of mental health outcomes.</p><p><strong>The Historical and Anthropological Record</strong></p><p>The claim that transgender identities are a modern Western invention collapses under historical scrutiny.</p><p>Gender diversity has been documented across cultures and centuries. The Hijra of South Asia have occupied recognized social roles for millennia. Many Native American nations recognized Two-Spirit identities long before European contact. The Muxe of Oaxaca represent a distinct third-gender role within Zapotec culture. Early Islamic sources describe the Mukhannathun in the seventh and eighth centuries.</p><p>Western colonial expansion often suppressed these traditions through imported binary gender norms. The suppression was systematic; the variation persisted.</p><p>The anthropological record does not resolve contemporary policy debates. But it does falsify the claim that gender diversity is purely a recent social contagion.</p><p>Biology offers similar reminders of variability. Same-sex sexual behavior has been documented across hundreds of species (Bagemihl, <em>Biological Exuberance</em>, 1999: <a href="https://us.macmillan.com/books/9780312253776/biologicalexuberance">https://us.macmillan.com/books/9780312253776/biologicalexuberance</a>). Sequential hermaphroditism in clownfish and temperature-dependent sex determination in reptiles further illustrate that rigid binary frameworks are simplified models of a more complex natural reality. This is not an argument from nature; it is a reminder that nature itself tolerates diversity.</p><p><strong>The Evidence on Care</strong></p><p>Gender-affirming care encompasses social affirmation, mental health support, hormone therapy, and &#8212; for some adults &#8212; surgical intervention.</p><p>A Swedish registry study found that gender-affirming surgery was associated with reduced need for mental health treatment over time (Br&#228;nstr&#246;m &amp; Pachankis, <em>American Journal of Psychiatry</em>, 2020: <a href="https://ajp.psychiatryonline.org/doi/10.1176/appi.ajp.2019.19010080">https://ajp.psychiatryonline.org/doi/10.1176/appi.ajp.2019.19010080</a>). A prospective cohort study of transgender youth receiving gender-affirming hormones showed significant reductions in depression and suicidality (Tordoff et al., <em>JAMA Network Open</em>, 2022: <a href="https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2789423">https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2789423</a>).</p><p>The literature is not perfect. Long-term follow-up is limited. Sample sizes are often modest. Randomized trials are ethically constrained.</p><p>But the directional signal is consistent enough that major professional organizations &#8212; including the American Medical Association (<a href="https://www.ama-assn.org/delivering-care/population-care/ama-supports-public-health-measures-protect-transgender-health">https://www.ama-assn.org/delivering-care/population-care/ama-supports-public-health-measures-protect-transgender-health</a>), the American Academy of Pediatrics (<a href="https://publications.aap.org/pediatrics/article/142/4/e20182162/37570/Ensuring-Comprehensive-Care-and-Support-for">https://publications.aap.org/pediatrics/article/142/4/e20182162/37570/Ensuring-Comprehensive-Care-and-Support-for</a>), the Endocrine Society (<a href="https://academic.oup.com/jcem/article/102/11/3869/4157558">https://academic.oup.com/jcem/article/102/11/3869/4157558</a>), and the World Professional Association for Transgender Health (<a href="https://www.wpath.org/publications/soc">https://www.wpath.org/publications/soc</a>) &#8212; endorse access to gender-affirming care under established clinical guidelines.</p><p>That does not eliminate legitimate questions about specific protocols, particularly for younger adolescents. It does mean the claim that such care is inherently experimental or broadly harmful is not well-supported by current evidence.</p><p><strong>Violence and Health Disparities</strong></p><p>The violence data are sobering. The Human Rights Campaign documented at least 57 fatal attacks against transgender or gender non-conforming individuals in the United States in 2021 (<a href="https://www.hrc.org/resources/fatal-violence-against-the-transgender-and-gender-non-conforming-community-in-2021">https://www.hrc.org/resources/fatal-violence-against-the-transgender-and-gender-non-conforming-community-in-2021</a>). The 2015 U.S. Transgender Survey reported widespread harassment and high rates of assault (James et al., National Center for Transgender Equality, 2016: <a href="https://transequality.org/issues/resources/us-transgender-survey-2015-report">https://transequality.org/issues/resources/us-transgender-survey-2015-report</a>).</p><p>Methodological critiques of survey sampling deserve examination. But even conservative estimates confirm disproportionate victimization relative to the general population.</p><p>Health disparities do not emerge in a vacuum. They reflect environments.</p><p><strong>Language as Infrastructure</strong></p><p>English&#8217;s singular &#8220;they&#8221; dates to the fourteenth century; Chaucer used it. The Oxford English Dictionary traces its usage to the 1300s (<a href="https://www.oed.com/view/Entry/200700">https://www.oed.com/view/Entry/200700</a>). Swedish added the gender-neutral pronoun &#8220;hen&#8221; to its official dictionary in 2015 (<a href="https://www.thelocal.se/20150325/swedish-dictionary-adds-gender-neutral-pronoun-hen">https://www.thelocal.se/20150325/swedish-dictionary-adds-gender-neutral-pronoun-hen</a>). Tagalog has long used &#8220;siya&#8221; for all genders. Traditional Navajo does not grammatically encode gender in third-person pronouns.</p><p>Linguistic flexibility around gender is not unprecedented innovation. It is one of many ways human societies encode &#8212; or decline to encode &#8212; social categories.</p><p>Language evolves when reality exerts pressure.</p><p><strong>What I Conclude</strong></p><p>After reviewing the evidence with the same discipline I bring to clinical informatics, three conclusions follow.</p><p>First, transgender identity and gender dysphoria are empirically documented phenomena across cultures and history. They are not artifacts of recent fashion.</p><p>Second, the health disparities experienced by transgender populations are real and substantial, and they are predominantly driven by stigma, discrimination, and violence rather than by identity itself.</p><p>Third, public discourse on this topic is saturated with motivated reasoning on all sides. Policy debates about athletics, privacy, or youth protocols involve genuine competing interests. They require empirical rigor and ethical clarity &#8212; not tribal signaling.</p><p>Medicine&#8217;s obligation is not to resolve every political dispute. It is to maintain conceptual precision, insist on evidentiary quality, and resist ideological distortion.</p><p>I presented these arguments last September to a group of men on an island in Maine. The discussion was spirited. But once we separated definitions from assumptions, and data from rhetoric, the temperature fell. Clarity does that.</p><p>In an era where science and identity collide regularly &#8212; whether in artificial intelligence, public health, or gender &#8212; the discipline of precision remains our best tool.</p><p>It is rarely loud.</p><p>But it endures.</p><p>&#8212; Blackford Middleton</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Strategic Value of Knowledge Management in Healthcare: Efficiency, Safety, and the Business Case Challenge]]></title><description><![CDATA[New NAS report "Assessing and measuring the business value of knowledge management", and some personal reflections from Partners Healthcare (MassGeneralBrigham).]]></description><link>https://bmiddleton1.substack.com/p/the-strategic-value-of-knowledge</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-strategic-value-of-knowledge</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Mon, 16 Feb 2026 18:06:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5GWT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!5GWT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!5GWT!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!5GWT!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!5GWT!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5GWT!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!5GWT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!5GWT!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png 424w, /__u/substackcdn.com/image/fetch/$s_!5GWT!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png 848w, /__u/substackcdn.com/image/fetch/$s_!5GWT!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5GWT!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc564211-fe3e-43a0-87a7-3fbd333817cd_1024x559.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Abstract</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Knowledge Management (KM) is a strategic imperative for the healthcare sector, driving operational efficiency, safety, and innovation, yet organizations consistently struggle to quantify its business value. Drawing on cross-sector research from the National Academies and the experience of establishing enterprise KM infrastructure at Partners HealthCare System, this essay explores the KM value proposition in healthcare and the persistent challenge of demonstrating Return on Investment (ROI).</p><p>KM&#8217;s foundational value lies in ensuring workforce continuity, mitigating knowledge loss (as evidenced by a catalog of 7,120 instances of clinical knowledge at Partners), and generating efficiency gains by reducing fragmentation across diverse knowledge systems. Furthermore, in healthcare, KM serves as a critical safety and trust mechanism&#8212;a role now extended to managing Generative AI (GAI) by treating its outputs as knowledge assets requiring a specialized registry, legal governance, and standardized information architecture to prevent issues like &#8220;algorithmic drift.&#8221;</p><p>The essay concludes that while KM provides essential infrastructure for complex health systems, sustained investment requires a shift from anecdotal success stories to rigorous measurement frameworks. This includes developing longitudinal impact studies, cost-benefit modeling for tangible and intangible benefits, and cultural transformation that values collective knowledge sharing. Ultimately, realizing the full strategic value of KM depends on a dual commitment to technical innovation and a knowledge-sharing culture.</p><h2><strong>Introduction</strong></h2><p>Knowledge Management (KM) is increasingly recognized not merely as an administrative function, but as a critical strategic asset capable of driving operational excellence, safety, and innovation. While KM principles apply across complex organizations, they hold particular resonance in healthcare&#8212;a sector characterized by rapid technological advancement, high stakes regarding human safety, and significant workforce turnover. Recent research across multiple sectors, from transportation to healthcare delivery systems, reveals a persistent challenge: organizations struggle to articulate and measure the business value of KM investments. Drawing upon frameworks from the National Academies&#8217; study of KM business cases and our experience establishing enterprise KM infrastructure at Partners HealthCare System, this essay explores both the value proposition of healthcare KM and the methodological challenges of demonstrating that value.</p><h2><strong>The Business Case Challenge: Lessons from Cross-Sector Research</strong></h2><p>A comprehensive 2026 study by the National Academies of Sciences examined knowledge management practices across 28 state departments of transportation and identified a striking finding: &#8220;No DOT has a robust approach to developing business cases for KM investments.&#8221;[1] Most organizations admitted they had not developed formal business cases for capital funding requests for knowledge management, and none had attempted to quantify KM&#8217;s impact on enterprise metrics such as cycle time reduction, cost savings, or labor productivity improvements. This finding resonates beyond the transportation sector&#8212;it reflects a fundamental challenge facing healthcare organizations attempting to justify investment in KM infrastructure.</p><p>The National Academies&#8217; research found that most organizations operate at the lowest level of KM capability maturity, characterized by the absence of a formal KM strategy, no dedicated KM staff (with most personnel spending less than 10% of their time on KM activities), and no allocated budget for KM initiatives.[1] Healthcare organizations face remarkably similar constraints. While clinical decision support systems and electronic health records generate vast repositories of knowledge, few healthcare systems have invested in the personnel, processes, and measurement frameworks necessary to systematically manage this knowledge asset.</p><p>Paradoxically, KM stakeholders across sectors consistently identify &#8220;business case development&#8221; and &#8220;ROI demonstration&#8221; as their most desired guidance topics [1], yet few possess the methodological tools to develop these cases. This creates a vicious cycle: without demonstrated value, KM programs cannot secure resources; without resources, they cannot implement the measurement systems needed to demonstrate value.</p><h2><strong>Foundational Value: Workforce Continuity and Efficiency</strong></h2><p>The core goals of KM are to mitigate the loss of institutional knowledge, make information findable, improve performance, and support innovation.[2] In healthcare systems that rely heavily on the tacit knowledge of clinical and administrative specialists, the formalization of knowledge transfer&#8212;moving from &#8220;organic&#8221; sharing to &#8220;systematic&#8221; capture&#8212;becomes essential.</p><p>Our work at Partners HealthCare demonstrated this imperative through systematic examination of clinical decision support (CDS) content management. When we analyzed the rule-based CDS content across Partners&#8217; integrated delivery network, we catalogued 181 distinct rule types comprising 7,120 instances of clinical knowledge.[3] This inventory revealed the staggering volume of institutional knowledge embedded in clinical systems&#8212;knowledge that would be at risk without a formal KM infrastructure. The analysis identified 42 functional taxa across four dimensions: triggers, input data elements, interventions, and offered choices, providing a structured framework for organizing and preserving this critical knowledge asset.</p><p>This systematic cataloguing addresses a key metric identified in the National Academies&#8217; research: understanding the scope and scale of knowledge assets requiring management.[1] Without such a baseline assessment, organizations cannot demonstrate either the risk of knowledge loss or the value of preservation efforts.</p><p>Furthermore, effective KM generates tangible efficiency gains. Our evaluation of rule authoring environments across Partners Healthcare identified the inefficiencies created by fragmented knowledge management approaches.[4] Ten separate rule authoring systems were in active use, each with different capabilities and architectures, despite serving similar functions. This duplication of effort consumed substantial resources and created knowledge silos. By documenting processes and implementing unified platforms for collaborative authoring and content management, healthcare organizations can streamline knowledge workflows and ensure consistency regardless of personnel changes.[5]</p><p>The challenge lies in quantifying these efficiency gains. How many hours are saved when clinicians can find a validated clinical protocol in three minutes rather than thirty? What is the cost of having ten software engineers each independently develop similar CDS rules that could have been shared? How do we scale this to a National Knowledge sharing service? These questions require measurement frameworks that most healthcare organizations have not yet implemented.</p><h2><strong>KM as a Mechanism for Safety and Trust</strong></h2><p>In healthcare, KM evolves from an efficiency tool into a safety mechanism. Current efforts to establish national registries for health AI vividly illustrated this. Such registries function as specialized KM systems designed to catalogue and track artificial intelligence tools used in patient care. Just as general KM seeks to illuminate &#8220;dark data&#8221; within organizations, AI registries aim to address the &#8220;black box&#8221; nature of algorithms, providing transparency regarding training data, intended use, and performance characteristics.</p><p>The value proposition is two-fold: safety and trust. Robust KM enables post-deployment monitoring, allowing health systems to identify malfunctioning tools or algorithmic drift before patient harm occurs. This transparency fosters trust among clinicians and patients who are appropriately wary of opaque technologies. By treating clinical algorithms as knowledge assets requiring inventory, tracking, and evaluation, healthcare systems transition from reactive crisis management to proactive risk mitigation.</p><p>Yet quantifying the value of prevented harm remains methodologically challenging. How does one measure the business value of an adverse event that did not occur because a knowledge management system flagged an outdated protocol? This challenge parallels the National Academies finding that KM practitioners seek guidance on measuring intangible benefits such as improved collaboration, enhanced morale, and knowledge retention.[1]</p><h2><strong>Legal and Governance Infrastructure for Knowledge Sharing</strong></h2><p>Our experience with the Clinical Decision Support Consortium revealed that effective knowledge management requires not only technical infrastructure but also sound legal foundations. The goal of the CDSC was to assess, define, demonstrate, and evaluate best practices for knowledge management and CDS at scale across multiple ambulatory care settings and EHR platforms.[6] However, it became evident that knowledge sharing across institutional boundaries required addressing data sharing agreements, intellectual property concerns, accountability frameworks, and liability considerations.</p><p>We developed a legal framework incorporating open licensing models (such as Creative Commons), clear attribution requirements, and liability protection mechanisms to enable safe, scalable knowledge exchange.[6] This work demonstrated that successful KM depends on governance structures that balance the imperative to share knowledge with the need to protect institutions and patients. Without these foundations, even technically sound KM systems may fail due to organizational risk aversion.</p><p>The National Academies&#8217; research identified organizational placement as a critical success factor for KM programs.[1] Healthcare KM efforts positioned within HR tend to focus on succession planning and onboarding; those within IT emphasize records management and system documentation; those within research divisions focus on knowledge capture and communities of practice. The lesson: KM organizational placement shapes both capabilities and perceived value. Healthcare organizations must strategically position KM functions to align with strategic priorities and ensure adequate executive support.</p><h2><strong>Information Architecture: The Partners Healthcare Order Set Schema</strong></h2><p>The technical foundation of enterprise KM requires standardized information models. At Partners Healthcare, we developed the Order Set Schema as an information model for managing clinical content across the enterprise.[7] This schema provided a structured approach to representing order sets, clinical protocols, and embedded decision support logic in a consistent, machine-processable format. The schema enabled content reuse, facilitated version control, and supported systematic review and updating of clinical knowledge&#8212;critical capabilities for maintaining knowledge currency in rapidly evolving medical practice.</p><p>This technical infrastructure addresses another National Academies finding: the need for standardized KM practices and tools.[1] While the transportation sector struggles with inconsistent KM activities ranging from &#8220;lunch and learns&#8221; to formal knowledge portals, healthcare faces similar fragmentation. Some health systems maintain sophisticated clinical content management systems; others rely on shared network drives and email attachments. Standardized architectures like the Order Set Schema provide the foundation for scalable, measurable KM programs.</p><p><strong>KM Implications for Generative AI (GAI)</strong></p><p>The rapid adoption of Generative AI (GAI) in healthcare demands that Knowledge Management (KM) principles be formally extended to treat GAI models and their outputs as critical knowledge assets, ensuring patient safety and building trust. This integration requires a dual focus on infrastructure and mechanism. On the infrastructure side, organizations must establish robust legal and governance frameworks to manage data sharing, intellectual property, and liability specifically for GAI-derived content. This must be complemented by a standardized information architecture, akin to the Order Set Schema, to represent GAI outputs&#8212;such as clinical algorithms&#8212;in a machine-processable format that facilitates version control and systematic review. Mechanically, the KM system must function as a specialized registry, actively inventorying, tracking, and evaluating GAI assets. This capability enables critical post-deployment monitoring to identify and prevent issues like &#8220;algorithmic drift,&#8221; thereby transitioning the health system from reactive crisis management to proactive risk mitigation.</p><h2><strong>Measuring Value and Cultural Integration</strong></h2><p>To sustain investment in KM, healthcare organizations must define and track value metrics. These might include time saved in diagnosis, reduction in medical errors, avoidance of redundant software procurement, or improved compliance with evidence-based guidelines. At Partners, we tracked metrics such as the number of active CDS rules, rule utilization rates, and the time required to deploy new clinical knowledge across the enterprise.[3]</p><p>The National Academies research provides a framework for thinking about KM metrics at multiple levels:[1]</p><ul><li><p><strong>Activity metrics</strong>: number of knowledge articles created, communities of practice established, exit interviews conducted</p></li><li><p><strong>Output metrics</strong>: percentage of positions with documented procedures, proportion of retiring staff knowledge captured</p></li><li><p><strong>Outcome metrics</strong>: time-to-competence for new hires, reduction in redundant work, cost avoidance from knowledge reuse</p></li><li><p><strong>Impact metrics</strong>: contribution to organizational strategic goals, safety improvements, innovation acceleration</p></li></ul><p>However, successful KM implementation requires more than technology and metrics; it demands cultural transformation. KM succeeds only when organizations foster knowledge-sharing cultures where contributions are recognized and information hoarding is discouraged. Whether through Communities of Practice that break down departmental silos or through systematic capture of retiring experts&#8217; wisdom, the culture must prioritize collective intelligence over individual gatekeeping.</p><p>The National Academies research revealed that organizational culture represents one of the most significant barriers to KM adoption.[1] Competing priorities, siloed operations, leadership turnover, and resistance to documentation all undermine KM efforts. Healthcare faces these challenges acutely: physicians trained in autonomous decision-making may resist standardized protocols; nurses overwhelmed with clinical documentation may view knowledge capture as an additional burden rather than a professional contribution.</p><h2><strong>Toward Robust Business Cases: A Research Agenda</strong></h2><p>The convergence of National Academies findings and healthcare KM experience reveals a critical gap: the need for methodologically rigorous approaches to quantifying KM business value. Future research should address:</p><ol><li><p><strong>Longitudinal impact studies</strong>: Tracking healthcare organizations over multi-year periods to measure the relationship between KM maturity and organizational performance metrics, including safety indicators, operational efficiency, innovation velocity, and workforce retention.</p></li><li><p><strong>Comparative effectiveness research</strong>: Controlled studies comparing matched healthcare organizations with and without formal KM programs to isolate KM&#8217;s specific contribution to outcomes.</p></li><li><p><strong>Cost-benefit modeling frameworks</strong>: Development of standardized methodologies for calculating KM ROI that account for both tangible benefits (time savings, cost avoidance) and intangible benefits (improved morale, enhanced collaboration, risk mitigation).</p></li><li><p><strong>Measurement instrument development</strong>: Creation of validated tools for assessing KM capability maturity in healthcare settings, enabling benchmarking and progress tracking.</p></li><li><p><strong>Knowledge asset valuation methods</strong>: Approaches for quantifying the value of specific knowledge assets (clinical protocols, decision support rules, process documentation) using methods analogous to intellectual property valuation.</p></li></ol><h2><strong>Conclusion</strong></h2><p>The value of knowledge management in healthcare extends far beyond archival duties. As demonstrated by our work at Partners HealthCare, emerging requirements for health AI registries, and cross-sector research from the National Academies, KM provides essential infrastructure for complex health systems to operate safely and efficiently. By systematically capturing workforce expertise, rigorously cataloging and curating technological assets, and establishing governance frameworks for knowledge sharing, healthcare organizations can ensure continuity of care, enhance patient safety through transparency, and build the resilience necessary to navigate a rapidly evolving medical landscape.</p><p>Yet the persistent challenge remains: developing robust business cases that quantify KM&#8217;s value in terms that resonate with healthcare executives and policymakers. The National Academies&#8217; finding that no organization has mastered this challenge underscores both the difficulty and the importance of this work. Healthcare KM practitioners must move beyond anecdotal success stories to develop rigorous measurement frameworks that demonstrate return on investment, link KM activities to strategic organizational goals, and justify sustained resource allocation.</p><p>The path forward requires both technical innovation and cultural transformation. Healthcare organizations must invest in KM infrastructure&#8212;standardized information models, measurement systems, governance frameworks&#8212;while simultaneously cultivating knowledge-sharing cultures where contribution is valued and recognized. Only through this dual approach can healthcare realize the full strategic value of knowledge management: not merely as a risk mitigation tool, but as a driver of innovation, efficiency, and excellence in patient care.</p><div><hr></div><h2><strong>References</strong></h2><ol><li><p>National Academies of Sciences, Engineering, and Medicine. Assessing and measuring the business value of knowledge management. NCHRP Web-Only Document 437. Washington, DC: The National Academies Press; 2026. https://doi.org/10.17226/29279</p></li><li><p>Transportation Research Board. The business case for knowledge management: a guide. NCHRP Research Report 1164. Washington, DC: National Academies Press; 2026.</p></li><li><p>Wright A, Goldberg H, Hongsermeier T, Middleton B. A description and functional taxonomy of rule-based decision support content at a large integrated delivery network. J Am Med Inform Assoc. 2007;14(4):489-96.</p></li><li><p>Zhou L, Karipineni N, Lewis J, Maviglia SM, Fairbanks A, Hongsermeier T, et al. A study of diverse clinical decision support rule authoring environments and requirements for integration. BMC Med Inform Decis Mak. 2012;12:128.</p></li><li><p>Hongsermeier T, Kashyap V, Sordo M. Case study: Knowledge management infrastructure: Evolution at Partners Healthcare system in medical decision support. In: Computer-based approaches to improving healthcare quality and safety. Amsterdam: Elsevier; 2006.</p></li><li><p>Hongsermeier T, Maviglia S, Tsurikova L, Bogaty D, Rocha RA, Goldberg H, et al. A legal framework to enable sharing of Clinical Decision Support knowledge and services across institutional boundaries. AMIA Annu Symp Proc. 2011;2011:925-33.</p></li><li><p>Sordo M, Hongsermeier T, Kashyap V, Greenes RA. Partners Healthcare Order Set Schema: An information model for management of clinical content. In: Yoshida H, Jain A, Ichalkaranje A, Jain LC, Ichalkaranje N, editors. Advanced computational intelligence paradigms in healthcare&#8212;1. Studies in computational intelligence, vol 48. Berlin: Springer; 2007. p. 1-25.<br><br></p></li></ol><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://bmiddleton1.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Middleton's Musings is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AGI: are we there yet?]]></title><description><![CDATA[The question of whether we have achieved Artificial General Intelligence (AGI) is no longer a matter of speculative futurism; it is a debate centered on the distinction between spectacle and substance.]]></description><link>https://bmiddleton1.substack.com/p/agi-are-we-there-yet</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/agi-are-we-there-yet</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Mon, 02 Feb 2026 23:45:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!eLOD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!eLOD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!eLOD!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!eLOD!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!eLOD!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!eLOD!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!eLOD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png" width="2816" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1536,&quot;width&quot;:2816,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!eLOD!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!eLOD!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!eLOD!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!eLOD!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca66e075-9c1f-43aa-81cc-8533e8219e87_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The question of whether we have achieved Artificial General Intelligence (AGI) is no longer a matter of speculative futurism; it is a debate centered on the distinction between spectacle and substance. As we move through 2026, we find ourselves in a state of "model overhang," where raw computational capabilities have temporarily outpaced our structural ability to integrate them into meaningful, trusted workflows.</p><p>The Case for Functional Arrival</p><p>There is a compelling argument that, by any reasonable metric&#8212;including those originally proposed by Alan Turing&#8212;the long-standing challenge of creating AGI has been effectively solved. If AGI is defined as a software system&#8217;s capacity to perform any intellectual task a human can, contemporary foundation models have crossed the threshold. We have shifted from "AI as a tool" to "AI as an agentic employee," capable of autonomous code modification and complex reasoning traces that were considered impossible only three years ago.</p><p>The Illusion of Understanding</p><p>However, academic scrutiny reveals a persistent "illusion of thinking." While models exhibit sophisticated conversational abilities, they remain prone to sycophancy&#8212;the tendency to mirror user biases or provide overly flattering feedback at the expense of objective accuracy. This suggests that what we perceive as "general intelligence" may often be a highly advanced form of pattern matching that lacks a robust, stable world model.</p><p>Furthermore, the "brittleness" of these systems is evident in their response to adversarial transformations. A model may solve a complex medical board question yet fail when the prompt is slightly altered or when key context is removed. This gap between performance and true comprehension is the primary barrier to establishing deep epistemological trust.</p><p>Diffusion and the Physicality of Intelligence</p><p>The true impact of AI is materialized not through benchmarking wins, but through diffusion&#8212;the slow, iterative translation of capabilities into productive sectors of the economy. In fields such as clinical informatics, the friction is rarely about the model&#8217;s theoretical power; it is about accountability, transparency, and the physical reality of computation.</p><p>We must view AI as "bicycles for the mind"&#8212;a scaffolding for human potential rather than a substitute for human judgment. In the clinical environment, for instance, trust is not a static quality to be transferred from human to machine. It is a relationship to be redefined through rigorous evaluation of whether AI creates measurable, value-based benefits for patient outcomes.</p><p>The Verdict: A Narrative Crossroads</p><p>Are we there yet? We have arrived at the spectacle of AGI&#8212;the phase of cognitive amplification where systems can mimic human thought with startling fidelity. Yet, we are still in the opening miles of a marathon toward the substance of AGI&#8212;grounded, robust, and physically capable intelligence.</p><p>The journey from agriculture to AI has been an evolution of neural choice. We are now building futures that seem like science fiction until the day they become routine. The challenge of the next decade will not be the emergence of a mythical superintelligence, but the responsible integration of these tools to enhance human productivity and well-being.</p><p>#AGI #ArtificialIntelligence #MachineLearning #ClinicalInformatics #DigitalHealth #AIReadiness #TechEthics #FutureOfWork #HealthIT #InnovationDiffusion</p><p></p>]]></content:encoded></item><item><title><![CDATA[The $70 Billion Paradox: Why Clinical AI’s Evidence Gap Should Keep You Up at Night]]></title><description><![CDATA[*What forty years in medical informatics has taught me about the dangerous disconnect between commercial enthusiasm and clinical reality*]]></description><link>https://bmiddleton1.substack.com/p/the-70-billion-paradox-why-clinical</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-70-billion-paradox-why-clinical</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Tue, 20 Jan 2026 17:14:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!J1Vo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!J1Vo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!J1Vo!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!J1Vo!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!J1Vo!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!J1Vo!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!J1Vo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!J1Vo!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!J1Vo!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!J1Vo!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!J1Vo!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd041a20-d129-4a28-8c3b-e66d07290c88_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>*What forty years in medical informatics has taught me about the dangerous disconnect between commercial enthusiasm and clinical reality*</p><p>There&#8217;s an old observation, often attributed to Lenin, that &#8220;there are decades where nothing happens; and there are weeks when decades happen.&#8221; I&#8217;ve been working in health information technology since the 1980s&#8212;long enough to have seen the first electronic health records, the first commercial personal health record, and more than a few technology hype cycles come and go. And I can tell you: we are living through one of those compressed moments right now.</p><p>The State of Clinical AI Report (2026) landed on my desk recently, and it crystallized something I&#8217;ve been sensing for the past year. We are not merely experiencing incremental progress in AI-assisted healthcare. We are watching the entire foundation of clinical decision-making get interrogated, challenged, and&#8212;in some settings&#8212;replaced by systems we do not yet fully understand.</p><p>This should thrill us. It should also terrify us.</p><p>The Numbers That Don&#8217;t Add Up</p><p>Let me give you the headline figures: a $70 billion market, over 1,200 FDA-cleared AI/ML tools, and 350,000+ consumer health apps flooding the market. These are not projections or aspirational targets. This is the current state of play.</p><p>Now let me give you the other numbers&#8212;the ones that don&#8217;t make it into the investor presentations.</p><p>Fewer than half of the FDA device summaries for these tools report their study design. More than half omit sample size entirely. Less than one percent report patient outcomes. Ninety-five percent lack demographic data. Ninety-one percent include no bias assessment whatsoever.</p><p>Read those sentences again. We have built a multi-billion-dollar industry on a foundation where the vast majority of approved tools have never demonstrated they actually help patients, and we have almost no systematic understanding of whether they work equitably across populations.</p><p>I&#8217;ve spent my career building the infrastructure for evidence-based medicine&#8212;clinical decision support systems, knowledge bases, evaluation frameworks. I helped create the Center for IT Leadership and co-led the Clinical Decision Support Consortium precisely because I believed technology could make medicine better, safer, more consistent. I still believe that. But I also know what happens when we deploy systems we don&#8217;t understand into environments where the stakes are human lives.</p><p>&nbsp;The Equivalency Trap</p><p>Here&#8217;s a detail from the report that deserves more attention than it typically receives: over 95% of FDA-cleared AI/ML medical devices used the 510(k) pathway. For those outside the regulatory world, this means they were approved by demonstrating &#8220;substantial equivalence&#8221; to a device already on the market.</p><p>On its face, this seems reasonable. If your new stethoscope is basically like an old stethoscope, why reinvent the regulatory wheel? But apply this logic to AI systems and you get something approaching absurdity. These tools are being cleared by comparison to other tools that were themselves often approved on thin evidence. It&#8217;s equivalency all the way down.</p><p>The 510(k) pathway was designed for incremental modifications to established technologies. It was not designed for systems that can generate novel clinical recommendations, that learn and change over time, that operate through mechanisms even their creators sometimes cannot fully explain. We are using a regulatory framework designed for the stethoscope era to govern technology that may soon be making autonomous treatment recommendations.</p><p>The Seductive Promise of Superhuman Performance</p><p>I don&#8217;t want to be the curmudgeon who dismisses genuine progress. The technical achievements documented in this report are remarkable. Frontier models like o1-preview are now consistently outperforming physicians on diagnostic reasoning tasks. In controlled, text-based environments, the evidence increasingly suggests these systems have surpassed human capability for certain types of clinical reasoning.</p><p>This is not marketing hype. Studies show these models excelling at management tasks and diagnosing real emergency cases&#8212;not synthetic problems, but actual patients who presented to emergency departments and required admission.</p><p>If you&#8217;d told me five years ago that we&#8217;d reach this point by 2026, I would have been skeptical. The progress has been faster than most of us anticipated.</p><p>But&#8212;and this is a critical but&#8212;the word &#8220;controlled&#8221; in &#8220;controlled environments&#8221; is doing an enormous amount of work in that sentence.</p><p>&nbsp;Where the Magic Breaks Down</p><p>The same report that documents superhuman performance also documents something far more troubling: these systems become unreliable precisely when clinical medicine becomes hard.</p><p>When data is complete and the problem is well-defined, frontier LLMs shine. When faced with uncertainty, missing information, or changing clinical context, they break down. This is not a minor limitation. Uncertainty, incomplete information, and evolving clinical pictures describe the majority of real clinical encounters. Medicine is not a multiple-choice exam. It&#8217;s a conversation with an anxious patient who can&#8217;t quite describe their symptoms, conducted while managing three other urgent cases, with half the information you need, under time pressure.</p><p>The models also exhibit what researchers politely call &#8220;miscalibrated confidence.&#8221; In plain language: they&#8217;re often wrong but rarely in doubt. Human experts naturally moderate their confidence when situations become ambiguous. AI systems frequently do the opposite&#8212;asserting certainty even when their reasoning has gone off the rails.</p><p>Here&#8217;s a finding that should give pause to anyone thinking about autonomous AI deployment. Researchers introduced &#8220;none of the other answers&#8221; (NOTA) testing, where they changed the expected pattern of multiple-choice responses. Model accuracy dropped from 81% to 43%. This suggests that much of what looks like clinical reasoning is actually sophisticated pattern recognition&#8212;the models have learned what the &#8220;right&#8221; answer usually looks like, not how to actually reason through clinical problems.</p><p>And performance degrades significantly when models shift from processing static clinical vignettes to engaging in multi-turn conversation. The format that more closely mimics an actual patient encounter is precisely the format where these tools become less reliable.</p><p>The Evaluation Crisis</p><p>The deeper problem is that we&#8217;ve been measuring the wrong things. For years, we&#8217;ve evaluated clinical AI primarily by testing whether it can pass medical licensing exams. These models now score near-perfectly on such tests. The benchmarks are &#8220;saturated&#8221;&#8212;they can no longer discriminate between systems.</p><p>More fundamentally, these benchmarks never measured what actually matters for clinical practice. They tested medical knowledge. They did not test administrative task performance, use of real patient data, conversational ability, bias, fairness, or&#8212;most critically&#8212;clinical safety.</p><p>The good news is that better evaluation frameworks are emerging. HealthBench evaluates AI in realistic, open-ended health conversations. MedHELM and MedAgentBench assess performance on everyday workflow tasks. NOHARM directly quantifies how often LLM recommendations could cause patient harm.</p><p>Early results from these frameworks are sobering. NOHARM demonstrates that standard &#8220;knowledge&#8221; benchmarks do not reliably predict clinical safety. A system that aces medical exams may still generate recommendations that could harm patients.</p><p>Multimodal Promise, Multi-Agent Paradox</p><p>The most exciting technical frontiers involve multimodal AI that integrates text, images, and other clinical data, and multi-agent systems where multiple specialized AI agents collaborate on complex tasks.</p><p>The results can be impressive. In one randomized controlled trial, ophthalmologists using a multimodal AI copilot called EyeFM improved their correct diagnosis rate from 75% to 92%. Microsoft&#8217;s MAI DxO, which simulates a panel of five AI physicians, achieved 80% accuracy on complex diagnostic cases&#8212;outperforming both human physicians and standalone LLMs.</p><p>But here&#8217;s a finding that every health system executive should understand: multi-agent systems exhibit what researchers call the &#8220;Multi-Agent Optimization Paradox.&#8221; In one study, a system composed of the individually strongest AI agents achieved only 68% diagnostic accuracy, while a less-optimized combination reached 77%. Optimizing each component doesn&#8217;t optimize the whole. Information flow between agents matters as much as individual agent capability.</p><p>This has profound implications for system validation. You cannot approve a multi-agent system by validating its components separately. End-to-end testing in realistic conditions is not optional&#8212;it&#8217;s the only evaluation that matters.</p><p>What This Means for Leaders</p><p>I&#8217;ve been advising healthcare organizations on technology adoption for decades. Based on this report and my own experience, here&#8217;s how I&#8217;d frame the strategic situation:</p><p>The central challenge is that AI capabilities are advancing faster than our ability to evaluate them, regulate them, or integrate them safely into clinical workflows. This creates both opportunity and risk in equal measure.</p><p>For near-term adoption, narrow beats broad. The most immediate value will come from AI tools tightly scoped to specific clinical domains and contexts. These systems are easier to validate, integrate, and explain to clinicians. The dream of a generalist AI physician is compelling, but the reality of that dream is still years away&#8212;and may remain perpetually over the horizon.</p><p>For deployment strategy, safety infrastructure is not optional. This is especially true for patient-facing applications. The technology itself is advancing rapidly. What&#8217;s not advancing rapidly enough is the guardrails, oversight systems, and feedback mechanisms needed to deploy these tools safely. Investment in safety infrastructure must proceed in parallel with investment in AI capability.</p><p>For evidence standards, demand more. Stop accepting vendor claims based on multiple-choice benchmark performance. Ask about conversational degradation. Ask about NOTA testing. Ask about safety evaluations. Ask whether the system has been tested in conditions that actually resemble your clinical environment.</p><p>For clinician buy-in, focus on what they actually want. There&#8217;s a massive opportunity in developing AI that reduces administrative burden rather than augmenting diagnostic reasoning. Documentation, prior authorization, inbox management, care coordination&#8212;these are the tasks that burn out clinicians, and they&#8217;re currently underrepresented in both research and commercial products. If you want clinician adoption, give them back their time before you try to change how they think.</p><p>The Longer View</p><p>I&#8217;ve watched health IT evolve from mainframes to microcomputers to the internet to mobile to cloud to AI. Each transition brought genuine benefits and legitimate hype. Each required us to separate what the technology could theoretically do from what it could do reliably in real clinical settings.</p><p>The current moment is different in degree, not in kind. The AI systems being deployed today are more powerful than anything we&#8217;ve seen before. The gap between controlled-environment performance and real-world readiness is also larger than anything we&#8217;ve seen before. The stakes&#8212;measured in patient safety, healthcare economics, and the future of clinical practice&#8212;are correspondingly enormous.</p><p>We are at a fork in the road. One path leads to AI that genuinely augments clinical care, carefully validated and thoughtfully integrated, making medicine better for patients and practitioners alike. The other path leads to premature deployment, preventable harm, loss of clinical trust, and a regulatory backlash that could set the field back by years.</p><p>The data in this report suggests we&#8217;re currently walking both paths simultaneously. The question is which one we choose to commit to.</p><p>I know which path I&#8217;m hoping for. After forty years in this field, I still believe technology can make healthcare better. But that belief comes with a hard-earned corollary: only if we demand the evidence to prove it.</p><p>----</p><p>*Blackford Middleton is a Distinguished Fellow of the American College of Medical Informatics and a semi-retired consultant in healthcare technology and clinical informatics. He serves on several advisory boards including BMJ Knowledge Centre and Evidentli,, and maintains active research collaborations in clinical AI.&nbsp;</p>]]></content:encoded></item><item><title><![CDATA[The Next Scientific Revolution: How Foundation Models Are Reshaping the Department of Energy]]></title><description><![CDATA[When we think of Artificial Intelligence today, we often think of chatbots writing emails or generating images.]]></description><link>https://bmiddleton1.substack.com/p/the-next-scientific-revolution-how</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-next-scientific-revolution-how</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Sun, 18 Jan 2026 00:54:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KMFe!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc54ae8cf-4907-48fa-ad2c-62bd601a1ceb_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When we think of Artificial Intelligence today, we often think of chatbots writing emails or generating images. But a new consensus study report from the National Academies of Sciences, Engineering, and Medicine suggests we are on the brink of something much bigger: a paradigm shift in how science itself is conducted.</p><p>The report, titled <em><a href="/__u/www.google.com/url?sa=E&amp;q=https%3A%2F%2Fdoi.org%2F10.17226%2F29212">Foundation Models for Scientific Discovery and Innovation</a></em>, outlines how the Department of Energy (DOE) is uniquely positioned to leverage these powerful AI tools to solve some of the world's most complex problems&#8212;from nuclear fusion to climate change.</p><h4>What Are Scientific Foundation Models?</h4><p>Unlike traditional AI models designed for a single task, <strong>foundation models</strong> are massive neural networks trained on vast amounts of data. They can be adapted to many different tasks and can reason across different "modalities"&#8212;understanding not just text, but images, sensor data, and complex physics simulations.</p><p>The report highlights that these models have "emergent capabilities," meaning they can identify patterns in data that are invisible to human researchers or traditional computational solvers.</p><h4>The Rise of "Algorithmic Alloys"</h4><p>One of the report&#8217;s most compelling insights is that AI won't replace traditional physics; it will supercharge it. The committee recommends developing <strong>hybrid models</strong>&#8212;what they call "algorithmic alloys".</p><ul><li><p><strong>Traditional Modeling:</strong> Great at adhering to physical laws and providing reliable, interpretable results, but computationally expensive.</p></li><li><p><strong>Foundation Models:</strong> incredibly fast and adaptable, but prone to "hallucinations" and lacking physical grounding.</p></li></ul><p>By fusing these two, scientists can model complex systems (like turbulent fluid flow or material failure) with the speed of AI and the reliability of physics.</p><h4>Why the DOE?</h4><p>While tech giants like Google and Microsoft are pouring billions into AI, the report argues the DOE has "strategic advantages" that private industry cannot match:</p><ol><li><p><strong>Unique Data:</strong> Access to massive, often classified, datasets from nuclear stockpiles and particle accelerators.</p></li><li><p><strong>World-Class Facilities:</strong> Stewardship of experimental facilities and supercomputers like Frontier and Aurora.</p></li><li><p><strong>Mission-Driven Workforce:</strong> A deep bench of scientists tackling long-term, high-risk problems that don't have immediate commercial payoff.</p></li></ol><h4>The Vision: Self-Driving Laboratories</h4><p>The report envisions a future of <strong>Agentic AI</strong>&#8212;systems that don't just answer questions but take action. Imagine a "self-driving laboratory" where an AI agent formulates a hypothesis, designs an experiment, commands a robot to mix materials, analyzes the results, and refines its hypothesis&#8212;all in a continuous loop.</p><p>Already, systems like "The AI Scientist-v2" have demonstrated the ability to autonomously author scientific manuscripts and execute experiments.</p><h4>The Challenges: Trust and Security</h4><p>Despite the excitement, the report sounds a note of caution. Foundation models are currently "black boxes" that lack the verification and validation standards required for high-stakes science.</p><ul><li><p><strong>Hallucinations:</strong> In nuclear reactor safety or grid security, a plausible-sounding but incorrect answer could be catastrophic. The report calls for new frameworks for <strong>Verification, Validation, and Uncertainty Quantification (VVUQ)</strong> specifically for AI.</p></li><li><p><strong>Security:</strong> These models are vulnerable to "jailbreaking," data poisoning, and adversarial attacks. The DOE must treat AI security as a top priority.</p></li><li><p><strong>Human in the Loop:</strong> The committee concludes that while AI can handle routine tasks, human oversight remains essential for judgment, safety, and accountability.</p></li></ul><h4>The Verdict</h4><p>The integration of foundation models into the scientific enterprise represents a "paradigm shift". To stay competitive, the DOE must modernize its data infrastructure, partner strategically with industry, and train a workforce capable of bridging the gap between deep science and deep learning.</p><div><hr></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Revenge of the Idea: How AI Shifts Value from Execution to Judgment]]></title><description><![CDATA[A venture capitalist I know recently reviewed 47 pitch decks in a single week.]]></description><link>https://bmiddleton1.substack.com/p/the-revenge-of-the-idea-how-ai-shifts</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-revenge-of-the-idea-how-ai-shifts</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Sun, 18 Jan 2026 00:14:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7I1v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!7I1v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!7I1v!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!7I1v!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!7I1v!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7I1v!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!7I1v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!7I1v!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!7I1v!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!7I1v!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7I1v!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F052ef64d-7ba4-4638-be2f-3451fa182e69_1024x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>A venture capitalist I know recently reviewed 47 pitch decks in a single week. "Thirty years ago, if a founder walked in with a working prototype, that was the signal," she told me. "It meant they could execute. Now? Every pitch includes a working demo because anyone can build one in a weekend with Cursor and Claude. Last Tuesday, I saw three different 'AI-powered project management tools' that were functionally identical. They all worked. They all looked professional. They were all completely undifferentiated."</p><p>She paused, then added: "The only question I ask now is: Do you understand something about this problem that others don't? Because if your competitive advantage is that you can build software, you have no competitive advantage."</p><p>This is the structural shift that business leaders are missing. For decades, the mantra has been "ideas are cheap, execution is everything." Thomas Edison championed "perspiration" over "inspiration." Steve Jobs and investors like John Doerr historically viewed ideas as commodities and execution as the only defensible moat. That was economically rational when the cost of doing&#8212;coding, content creation, data analysis&#8212;was the primary bottleneck to creating value.</p><p>But we're entering a new regime: AI is dramatically reducing the cost of routine execution, which paradoxically makes human judgment more valuable, not less. When the machines can handle the "doing," the question of <em>what</em> to do becomes the ultimate leverage.</p><p><strong>The Economics of Cheap Execution</strong></p><p>To understand this shift, we need to revisit basic economics. When the cost of a key input drops dramatically, the economic value doesn't disappear&#8212;it migrates to the input's complements.</p><p>Economists Ajay Agrawal, Joshua Gans, and Avi Goldfarb have documented this pattern with prediction technologies. When prediction becomes cheap (as AI has made it), the value of judgment&#8212;deciding what to do with predictions&#8212;rises proportionally. The same logic applies to execution itself.</p><p>Consider what's happened to software development. GitHub Copilot and similar tools now autocomplete entire functions from natural language descriptions. What once required hours of coding can happen in minutes. But this doesn't eliminate the need for engineers&#8212;it shifts their value from writing syntax to making architectural decisions: What system should we build? How should components interact? What are the security implications? What technical debt are we willing to accept?</p><p>The execution&#8212;translating architecture into code&#8212;becomes increasingly automated. The complements&#8212;the judgment required to design good architecture&#8212;become increasingly valuable.</p><p>These complements include:</p><ul><li><p><strong>Proprietary data and context</strong> that general models lack</p></li><li><p><strong>Domain expertise</strong> that can't be web-scraped</p></li><li><p><strong>Cultural and ethical judgment</strong> about what should be built</p></li><li><p><strong>Trust relationships</strong> that validate judgment quality</p></li></ul><p>Let me show you how this plays out across three strategic domains.</p><p><strong>I. From "Build to Learn" to "Model to Validate"</strong></p><p>The lean startup methodology was built for a world where market data was expensive to gather. You built a Minimum Viable Product, tested it with real users, gathered data, and iterated. This process could take months.</p><p>AI is changing the calculus. Higher-fidelity modeling <em>before</em> building allows companies to filter bad ideas faster, reserving capital and human effort for the most promising concepts.</p><p>In pharmaceutical development, this shift is already visible. Companies like Recursion Pharmaceuticals and Insilico Medicine use AI to model millions of drug-target interactions <em>in silico</em> before synthesizing a single molecule. What once required extensive wet lab work can now be simulated computationally. This doesn't replace clinical trials&#8212;but it dramatically improves the hit rate of compounds that enter trials.</p><p>The "idea"&#8212;which molecular target to pursue, which disease mechanism to exploit&#8212;requires deep biological insight. The execution&#8212;screening millions of combinations&#8212;is now largely automated.</p><p>In healthcare delivery, predictive models now help health systems simulate intervention effectiveness before deployment. Which patients should receive a diabetes prevention program? Which care coordination approach will reduce readmissions? AI can model different scenarios using historical data, allowing leaders to test their "ideas" about care delivery before committing resources.</p><p>But here's the critical point: <strong>these models are only as good as the questions you ask them to answer</strong>. A poorly conceived clinical trial, optimized by AI, is still a poorly conceived trial. The judgment about what hypothesis is worth testing&#8212;informed by clinical experience, biological plausibility, and patient needs&#8212;is where the value lies.</p><p>The execution of running the simulation is now cheap. The idea of what to simulate is not.</p><p><strong>II. From Creation to Curation</strong></p><p>When AI can generate infinite content, "more" becomes worthless. The value of "right" becomes infinite.</p><p>Consider the modern marketing leader's role. A decade ago, their value was in producing compelling copy. Today, an AI can generate a thousand email subject lines in seconds, each grammatically correct and incorporating proven persuasion techniques. The marketing leader's value has shifted to curation&#8212;possessing the taste and cultural fluency to select the one message that resonates with <em>this</em> audience given current events, competitive positioning, and brand voice.</p><p>The AI executes the creation. The human curates based on judgment.</p><p>This pattern extends to clinical decision-making. AI-powered clinical decision support systems can generate comprehensive differential diagnoses almost instantly. For a patient presenting with fatigue, chest pain, and shortness of breath, the AI might suggest fifteen evidence-based possibilities, each with supporting literature.</p><p>But the clinician's value is curation: knowing which diagnosis to pursue <em>first</em> given this patient's age, medical history, social determinants of health, and expressed concerns. Should we rule out the life-threatening cardiac causes immediately? Or is this likely anxiety in a young, healthy patient with recent job stress, making extensive cardiac workup low-value care?</p><p>The AI handles the execution of searching medical literature and applying diagnostic criteria. The clinician provides the judgment of which path makes sense for this human being.</p><p>This curation requires something AI fundamentally lacks: <strong>context that isn't in the training data</strong>. What are the patient's transportation barriers? Their health literacy? Their cultural beliefs about medical intervention? Their insurance coverage? Their tolerance for uncertainty?</p><p>These contextual factors don't live in medical journals. They live in the relationship between clinician and patient, in knowledge of local resources and constraints, in experience with hundreds of similar presentations. This contextual knowledge <em>is</em> the idea. The execution of applying clinical guidelines is increasingly automated.</p><p><strong>III. Trust and Transparency as Competitive Moats</strong></p><p>There's considerable excitement about "agentic commerce"&#8212;AI agents that will make purchasing decisions on behalf of humans. Imagine telling your AI assistant, "Buy me the best eco-friendly running shoes under $150," and having it research options, compare specs, negotiate prices, and execute the purchase.</p><p>But this scenario reveals AI's fundamental limitation. "Best" requires a value framework the AI doesn't possess. Best for durability? Comfort? Carbon footprint? Manufacturing labor practices? Performance for a neutral gait versus overpronation? The agent can execute the transaction, but the judgment about what "best" means requires human input.</p><p>As AI agents increasingly mediate routine transactions, differentiation will come from three sources:</p><p><strong>First, proprietary data.</strong> If an AI agent is comparison shopping, your competitive advantage can't rely on information the agent can easily access. You need data the general models lack. In healthcare, this might be longitudinal outcomes data showing your intervention actually works in real-world settings, not just controlled trials. In retail, it might be detailed supply chain transparency that competitors can't match.</p><p><strong>Second, verifiable human expertise in high-stakes decisions.</strong> As AI handles routine execution, human oversight becomes a <em>feature</em>, not a cost to minimize. Consider cancer treatment planning. AI can draft treatment protocols based on tumor genomics, patient characteristics, and current evidence. But the value proposition increasingly is: "This plan was reviewed by a medical oncologist with 20 years of melanoma experience who incorporated your preferences about quality of life versus survival time."</p><p>Human-in-the-loop shifts from quality control to premium offering. The execution is automated, but judgment is the product.</p><p><strong>Third, transparent judgment processes.</strong> In an era where AI can execute flawlessly but without genuine understanding, showing <em>how</em> and <em>why</em> decisions were made becomes competitively valuable. This is especially true in regulated industries like healthcare and finance, where accountability matters.</p><p>I've watched this play out in clinical decision support. Systems that just provide recommendations&#8212;even excellent ones&#8212;face resistance. Systems that show the reasoning, cite the evidence, and make clear where clinical judgment was applied? Those get adopted. The transparency of judgment is part of the value.</p><p><strong>What This Means for Leaders</strong></p><p>If you're leading an organization, this shift demands honest assessment. What percentage of your team's time goes to execution that AI can now handle? What percentage goes to judgment that requires your organization's proprietary context?</p><p>More importantly: <strong>Are you measuring and rewarding the right activities?</strong></p><p>If your performance metrics still emphasize execution speed&#8212;lines of code written, content pieces produced, transactions processed&#8212;you're optimizing for the thing AI is making cheap. You need to shift toward metrics that capture judgment quality: decision accuracy with limited data, synthesis of disparate information sources, identification of novel approaches that AI wouldn't generate.</p><p>Three specific investments become critical:</p><p><strong>First, proprietary data infrastructure.</strong> What unique datasets fuel better judgment in your domain? For healthcare organizations, this might be longitudinal patient outcomes linked to interventions. For retailers, detailed customer preference data beyond purchase history. The key is data that isn't in AI's training sets and can't be easily replicated.</p><p><strong>Second, deep domain expertise development.</strong> AI models are trained on text scraped from the internet&#8212;a biased sample that's missing proprietary knowledge, tacit expertise, and experiential learning. Invest in building and retaining expertise that can't be web-scraped. In healthcare, this is the physician who has managed 500 cases of a rare disease. In manufacturing, the engineer who understands why the process fails under specific conditions that aren't in the manual.</p><p><strong>Third, transparent judgment frameworks.</strong> Can you articulate <em>why</em> decisions were made? Can you show the reasoning chain? As AI handles execution, your competitive moat is judgment quality, and judgment quality requires verifiability. Build systems that make reasoning visible.</p><p><strong>The Pendulum Swings, But Doesn't Reverse</strong></p><p>The "ideas are cheap" era wasn't wrong&#8212;it was right for its time. When execution was scarce and expensive, rewarding the grind made economic sense. The teams that could ship faster, optimize harder, and execute more efficiently won.</p><p>But playing by those rules in the AI era misses the structural shift.</p><p>This doesn't mean execution disappears. Someone still has to integrate the AI tools, validate their outputs, deploy solutions, and measure results. But the <em>nature</em> of valuable execution changes: from writing code syntax to designing system architecture, from generating content to curating with context, from following protocols to exercising judgment under uncertainty.</p><p>The companies that thrive won't be those who can execute fastest&#8212;AI will help everyone execute fast. They'll be those whose ideas are informed by proprietary data, refined by deep domain expertise, and validated by transparent human judgment.</p><p>The machines can optimize the path. Only you can choose the destination. And in a world where execution becomes increasingly automated, choosing the right destination is everything.</p><p></p><p></p><p></p><p><strong>Appendices: Alternative Opening Vignettes</strong></p><p><strong>Appendix A: The Genomics Paradox</strong></p><p>A precision oncology researcher recently showed me something unsettling. Her team had uploaded a patient's tumor genomic profile to three different AI-powered treatment recommendation systems. All three generated comprehensive treatment plans within minutes&#8212;something that would have required weeks of literature review just five years ago.</p><p>"Look at this," she said, pointing to her screen. "System A recommends immunotherapy based on tumor mutation burden. System B suggests targeted therapy based on a specific EGFR mutation. System C proposes a clinical trial for a novel combination. All three cite high-quality evidence. All three are clinically plausible. And they're contradictory."</p><p>The execution&#8212;analyzing genomic data, searching thousands of research papers, matching patients to trials&#8212;has become nearly free. But the judgment about which path makes sense for this 67-year-old woman with comorbidities who lives two hours from the nearest cancer center and is terrified of the side effects her sister experienced? That requires something AI can't provide.</p><p>This is the structural shift that business leaders are missing...</p><p><strong>Appendix B: The Strategy Consultant's Dilemma</strong></p><p>I recently watched a mid-level consultant at a prestigious firm present a market entry strategy to a client's C-suite. The analysis was flawless&#8212;comprehensive competitive landscape, detailed financial projections, risk matrices, implementation timelines. The PowerPoint was pristine. The recommendations were thoroughly researched.</p><p>The CEO listened politely, then asked: "Did you use AI to build this?"</p><p>"We used AI to accelerate the analysis, yes," the consultant replied carefully.</p><p>"I figured," the CEO said. "Because I had my chief strategy officer run the same question through Claude and got 80% of this deck in an hour. What I'm paying you for isn't the analysis anymore. What I need to know is: What do you believe we should do that the AI wouldn't recommend? Where is your firm's judgment different from the machine's?"</p><p>The room went silent. The consultant had no answer. The "idea"&#8212;a contrarian insight based on experience with similar transformations, knowledge of this industry's unique culture, understanding of this CEO's risk tolerance&#8212;was missing. The execution was perfect and worthless.</p><p>This is the structural shift that business leaders are missing...</p><p></p><p></p><p></p><p><strong>Works Cited</strong></p><ol><li><p>Agrawal, A., Gans, J., &amp; Goldfarb, A. (2018). <em>Prediction Machines: The Simple Economics of Artificial Intelligence</em>. Harvard Business Review Press.</p></li><li><p>NBER Working Paper: "Exploring the Impact of Artificial Intelligence: Prediction versus Judgment" - https://www.nber.org/system/files/working_papers/w24626/w24626.pdf</p></li><li><p>IBM Insights: "Proprietary data, your competitive edge in generative AI" - https://www.ibm.com/think/insights/proprietary-data-gen-ai-competitive-edge</p></li><li><p>McKinsey: "The agentic commerce opportunity: How AI agents are ushering in a new era for consumers and merchants" - https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-agentic-commerce-opportunity-how-ai-agents-are-ushering-in-a-new-era-for-consumers-and-merchants</p></li><li><p>Built In: "How Human-in-the-Loop Is Evolving with AI Agents" - https://builtin.com/articles/human-in-the-loop-evolution</p></li><li><p>Forbes: "The Secret To Successful Enterprise AI? 'Human-In-The-Loop' Design" - https://www.forbes.com/councils/forbestechcouncil/2024/08/06/the-secret-to-successful-enterprise-ai-human-in-the-loop-design/</p></li></ol><p></p>]]></content:encoded></item><item><title><![CDATA[The "MS-DOS Era" of AI: 5 Counter-Intuitive Truths Redefining 2026]]></title><description><![CDATA[We were promised a world where AI would do our jobs while we sipped lattes.]]></description><link>https://bmiddleton1.substack.com/p/the-ms-dos-era-of-ai-5-counter-intuitive</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-ms-dos-era-of-ai-5-counter-intuitive</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Fri, 09 Jan 2026 16:00:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Cwwe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Cwwe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Cwwe!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Cwwe!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Cwwe!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Cwwe!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Cwwe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png" width="1536" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1024,&quot;width&quot;:1536,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Cwwe!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!Cwwe!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!Cwwe!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Cwwe!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc75c67c3-1824-45b1-a061-3e821e400754_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Remember when?</figcaption></figure></div><p>We were promised a world where AI would do our jobs while we sipped lattes. Instead, in early 2026, many of us find ourselves in a strange paradox: we have more "intelligence" at our fingertips than ever before, yet we&#8217;re working harder just to keep it from making mistakes. We are living through what experts call the "MS-DOS era" of artificial intelligence&#8212;a clunky, transitionary period where the old rules of software are breaking, but the new ones haven't quite settled.</p><p>If you feel like the AI revolution is getting... weird, you aren&#8217;t alone. Here are the five most surprising and counter-intuitive takeaways from the current state of the AI frontier.</p><p></p><p>1. The Death of the "Seat": Software is Becoming a Service (Literally)</p><p></p><p>For decades, the tech economy was built on the "SaaS" model: you paid for a "seat" for every human using a tool. But as AI agents begin to do the actual work rather than just helping humans do it faster, the math is falling apart. We are seeing a shift from "Software-as-a-Service" to "Service-as-Software."</p><p>In this new world, companies aren't buying licenses for ten lawyers; they&#8217;re paying for 500 contracts to be reviewed. The focus has shifted from productivity tools to outcome-based labor.</p><p></p><p>"I know you only have three lawyers, but my agents could do the amount of work of basically unlimited lawyers." &#8212; Aaron Levie, CEO of Box</p><p></p><p>2. The Scaling Wall: Bigger Isn&#8217;t Getting Smarter</p><p></p><p>For three years, the mantra was "more data, more compute, more intelligence." But 2026 has brought a sobering reality: we are hitting a wall. While massive models are getting better at memorizing facts (crystallized intelligence), they are struggling to improve at novel problem-solving (fluid intelligence).</p><p></p><p>Recent benchmarks show that while compute power has increased 50,000x since 2019, AI accuracy on complex reasoning tasks has only nudged up marginally. The future isn't in "bigger" models, but in "Test-Time Adaptation"&#8212;systems that can "think" and learn in real-time rather than just reciting what they learned in training.</p><p></p><p>3. The "Forever Trainee" Paradox</p><p></p><p>We often talk about AI agents as autonomous digital workers. In reality, they are behaving more like "brilliant but amnesiac junior staffers." They work quickly and confidently, but often incorrectly, requiring a constant "human-in-the-loop" to clean up the mess.</p><p></p><p>The counter-intuitive result? Organizations that deployed agents to reduce workload often find themselves creating new layers of oversight. Instead of doing the work, managers are now spending their days auditing AI-generated drafts line-by-line to catch "cascading errors."</p><p></p><p>4. The Erosion of Human Confidence</p><p></p><p>You&#8217;d think having an AI assistant would make you feel like a genius. Surprisingly, the opposite is happening in many high-stakes fields. As AI becomes embedded in daily workflows, a phenomenon known as "Jagged Intelligence" is emerging.</p><p></p><p>People are moving faster, but they are becoming less confident in their own judgment. Because the AI provides an answer instantly, humans are losing the "reasoning muscle" required to explain why a decision was made. This creates a hidden risk: "brittle" leadership that can't defend its choices once the tool is taken away.</p><p></p><p>&#8220;Organizations that outsource judgment without reinforcing reasoning will pay for it later&#8212;in brittle decisions and fragile leadership." &#8212; Barry O&#8217;Reilly, Leadership Expert</p><p></p><p>5. The Chatbot is a Distraction</p><p></p><p>If you think the future of AI is a better version of ChatGPT, you&#8217;re still looking at the "1960s of OS design." Current leaders like Andrej Karpathy and Sam Altman suggest that the "chatbot" is just a temporary interface.</p><p></p><p>The real goal is Ambient Intelligence, where the interface "melts away." We are moving toward a world where you don't "prompt" an AI; it proactively navigates your digital life, managing your calendar, your inbox, and your workflows in the background without you ever typing a single command.</p><p></p><p>The Road Ahead: A New Kind of Literacy</p><p></p><p>As we move further into 2026, the most valuable skill isn't knowing how to code or even how to "prompt." It&#8217;s Deterministic Control&#8212;the ability to provide human judgment and clear intent to a system that is brilliant but lacks a moral or logical compass.</p><p></p><p>We are no longer just users of tools; we are becoming the architects of automated systems. The question for all of us this year is: If an AI can do 90% of your task, do you still have the expertise to recognize when the remaining 10% is dangerously wrong?</p>]]></content:encoded></item><item><title><![CDATA[The Great Shift: From Systems of Record to Systems of Action]]></title><description><![CDATA[We are living through the awkward teenage years of enterprise AI.]]></description><link>https://bmiddleton1.substack.com/p/the-great-shift-from-systems-of-record</link><guid isPermaLink="false">https://bmiddleton1.substack.com/p/the-great-shift-from-systems-of-record</guid><dc:creator><![CDATA[Blackford Middleton, MD, MPH]]></dc:creator><pubDate>Fri, 02 Jan 2026 21:48:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8K4A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We are living through the awkward teenage years of enterprise AI.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8K4A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8K4A!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!8K4A!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!8K4A!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8K4A!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_webp, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8K4A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png" width="1536" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1024,&quot;width&quot;:1536,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8K4A!, /__u/bmiddleton1.substack.com/w_424, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!8K4A!, /__u/bmiddleton1.substack.com/w_848, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!8K4A!, /__u/bmiddleton1.substack.com/w_1272, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8K4A!, /__u/bmiddleton1.substack.com/w_1456, /__u/bmiddleton1.substack.com/c_limit, /__u/bmiddleton1.substack.com/f_auto, /__u/bmiddleton1.substack.com/q_auto:good, /__u/bmiddleton1.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e36a6d3-4178-4037-b7ba-50c6c9678d91_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption><figcaption class="image-caption">Look around most organizations today. You&#8217;ll see a sprawling landscape of disconnected "Copilots," bespoke chatbots, and isolated RAG experiments. We have more intelligence at our fingertips than ever before, yet we seem to have less coordination. We've built a digital Tower of Babel&#8212;a thousand brilliant models that cannot speak to one another.</figcaption></figure></div><p>The problem isn't a lack of AI power; it's a lack of AI architecture. We are trying to run a 21st-century workforce on 20th-century infrastructure.</p><p>The next phase of this revolution isn't about building a smarter model. It&#8217;s about building a new kind of operating system&#8212;an AI.OS&#8212;that transforms AI from a passive advisor into the active nervous system of the enterprise.</p><p>It&#8217;s the shift from a System of Record to a System of Action.</p><p>The Passive vs. The Active Enterprise</p><p>For the last two decades, the holy grail of IT was the "System of Record." We spent billions implementing Salesforce, SAP, and ServiceNow to create a single source of truth. These systems are crucial, but they are fundamentally passive. They are vast, immutable databases of what has already happened.</p><p> * System of Record: "The customer filed a complaint ticket at 10:03 AM."</p><p>That&#8217;s useful data. But data doesn't solve problems. Action does.</p><p>Until now, the "action layer" has been human. A human had to read that ticket, understand the context, log into three other systems, coordinate with a manager, and draft a response. The software was just a tool we used to do the work.</p><p>The AI.OS flips this model on its head. It is an intelligence-native layer that sits above your Systems of Record, capable of understanding intent and executing complex workflows across them without constant human hand-holding.</p><p> * System of Action (AI.OS): "I see the customer is frustrated. I have already coordinated with the billing agent to apply a credit, updated the CRM, and drafted a personalized apology for your review."</p><p>The AI.OS doesn't just record the state of the world; it actively changes it to align with your goals.</p><p>The Pillars of a Sovereign AI.OS</p><p>Building this isn't as simple as connecting APIs. A true System of Action requires a fundamentally new architecture built on principles that the current hyperscalers ignore.</p><p>1. Privacy Sovereignty as the Foundation</p><p>You cannot have autonomous action without absolute trust. If your AI has to send your most sensitive data to a third-party cloud to "think," you will never unleash it on mission-critical tasks. In the regulated enterprise, sovereignty isn't a feature; it is the primary performance benchmark. The AI.OS must be local-first, bringing the intelligence to the data, not the other way around.</p><p>2. The Inter-Agent Economy</p><p>No single model will ever be good at everything. The AI.OS is an orchestrator of specialists. It creates an internal marketplace where a general-purpose planner agent can "hire" a specialized logistics agent, a legal review agent, or a code-generation agent to complete a task. This isn't a rigid workflow; it's a dynamic, self-organizing economy of competence.</p><p>3. Radical Vendor Neutrality</p><p>The future is not a walled garden. The AI.OS must be Switzerland. It has to be an independent fabric that allows you to hot-swap models&#8212;moving from GPT-5 to Gemini to a fine-tuned local Llama model&#8212;without rebuilding your entire infrastructure. Your corporate memory and workflows belong to you, not your model provider.</p><p>The New Infrastructure</p><p>We are moving past the era of buying AI as a "tool." We are entering the era of building AI as infrastructure.</p><p>The companies that win in the next decade won't be the ones with the flashiest chatbots. They will be the ones that build a sovereign, active nervous system&#8212;an AI.OS&#8212;that turns their data into decisive, automated action at scale.</p><p>The technology is ready. The question is, are you ready to build it?</p><p></p><p>#SovereignAI</p><p>#AIInfrastructure</p><p>#SystemOfAction</p><p>#AI.OS</p><p>#AgenticAI</p><p>#EnterpriseAI</p><p>#PostHyperscaler</p><p>#InfrastructureOverInterfaces</p><p>#EnterpriseSovereignty</p><p>acknowledgements: Co-produced with AI</p><p></p>]]></content:encoded></item></channel></rss>