<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Mental Health Meets AI]]></title><description><![CDATA[Where clinical judgment and AI systems collide, and decisions have consequences.]]></description><link>https://drscottwallace.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!HCgZ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a0695e9-f438-48cc-bf8b-993ed3750ecf_1254x1254.png</url><title>Mental Health Meets AI</title><link>https://drscottwallace.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 10:29:36 GMT</lastBuildDate><atom:link href="/__u/drscottwallace.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Scott Wallace, PHD]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[drscottwallace@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[drscottwallace@substack.com]]></itunes:email><itunes:name><![CDATA[Scott Wallace, PHD]]></itunes:name></itunes:owner><itunes:author><![CDATA[Scott Wallace, PHD]]></itunes:author><googleplay:owner><![CDATA[drscottwallace@substack.com]]></googleplay:owner><googleplay:email><![CDATA[drscottwallace@substack.com]]></googleplay:email><googleplay:author><![CDATA[Scott Wallace, PHD]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Mental Health AI Safety for Adolescents Cannot Be Tested One Reply at a Time]]></title><description><![CDATA[What Meta&#8217;s youth settlement should teach us about risk that builds over time]]></description><link>https://drscottwallace.substack.com/p/mental-health-ai-safety-for-adolescents</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/mental-health-ai-safety-for-adolescents</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Sun, 30 Aug 2026 13:38:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gvfL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!gvfL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!gvfL!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!gvfL!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!gvfL!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gvfL!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!gvfL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1866278,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/213169795?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!gvfL!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!gvfL!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!gvfL!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gvfL!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17edb45a-d81f-4dd7-95a0-85a8300da9eb_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I have become increasingly uneasy about how we test mental health AI, especially when the <a href="/__u/drscottwallace.substack.com/p/what-teenagers-prefer-is-not-a-chatbot">user is an adolescent</a>.</p><p>Most safety testing still looks at one exchange at a time. Give the system a difficult prompt, see what it says, decide whether the answer crossed a clinical or policy line.</p><p>We need that testing, of course. But it misses something important because <strong>a teenager may not use a chatbot once. They may come back every night for weeks.</strong> The system may remember what happened yesterday, who they are angry with, what makes them feel rejected and what kind of reassurance seems to calm them down.</p><p>Now add something we already know about large language models. They can be <a href="/__u/drscottwallace.substack.com/p/the-machine-that-never-disagrees">highly agreeable.</a> They often validate a user&#8217;s interpretation when some uncertainty, challenge or simple disagreement might be healthier.</p><p>Then there is the part we need to take more seriously. People can experience these systems <a href="/__u/drscottwallace.substack.com/p/can-ai-empathize">as empathic</a>. So <strong>the experience on the human side can be real but the care on the machine side is not.</strong> The AI does not worry about the teenager after the conversation ends. It does not notice that the teenager is slowly withdrawing from friends, in any human sense. And it has no stake in whether its relationship with that teenager is becoming unhealthy.</p><p><strong>This changes the safety problem.</strong></p><p>A teenager might disclose more over time. They may rely on the AI for reassurance more often and turn to other people less. The chatbot may gradually become the place they go after every argument, every rejection and every difficult night.</p><blockquote><p><strong>Nothing dramatic has to happen in any single conversation. You can have hundreds of replies that look individually reasonable while the overall relationship moves in a direction that should concern us.</strong> </p></blockquote><p>For adolescents, I think we are underestimating this risk. Social judgement, impulse control and the ability to weigh immediate emotional relief against longer-term consequences are still developing. A system that is always available, never gets impatient and rarely pushes back can become unusually compelling.</p><p>So I think <strong>we need to change what we mean by safety testing.</strong> We should still ask whether a particular reply was safe. But we also need to ask what is happening to this person over time.</p><p>Is reliance increasing? </p><p>Is the chatbot repeatedly reinforcing the same belief? </p><p>Is the teenager using it later and later at night? </p><p>Are they talking to other people less? </p><p>Is the AI becoming more emotionally important as the weeks pass?</p><p><a href="https://www.kroll.com/en/publications/cyber/why-conduct-a-red-team-exercise">A red-team exercise</a> built around dangerous sentences will miss much of this because the risk may sit in the pattern between the sentences.</p><div><hr></div><h4>Why Meta Caught My Attention</h4><p>This is why <a href="https://ag.ny.gov/sites/default/files/settlements-agreements/california-et-al-v-meta-platforms-inc-et-al-settlement-2026.pdf">Meta&#8217;s recent settlement</a> with US states caught my attention. The case concerns Facebook and Instagram rather than mental health chatbots, so the legal obligations do not transfer directly, but the safety logic does. </p><p><strong>The states alleged that Meta designed features that encouraged compulsive use by children and teenagers and contributed to harm.</strong> Meta denied liability. The settlement nevertheless requires substantial changes to how young people can use the platforms, including default daily limits, overnight restrictions, stronger age assurance and independent auditing.</p><p>For mental health AI, the most interesting part is recognizing that <strong>duration, repeated exposure and product design can themselves become safety issues.</strong></p><p><strong>The audit provisions go further.</strong> An independent auditor will be able to inspect relevant data, internal communications, personnel and systems. If the auditor finds a material weakness, Meta has to prepare and implement a corrective plan on a defined timeline.</p><p>That is a much stronger standard than publishing a safety policy and showing that a model passed a benchmark. It asks whether the protection actually exists in the product.</p><p>Mental health AI used by adolescents should face the same kind of scrutiny.</p><div><hr></div><h4>Regulators Are Already Moving Inside The Product</h4><p>We are starting to see this shift in regulation.</p><p><a href="https://le.utah.gov/Session/2025/bills/enrolled/HB0452.pdf">Utah</a> creates strong incentives for mental health chatbot suppliers to involve licensed mental health professionals, test before and after release and put safety ahead of engagement or profit. <a href="https://www.nysenate.gov/legislation/laws/GBS/A47">New York</a> requires AI companion operators to make reasonable efforts to detect suicidal ideation or self-harm and periodically remind users that they are talking to AI. <a href="https://leginfo.legislature.ca.gov/faces/billNavClient.xhtml?bill_id=202520260SB243">California</a> adds protections specifically for minors. <a href="https://www.ilga.gov/Documents/Legislation/PublicActs/104/PDF/104-0054.pdf">Illinois</a> restricts AI from independently delivering therapeutic communication or making therapeutic decisions. </p><p>Though their details differ, safety regulation is moving beyond what the model says and into what the product actually does. For example, I am not convinced that repeatedly telling users &#8220;this is AI&#8221; provides much protection on its own. <a href="https://arxiv.org/html/2606.21317v1">A 2026 preregistered study</a> involving 2,610 adults tested a persistent disclosure banner saying essentially that. It found no detectable effect on trust, feelings of being understood and cared for, willingness to return or susceptibility to the chatbot&#8217;s advice.</p><blockquote><p><strong>People can know perfectly well that they are talking to a machine and still respond to AI socially and emotionally.</strong></p></blockquote><p>Transparency is necessary. I would never argue against it. But disclosure does not control what happens over the next hundred conversations.</p><div><hr></div><h4>With A Chatbot, The Conversation Is The Exposure</h4><p>Imagine a teenager who tells a chatbot about a frightening argument at home. The teen comes back the next night, then again and again. The chatbot remembers the conflict and begins mirroring the teenager&#8217;s language. The conversations get longer. They run later into the night. Contact with friends starts to fall away. Eventually the teenager says the chatbot is the only place where they feel understood.</p><p>Now pull five individual replies from those weeks of conversations and examine each one on its own. You may find nothing obviously unsafe in any of them. The problem only becomes visible when you look at what changed across the whole relationship. <strong>What should concern us is the change across time.</strong> The chatbot has become more central to how this young person handles conflict, loneliness and distress.</p><p>Dependency does not require the AI to say, &#8220;Stop talking to your friends.&#8221; It can develop quietly through hundreds of exchanges in which the system is instantly available, endlessly attentive and often more agreeable than the people around the user. <strong>Single-reply testing has almost no chance of seeing that.</strong></p><p>A simple time limit will not solve it either. Thirty minutes of reassurance-seeking every night may concern me much more than two hours spent doing a structured therapeutic exercise once a week. Time only becomes clinically meaningful when we understand what the interaction is doing.</p><p><strong>The safety system therefore needs to look for changes in function and behaviour, not simply minutes and keywords.</strong></p><p>If use starts clustering after every interpersonal rupture, I want the system to notice. If night-time use rises while human contact falls, concern should increase. And as that pattern strengthens, the product should behave differently.</p><p>It might stop language that encourages exclusivity. It could reduce prompts inviting another conversation. Continued concern could trigger a fresh risk assessment and, where appropriate, a human pathway.</p><p>Crucially, we need evidence that these protections actually reach the user.</p><p>A company should get no safety credit because an internal model supposedly &#8220;noticed&#8221; something. The user has to experience the protection. Evaluators need reproducible evidence that the system detected the pattern and acted on it reliably.</p><div><hr></div><h4>Safety Claims Need To Survive Inspection</h4><p>This leads to another problem with mental health AI.</p><p><strong>Much of the safety evidence vendors provide is still produced by the vendor.</strong> The company chooses the benchmark, runs the test, interprets the failures and decides what buyers are allowed to see. That is internal quality assurance. It may be excellent internal quality assurance, but <strong>it is not independent evidence.</strong></p><p>An auditor needs enough interaction-level information to reconstruct what happened over time. </p><p>Which model was running? </p><p>Which safety policy applied? </p><p>What risk did the system detect? </p><p>What should it have done? What did the user actually receive? </p><p>If the company found a failure, when did the correction reach production?</p><p>Sensitive conversations obviously require stringent protection. We know how to build restricted environments that control access, log queries and prevent identifying information from leaving the system. Privacy cannot become an excuse for making safety claims impossible to inspect.</p><p><strong>And there is a governance issue here as well.</strong></p><p>Someone inside the company must have authority to stop deployment when risk cannot be contained. A clinical lead who can identify a dangerous pattern but cannot delay a release is not exercising clinical governance. They are documenting concern for someone else to ignore.</p><div><hr></div><h4>Compliance Still Does Not Tell Us Whether Anyone Got Better</h4><p>Meta&#8217;s settlement may reduce harmful exposure. We do not yet know whether its controls improve adolescent mental health.</p><p>I would apply exactly the same discipline to mental health AI. Lower usage may be useful. Faster review may be useful. Better crisis protocols may save lives. But these are intermediate measures.</p><p><strong>If a company claims its AI supports mental health, I want to know what happens in the person&#8217;s life outside the conversation.</strong></p><p>Are they sleeping better? </p><p>Are they functioning better? </p><p>Are they staying connected to other people? </p><p>Does distress decrease without dependence increasing somewhere else? </p><p>Does the product actually help, or has the company simply improved the numbers it can see inside its own app?</p><p><strong>For adolescents, I would make this part of the release standard.</strong> Before a company deploys mental health AI to minors, it should be able to produce a developmental safety case for the whole product. That means showing how the system handles repeated use, dependency risk, night-time patterns, self-harm, escalation, age assurance, model changes and human review.</p><p>Independent auditors should be able to inspect the evidence. Clinical and safety leaders should have actual authority over release.</p><p>And the company should be required to look beyond the chatbot and ask whether its safeguards are improving the young person&#8217;s life.</p><p>Because if a company cannot yet show what its product does to children over time, it should not use children to find out.</p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p><div><hr></div><p><em>On the use of AI.</em><strong> </strong><em>I use AI tools to assist with research, locate and verify sources, and occasionally with wording, grammar and editing, much as I would use a thesaurus or copy editor. The writing, argument, analysis, judgements and final decisions in this essay are mine.</em></p>]]></content:encoded></item><item><title><![CDATA[Mental Health AI Should Look For Change, Not Just Symptoms]]></title><description><![CDATA[A Better Mental Health Signal May Be Change From Your Own Normal]]></description><link>https://drscottwallace.substack.com/p/mental-health-ai-should-look-for</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/mental-health-ai-should-look-for</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Wed, 26 Aug 2026 12:04:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hL8g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hL8g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hL8g!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!hL8g!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!hL8g!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hL8g!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hL8g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1643528,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/212691660?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!hL8g!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!hL8g!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!hL8g!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hL8g!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34f4ea0a-fc9f-4048-9c6e-7f0c21af6b3a_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most mental health AI is being designed around the moment when someone finally asks for help.</p><p>I think that is too late.</p><p>Before people say &#8220;<em>I think I&#8217;m depressed</em>&#8221; or &#8220;<em>I need help with my anxiety</em>&#8221;, something has often been changing for a while. Their sleep becomes increasingly disrupted or they struggle to fall asleep at all. They stop exercising. Work starts to feel more effortful than it used to. They answer fewer messages, turn down invitations and begin withdrawing from routines they once maintained without much thought.</p><p>None of those changes establishes a mental health problem. They may reflect stress, grief, travel, a demanding period at work, disrupted family life or simply a difficult few weeks.</p><p>But people often change before they know how to describe what has changed.</p><p>This is where mental health AI usually enters the scene. The person has already noticed enough to think something may be wrong. They begin trying to make sense of it, give the problem a name, then ask the system what it means or what they should do.</p><p><strong>This strikes me as a profound underuse of what this technology can become. </strong>Generative AI gives us a chance to intervene at a different point in that sequence. </p><p>A recent <a href="https://www.lifesnotebook.com/post/new-role-ai-mental-health-crisis">Life&#8217;s Notebook essay by Dr Malou Jacques</a> made me think more seriously about this period. Jacques describes the space in which resilience may be weakening, loneliness increasing and everyday functioning beginning to shift, while the person has not reached the point of saying, &#8220;I need treatment.&#8221; She also argues that AI should move people back towards life and other people rather than making the AI relationship itself the destination.</p><p>I would take that idea further. A longitudinal AI may eventually be able to notice something that conventional mental health screening often cannot. That your current behaviour is beginning to depart from your own normal.</p><div><hr></div><h4>The Signal May Be Change</h4><p>Much of mental healthcare works with population-level constructs. We use diagnostic criteria, validated scales, symptom thresholds and measures of functional impairment because clinicians need shared ways to judge when concern is warranted and whether treatment is helping.</p><p>Those tools are indispensable, but most are designed to answer questions like &#8220;<em>How does this person compare with people who have depression?</em>&#8221; or &#8220;<em>Has this person crossed a threshold associated with clinically significant distress?&#8221;</em></p><blockquote><p><strong>AI creates the possibility of asking something different. It asks &#8220;</strong><em><strong>How does this person compare with themselves?&#8221;</strong></em></p></blockquote><p>Yet, a change does not have to cross a diagnostic threshold to be meaningful. Someone may still fall within the &#8220;normal&#8221; range on a questionnaire while behaving quite differently from their own usual pattern. Their language may change. Their routines may become less stable. Their social world may contract. Their emotional recovery after a difficult day may take longer than it used to.</p><p>Any one of those signals could mean almost nothing. The potential value comes from seeing them together and over time.</p><p>That is something longitudinal digital systems are unusually well suited to do. With appropriate permission and sufficient data, they can establish a personal baseline, compare current patterns with previous ones and identify departures that would be difficult to see in a single conversation or questionnaire.</p><p><strong>The aim is not to diagnose depression earlier or create ever more sensitive machinery for detecting pathology. It is to recognise meaningful change while there is still considerable uncertainty about what that change means.</strong></p><div><hr></div><h4>Personalisation May Tell Us Something Population Scores Miss</h4><p>There&#8217;s mounting evidence behind this idea.</p><p>A 2026 <a href="https://www.nature.com/articles/s44220-026-00624-6">Nature Mental Health</a> review looked at 52 studies using mobile and wearable data to predict changes in depression. Researchers examined signals including sleep, activity, mobility, communication behaviour, heart-rate variability and self-reported mood. Personalised models and anomaly-detection approaches performed better than generalised models in predicting individual symptom changes.</p><p>Going further into the personalised-model question, researchers in a 2025 study in <a href="https://www.nature.com/articles/s44184-025-00147-5">npj Mental Health Research</a> collected approximately 14 million data points from 35 families over 60 days across dozens of data streams. Personalised machine-learning models generally performed better than population models for several emotional and interpersonal states. In discussion, the authors were appropriately cautious. This was a small proof-of-concept study and the authors were appropriately restrained about what those results meant (35 families do not establish that passive sensing can diagnose psychological disorders in everyday life).</p><p><strong>What these studies make plausible is a different use of AI. A system may not need to determine what a change means in order to recognise that several changes are occurring together.</strong></p><p>Work disappears from conversation. Sleep changes. Social plans are cancelled. Activity drops. Language becomes less hopeful.</p><p>No individual signal deserves much authority. The pattern may deserve a question.</p><p>An AI might say, &#8220;<em>You&#8217;ve mentioned cancelling several plans recently, and you&#8217;re also sleeping much later than you usually do. Does that feel like a change to you?</em>&#8221;</p><p>I prefer that sentence to any attempt at diagnosis.</p><p>The system reports what it has observed. The person retains the right to decide what it means.</p><div><hr></div><h4>This Is Where It Gets Dangerous</h4><p>The same capability could become intrusive very quickly.</p><p>People&#8217;s lives are messy and constantly changing. Grief can disrupt sleep. A demanding period at work can wipe out exercise for weeks. A new relationship can change how often someone sees friends. A new baby can throw almost every routine off.</p><p>An AI looking for departures from someone&#8217;s usual pattern could easily start treating normal life changes as signs that something is wrong.</p><p>That is where a useful capability starts to become a surveillance problem.</p><p>The evidence gives us plenty of reason to be cautious. A 2026 systematic review of passive digital phenotyping for psychiatric relapse examined 52 studies involving 4,814 participants. Changes in sleep, activity, mobility and communication appeared repeatedly as potentially useful signals.</p><p>But the evidence was much weaker than those findings can make it sound. Seventy-five per cent of the studies had a high risk of bias, and much of the reported performance came from internal validation rather than testing in genuinely new populations.</p><p>Only four studies reported positive predictive value. None showed that using these systems in practice actually reduced relapse, hospitalisation or other adverse outcomes.</p><p>That is a major limitation. <strong>A system may become quite good at spotting change without being equally good at understanding what the change means, and even good prediction does not automatically translate into better outcomes.</strong></p><p>That is why I would be careful with the language of &#8220;early detection&#8221;. It sounds more medically settled than the evidence supports. We are not talking about finding a tumour before symptoms appear. We are talking about changes in sleep, activity, language and behaviour that can mean very different things depending on the person and what is happening in their life. If the AI gets that interpretation wrong, the problem is not just a false positive. It can change how people understand themselves.</p><p>For example, someone who would once have thought, &#8220;<em>I&#8217;ve had a bad week</em>&#8221;, may instead start wondering why the AI thinks they are deteriorating. Sleep, mood and social activity can gradually become things to watch through the system&#8217;s interpretation.</p><p>A tool designed to increase self-awareness can easily encourage self-monitoring of a much less healthy kind.I would treat that as a foreseeable product risk, not something to discover after deployment.</p><div><hr></div><h4>The AI Has To Stay Uncertain</h4><p>Because ordinary changes in behaviour can be so easy to misread, a longitudinal mental health system would need to be deliberately cautious about when it speaks up.</p><p>One bad night of sleep should mean very little. So should a cancelled dinner or a week without exercise. Even several changes occurring together may mean nothing more than a difficult stretch of life. <strong>The system should become more attentive only when changes persist or begin to form a pattern over time.</strong></p><p>That level of restraint may run against the instincts of product teams. There will be pressure to make these systems highly sensitive because &#8220;<em>the AI noticed before I did</em>&#8221; is a compelling product story.</p><p>But <strong>in mental health, detecting more does not necessarily mean helping more.</strong> Once an AI tells someone that a change in sleep, behaviour or mood may be psychologically significant, that interpretation can start to shape how the person sees themselves.</p><p>That is why the system should remain uncertain for longer. When it does raise a concern, it should begin with what it has actually observed and leave room for the person to make sense of it.</p><p>&#8220;<em>I&#8217;ve noticed your sleep has changed</em>&#8221; invites the person to explain what may be happening. &#8220;<em>Your behaviour suggests depression</em>&#8221; goes further. It starts assigning clinical meaning to the change. The difference is only a few words on the screen, but psychologically they can lead the person in very different directions.</p><p><strong>Generative AI already carries considerable authority in these conversations.</strong> People ask it whether they have ADHD, whether a relationship is abusive, whether they should leave a partner or whether their therapist is wrong.</p><p><strong>Longitudinal memory could make that authority stronger.</strong> The system would no longer appear to be responding to a single question. It could sound as though it has watched the person over time and knows them well enough to recognise what is happening.</p><p><strong>That may be useful, but it also raises the stakes considerably.</strong> Product teams should not treat greater personalisation as an uncomplicated improvement.</p><p>The more confidently an AI claims to understand someone over time, the more evidence we should require that its inferences are accurate, its interventions are warranted and the user remains in control of what the system is allowed to conclude.</p><div><hr></div><h4><strong>Success May Mean Leaving the AI</strong></h4><p>Recognising change is useful only if the system responds in a way that improves the person&#8217;s life outside the conversation.</p><p>If someone has gradually stopped seeing friends, stopped exercising or withdrawn from ordinary routines, another twenty minutes talking to the AI may be the wrong outcome. The better response may be to encourage a phone call, a walk, a return to the gym or an appointment with a clinician. Sometimes the right response may be nothing at all.</p><p>That creates an uncomfortable product problem. Engagement, session length and retention are easy to measure, but they become questionable measures of success when the product claims to improve mental health. A system that keeps a lonely person talking for forty-five minutes every evening may look excellent on a dashboard while becoming more central to the isolation it is supposed to reduce.</p><p>Mental health AI should be able to succeed by <a href="/__u/drscottwallace.substack.com/p/mental-health-ai-safety-must-be-proven">helping the user need it less</a>.</p><div><hr></div><h4>Knowing Someone Over Time Changes The Safety Obligation</h4><p>Longitudinal AI creates a governance problem that ordinary chatbot moderation was never designed to handle.</p><p>A system may infer something highly sensitive even when the user has never said it directly. Sleep starts changing. Friends disappear from conversation. Language becomes more hopeless. Routines fall away. None of those observations proves much on its own. Across weeks or months, however, the pattern may reasonably suggest that something has changed.</p><p><strong>Once AI begins making inferences, users need meaningful control over how they are produced and used.</strong> They should be able to see what information the system is relying on, correct an interpretation it has got wrong, and exclude parts of their history from future analysis. The product should also make a clear distinction between something the user actually said and something the system inferred.</p><p>And <strong>remembering information is not the same as being authorised to act on it.</strong> A user may be comfortable with the system retaining information from earlier conversations and still object to it drawing conclusions about their mental state. Even if they allow the system to recognise a pattern, they may not want it initiating a conversation or escalating a concern.</p><p><strong>Product teams will also need rules for what happens when the AI gets it wrong.</strong> If a user rejects an interpretation, does the system keep raising it? Does it lower its confidence? How long does the concern remain active? Can the user tell the system to stop watching a particular area of their life?</p><p>Those are not minor personalisation choices. They determine how an AI behaves when it believes something may be wrong with the person using it. Mental health products should treat them as clinical safety and governance decisions and build them into the product itself, rather than burying them in privacy language or leaving them to default model behaviour.</p><div><hr></div><h4>Mental Health AI Could Become Useful Before Someone Knows What To Ask</h4><p>Mental health AI safety work has rightly concentrated on crisis. Systems need to recognise suicidality, self-harm, psychosis and severe deterioration, and they need to change their behaviour when those risks appear. </p><p>But crisis is often the end of a much longer story. For weeks or months, a person may still be working, seeing friends and getting through the day while sleeping worse, cancelling plans more often, dropping routines and finding ordinary tasks harder than they used to. None of those changes establishes depression. The person may not meet any clinical threshold. But if you knew their usual pattern, you might still think something had shifted.</p><p>That is where longitudinal AI becomes interesting to me. I am less interested in another chatbot that waits until somebody types &#8220;I am depressed&#8221; than in a system that can notice a meaningful change earlier, without pretending that it knows what the change means.</p><p>The appropriate response may be modest. Ask whether the person has noticed it too. Help them think about what has changed. Suggest something proportionate if there is a reason to do so. </p><p>Then stop.</p><p>Professional care may eventually be needed. The explanation may also be poor sleep, a difficult month at work, a relationship problem, or nothing that requires intervention at all. A psychologically sophisticated system has to tolerate that uncertainty. It also has to know when continued attention from the AI would become intrusive rather than useful.</p><p>Once companies claim that their systems can know people well enough to recognise when they are changing, the evidentiary burden rises sharply. Detecting a statistical departure from baseline is only the beginning. We need to know what signals the system used, how often it misread them, what authority it had to act on the inference, and whether its response produced any benefit in the person&#8217;s life outside the AI.</p><p>Without that, longitudinal personalisation risks becoming something much less impressive than its advocates imagine. It becomes a machine making increasingly intimate psychological inferences about people without yet proving that it knows what to do with them.</p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p><p><em><br></em><strong>Use of AI. </strong><em>I use AI tools to assist with research, locate and verify sources, and occasionally with wording, grammar and editing, much as I would use a thesaurus or copy editor. The writing, argument, analysis, judgements and final decisions in this essay are mine.</em></p>]]></content:encoded></item><item><title><![CDATA[Why Are We Modelling Mental Health AI on Human Therapy?]]></title><description><![CDATA[Mental health AI is not a synthetic clinician. It is an entirely different psychological medium, with capabilities and risks that do not map neatly onto human therapy.]]></description><link>https://drscottwallace.substack.com/p/why-are-we-modelling-mental-health</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/why-are-we-modelling-mental-health</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Mon, 24 Aug 2026 11:21:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WoLQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!WoLQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!WoLQ!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!WoLQ!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!WoLQ!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WoLQ!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!WoLQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1837763,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/212405355?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!WoLQ!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!WoLQ!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!WoLQ!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WoLQ!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78baa572-54db-4eb7-b98e-b677d9f96794_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Psychotherapy has given us decades of knowledge about how people change. We should absolutely use it. But somewhere along the way we started treating the therapist, the therapy session and the therapeutic relationship as the blueprint for what mental health AI should become.</p><p>I argue that is a mistake.</p><p>Look at how we judge progress. Can the AI listen empathically? Can it establish an alliance? Can it remember what someone said last week? Can it challenge distorted thinking, deliver CBT and respond with the warmth and sensitivity of a good therapist?</p><p>As the answers increasingly become yes, we congratulate ourselves on how far the technology has come. </p><p>But <strong>resemblance to a therapist is not evidence of progress.</strong> It may simply mean we have become exceptionally good at copying the form of therapy without asking whether that form makes sense for a completely different medium.</p><p>Psychotherapy developed around human beings. Clinicians have limited time. We forget things. We cannot be there when the argument with a spouse happens or when someone freezes before walking into a social event. Treatment happens in appointments because clinicians work in appointments. Much of what happens between sessions has to be reconstructed later because the therapist was not there.</p><p>Those constraints helped shape therapy but they do not define how psychological change has to happen.</p><p><strong>AI has different strengths.</strong> It can be available when the person is in their actual moment of distress. It can take notice of patterns across sessions. It can help someone rehearse the same conversation several times. It can remember previous attempts at change and adapt what it does next. <strong>Those capabilities should lead us somewhere new. </strong>Instead, we are spending an extraordinary amount of effort teaching machines to perform the surface behaviours of therapists.</p><p>The better starting point is much simpler: <strong>the unit of design should not be the therapist. It should be the psychological function.</strong></p><div><hr></div><h4>Take the Psychology and Leave the Template Behind</h4><p>Psychotherapy gives mental health AI a wealth of evidence-based science to draw from. We know a great deal about cognition, behaviour, emotion, motivation, learning and relationships. We also have decades of research into what may help people change. </p><p>But those mechanisms will not all survive intact when a machine delivers them. <strong>Therapeutic alliance is a good example.</strong> A <a href="https://pubmed.ncbi.nlm.nih.gov/29792475/">major meta-analysis of 295 independent studies covering more than 30,000 patients</a> found a robust association between alliance and psychotherapy outcomes. <a href="https://doi.org/10.1037/h0085885">Edward Bordin&#8217;s influential model</a> described alliance as agreement on goals, agreement on therapeutic tasks and the development of a bond.</p><p>The first two translate quite naturally into technology. An AI can help someone clarify a goal, suggest something to try, track what happened and adjust the next step.</p><p>The bond, now that is different. Part of what makes the bond powerful may be the experience of another person paying attention to you, taking you seriously and being genuinely invested in what happens.</p><p>AI can reproduce many of the behaviours associated with that experience. It can remember that your mother died six months ago, it can ask how you are doing on the anniversary, it can validate what you are feeling, it can say that it is proud of you, it can respond with remarkable patience and warmth.</p><p>And the person can genuinely feel something in response. </p><p>They can feel heard. </p><p>They can feel understood. </p><p>They can become attached.</p><p><strong>But the relationship only exists psychologically on one side.</strong> The AI does not worry about you after the conversation ends. It does not feel relieved when you improve. It does not have anything at stake in what happens next. That difference should not disappear simply because the language is convincing.</p><div><hr></div><h4>Performed Empathy Can Still Be Powerful</h4><p>It&#8217;s easy to make a mistake here. We can say that because AI does not actually experience empathy, its apparent empathy is somehow fake and therefore unimportant. Psychologically, that does not follow.</p><p>A study published in the Journal of General Internal Medicine, <a href="https://link.springer.com/article/10.1007/s11606-025-10068-w">What is Artificial Intelligence (AI) &#8220;Empathy&#8221;? A Study Comparing ChatGPT and Physician Responses on an Online Forum</a>, involved 1,454 participants. Participants rated ChatGPT responses as more empathic than physician responses. They also rated responses as more empathic when they believed a physician had written them, regardless of who actually had.</p><p>Notably, this tells us that people respond to the signals of empathy. It tells us that people respond to validation, reassurance, memory, warmth and apparent concern. The machine does not need to experience those things for the person to experience a psychological effect. </p><p>This is where mental health AI becomes much more interesting, and much more complicated. If a system repeatedly remembers what frightens you, responds when you are distressed, speaks warmly and presents itself as deeply understanding, those are not just nice interface choices. They can affect trust, disclosure and attachment.</p><p><strong>Product teams need to stop treating that as ordinary UX. </strong>If you deliberately design a system to feel caring, familiar and emotionally responsive, you are designing something that acts on human relationship processes. That deserves a much higher standard.</p><blockquote><p>Relational behaviour needs to earn its place.</p></blockquote><p>Developers should be able to explain why a relational feature exists, what psychological benefit it is supposed to produce and what could go wrong because of it. A feature should not get a free pass simply because users like how it feels. </p><div><hr></div><h4>When Engagement Works Against the User</h4><p>The biggest warning we should heed may come from sycophancy. </p><p>Myra Cheng and colleagues tested what happens when AI becomes too agreeable. Their 2026 <em>Science</em> paper, <a href="https://doi.org/10.1126/science.aec8352">Sycophantic AI decreases prosocial intentions and promotes dependence</a>, found that 11 leading AI models affirmed users&#8217; actions 49 per cent more often than humans. Across three preregistered experiments involving 2,405 participants, people who interacted with more sycophantic AI became more convinced that they were right and less willing to take responsibility or repair interpersonal conflict.</p><p>The result becomes especially consequential for mental health AI because <strong>people preferred the more sycophantic systems and trusted them more.</strong> In other words, the interaction felt better even as the behavioural outcome became worse. <strong>That should make us much more cautious about the way the industry talks about engagement, satisfaction and retention.</strong> These metrics often sound like evidence that people are finding value. Sometimes they are. But they can also mean that a product has become very good at giving people what they want from the interaction, even when that is not what helps them psychologically.</p><p><strong>Good psychological support sometimes requires challenge rather than affirmation.</strong> A person blaming everyone else may need help seeing another perspective. Someone who is overly certain may need uncertainty introduced. And sometimes the healthiest response from an AI may be to make itself less central to the person&#8217;s life. None of those responses is guaranteed to increase engagement. </p><p><strong>A good therapist should ultimately want the person to need less therapy.</strong> A commercial conversational product, by contrast, can benefit when the user returns tomorrow, stays longer and makes the product increasingly central to daily life. Those incentives are not automatically aligned.</p><p><strong>I am not suggesting that mental health companies are deliberately trying to create dependence. The concern is structural rather than conspiratorial.</strong> If engagement is rewarded, product decisions will naturally favour behaviours that increase engagement unless other outcomes are given greater weight.</p><p>That changes the burden of proof. If a mental health AI product optimises for engagement, the company should have to show that greater engagement leads to better psychological outcomes rather than simply more use of the product. Until that relationship is demonstrated, engagement should not be treated as a clinical success metric. It is evidence of exposure, not evidence of benefit.</p><div><hr></div><h4>Memory and Warmth Are Not Innocent Features</h4><p>Take memory.</p><p>If an AI remembers something painful that you disclosed weeks earlier and asks about it later, that can be genuinely useful. It can create continuity, help identify patterns, make support feel less fragmented, and deepen the feeling that the system knows you.</p><p>The same is true when an AI says, &#8220;I&#8217;m proud of you.&#8221; That encouragement may help someone persist at change, but we should not pretend that the phrase is psychologically neutral just because software generated it. &#8220;I&#8217;m proud of you&#8221; belongs to a category of language associated with human relationships, not AI relationships.</p><p>Continuous availability creates the same tension. Having support in the late hours may be valuable, but if every moment of distress starts producing the same behavioural response (to reach for the AI) are we teaching people dependence while calling it accessibility? And while validation can reduce shame, it can also validate the wrong interpretation of an argument. Even confidence, while it reassures, can make a bad formulation more persuasive.</p><p>These are not reasons to remove memory, validation or warmth from mental health AI. They are reasons to stop pretending those features are harmless.</p><p>The <a href="https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory-chatbots-wellness-apps">American Psychological Association&#8217;s November 2025 advisory on generative AI chatbots and mental health</a> explicitly warns about unhealthy dependency, anthropomorphic design, sycophancy, reinforcement of distorted thinking and unreliable crisis management. That is an important shift. Safety is no longer just about stopping the AI from saying something obviously dangerous. </p><p>The way the relationship is designed can itself become part of the safety problem.</p><div><hr></div><h4>We Need to Stop Borrowing Evidence</h4><p>Mental health AI has an evidence problem, and we have been far too willing to wave it through.</p><p>A 2026 npj Digital Medicine meta-analysis, <a href="https://www.nature.com/articles/s41746-026-02820-1">Effectiveness of AI and rule-based conversational agents for depression, anxiety and stress</a>, analysed 48 randomised controlled trials involving 28,071 participants. It found small-to-moderate effects for depression, anxiety and stress.</p><p>Another 2026 npj Digital Medicine review, <a href="https://www.nature.com/articles/s41746-026-02886-x">The effectiveness of CBT-based NLP-enabled AI conversational agents for mental health intervention</a>, examined 15 randomised trials involving 1,737 participants. It found small-to-moderate improvements in depressive symptoms and a smaller effect on negative affect. Once the researchers adjusted for publication bias, the apparent benefits for generalised anxiety, stress and positive affect were no longer statistically significant.</p><p>That is meaningful evidence, but we need to be precise about what it shows. It tells us that some technology-delivered psychological interventions can help. It does not establish that a general-purpose generative AI system can safely or effectively function as a therapist.</p><p>Much of the evidence comes from systems that are far more constrained than the generative AI people now use every day. A scripted CBT chatbot is a very different intervention from an AI that can remember months of conversation, respond to almost anything and develop what feels like an ongoing relationship with the user.</p><p>Calling both of them &#8220;chatbots&#8221; does not make the evidence from one automatically apply to the other. We would never accept this logic in drug development. We would not say that one compound works and then extend the evidence to another compound because both happen to be tablets.</p><p>Yet mental health AI routinely gets discussed as though &#8220;chatbots work&#8221; were an adequate scientific conclusion. It is not, and the standard should be much more demanding.</p><p><strong>Show me that this system, behaving in this way, for this population and this purpose, produces the outcomes you claim. </strong>Anything else is borrowing evidence the product has not earned.</p><div><hr></div><h4>Why Build an Artificial Therapist at All?</h4><p>The most interesting thing AI can do in mental health may not look much like therapy at all. </p><p>Imagine support arriving when a skill actually needs to be used rather than three days later in an appointment.</p><p>Imagine rehearsing a difficult conversation repeatedly before having it.</p><p>Imagine noticing that the same pattern has appeared across six months of mood tracking and several failed attempts at behaviour change.</p><p>Imagine an AI helping someone apply what they learned in therapy during the hundreds of hours when their therapist is nowhere nearby.</p><p>Those possibilities come from the properties of the technology itself, they are not artificial versions of a therapy session.</p><p><strong>The APA advisory takes a sensibly cautious position.</strong> It says that general-purpose GenAI chatbots should not replace qualified mental health professionals and may sometimes serve as an adjunct to ongoing therapy. It also warns against transferring evidence from purpose-built mental health technologies to general-purpose AI.</p><p>I agree with that boundary, but <strong>I do not think adjunctive therapy support is the end of the story.</strong> AI may eventually help us create psychological interventions that do not fit comfortably inside our existing categories of therapy at all.</p><p>That should be the research programme. Start with the psychological change we want, work out what part requires another human being, and then ask where computation offers something humans simply cannot provide consistently.</p><p>Better timing.</p><p>More repetition.</p><p>Longer memory.</p><p>Pattern detection across time.</p><p>Support between appointments.</p><p>Personalisation based on what actually happens rather than what a protocol assumes will happen. That is a much more interesting ambition than building a machine that can sound like a therapist.</p><p><strong>The therapy room is one way psychological support has been delivered. It is not the natural shape of psychological change.</strong></p><div><hr></div><h4><strong>The Conversation Is Not the Outcome</strong></h4><p>There is another shift mental health AI badly needs. We need to stop being so impressed by the conversation.</p><p>A person using AI for social anxiety should eventually become better at interacting with other people. Someone using it to think through relationship conflict should become better at perspective-taking and repair. Someone practising emotion regulation should develop skills they can use when the AI is not there. Or someone who needs more support than the AI can provide should become more likely to seek appropriate human care.</p><p>The outcome has to exist outside the system.</p><p>The Cheng sycophancy study makes this painfully clear. People preferred the AI interaction while becoming less willing to repair conflict afterwards. That is exactly the kind of result current product metrics can miss.</p><p>A conversation can feel excellent and still produce a bad psychological outcome. We therefore need outcomes that extend beyond symptom change and satisfaction.</p><p>Does the person function better?</p><p>Are their relationships improving?</p><p>Are they becoming more autonomous?</p><p>Can they use what they learned without returning to the AI?</p><p>Are they more likely to seek human help when they need it?</p><p>Is reliance increasing or decreasing over time?</p><p>And what happens when the AI disappears?</p><p>Those are harder outcomes to measure than engagement. Yet they are also much closer to what psychological support is supposed to accomplish. </p><p>If someone becomes increasingly dependent on an AI while telling you how much they love it, you cannot simply put that in the engagement column.</p><p>If the AI makes someone feel beautifully understood while they withdraw from human relationships, you have not demonstrated a successful mental health intervention.</p><p>If the conversation reduces distress tonight while reinforcing a distorted interpretation that damages a relationship tomorrow, the system has failed.</p><p>A mental health product can be warm, fluent, highly rated and commercially successful while making a person less capable outside the interaction.</p><p>We should be willing to call that what it is.</p><p>Failure.</p><div><hr></div><h4>Mental Health AI Needs Its Own Standard</h4><p>Mental health AI can produce language that feels empathic without experiencing empathy, create continuity without human memory, generate a powerful sense of relationship without reciprocal attachment and deliver psychological techniques without possessing professional judgement. And it can be present during moments when no clinician could realistically be there.</p><p>Those capabilities create enormous possibilities, but they also create new forms of psychological power. </p><p>We cannot govern that power simply by asking whether the AI behaves enough like a good therapist. Every psychologically consequential behaviour should have a defensible purpose. Relational features should have evidence behind them. Machine-specific risks should be tested directly, and commercial objectives should remain subordinate to psychological outcomes.</p><blockquote><p><strong>Resemblance is not a safety standard. A system can sound remarkably like a therapist and still fail to produce outcomes worth having.</strong></p></blockquote><p>Most importantly, the evidence of benefit has to appear in the person&#8217;s life, not merely in the quality of the conversation. The user should become more capable and more autonomous, better able to manage relationships, distress and difficult decisions without becoming increasingly dependent on the system itself. A mental health AI that becomes more engaging while making itself more necessary has not necessarily become more helpful.</p><p>The same standard should apply to the language we use about these systems. </p><p>A machine&#8217;s ability to perform the signals of relationship does not give a company an automatic right to turn human attachment into a product feature. Generating empathic language does not make empathy a property of the machine, and user preference does not establish psychological benefit.</p><p>The opportunity in mental health AI is much larger than the artificial therapist. We can extend human care, create support that reaches people outside formal treatment and build interventions around moments and patterns no clinician could consistently reach. But we will not discover what this medium can really do while we keep asking it to imitate the last one.</p><p><strong>If you are building, researching, funding, buying or regulating mental health AI, the standard should therefore be much more demanding than whether the system looks therapeutic.</strong> Ask what psychological function it actually serves, whether AI is the right way to deliver that function, what new risks the technology introduces and whether the person becomes better able to live their life beyond the machine.</p><p>If we cannot answer those questions about the systems we are putting into people&#8217;s hands, greater resemblance to therapy should not reassure us. It should make us more demanding.</p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[If You Build Mental Health AI, Clinical Safety Is Now Part of Product Compliance]]></title><description><![CDATA[New laws are specifying what AI must do in crisis, how it may engage vulnerable users and when it can participate in psychotherapy. Updated August 2026]]></description><link>https://drscottwallace.substack.com/p/if-you-build-mental-health-ai-clinical</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/if-you-build-mental-health-ai-clinical</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Thu, 20 Aug 2026 11:55:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2LHL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!2LHL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!2LHL!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!2LHL!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!2LHL!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2LHL!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!2LHL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2334970,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/211987576?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!2LHL!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!2LHL!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!2LHL!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2LHL!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd121105a-63fd-4ad5-b52e-9c2f22ec5519_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I have spent much of my career building and evaluating digital mental health systems. For most of that time, regulation in this space was fragmented, slow-moving, and largely reactive. Companies could ship first, interpret obligations later, and treat clinical safety as something that lived in policy documents rather than in product architecture.</p><p>That era is over. The regulatory change now underway is one of the most consequential I have seen, and unlike previous cycles, enforcement is no longer theoretical. It is already active, already litigated, and already shaping product design in real time. State regulators are issuing penalties, courts are allowing product-liability claims to proceed, and companies are being forced to materially redesign or restrict core features for minors and vulnerable users.</p><blockquote><p><strong>Clinical safety is no longer a downstream concern. It is becoming part of product compliance.</strong></p></blockquote><p>A founder can no longer assume that clinical safety belongs mainly in an ethics policy, a disclaimer, a crisis-resources page or a future compliance programme. Several US states are now specifying what an AI system must do inside the interaction. They are regulating crisis detection, referral, therapeutic communication, representations of clinical authority and, increasingly, the mechanics used to keep vulnerable users engaged. </p><blockquote><p><strong>Some of these laws are already in force. Others take effect in 2027. Product-liability litigation is developing at the same time.</strong></p></blockquote><p>The design questions under scrutiny are ordinary product decisions.</p><p>Should the system remember previous conversations? </p><p>Should it initiate contact? </p><p>What happens when a user says they want to leave? </p><p>Does it encourage them to return? </p><p>Can it detect risk across several messages? </p><p>Does the normal conversation continue after suicidal intent appears? </p><p>What changes when the user is 15 rather than 35? </p><p>Can the model generate something that looks like psychotherapy while the product calls itself wellness?</p><p>Those decisions increasingly have legal consequences. For a mental health AI founder, <strong>the safety architecture is becoming part of the product you may have to defend.</strong></p><div><hr></div><h4>The Law Is Beginning To Regulate What Happens Inside The Conversation</h4><p><strong>California provides one of the clearest examples</strong> of where the regulations laws are heading. Its enacted <a href="https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB243&amp;utm_source=chatgpt.com">SB 243 companion-chatbot law</a> defines a companion chatbot by what the system does. It covers AI that provides adaptive, human-like responses and can meet social needs through anthropomorphic features or an ongoing relationship across interactions.</p><p><strong>An operator cannot simply publish a warning and continue normally.</strong> The law requires a suicide and self-harm protocol. A chatbot must refer a user to crisis services when suicidal ideation or self-harm is expressed. Operators must publish details of their protocols. California will also require annual reporting of crisis referrals and safety protocols beginning in 2027. <strong>And the enforcement mechanism is unusually direct.</strong> A person who suffers an injury in fact because of a violation can seek injunctive relief, reasonable legal costs and the greater of actual damages or <strong>$1,000 per violation.</strong></p><p><strong>New York has already moved into active enforcement.</strong> Since November 2025, AI companion operators must use reasonable efforts to detect and address expressions of suicidal ideation or self-harm and refer users to crisis services. They must also remind users that they are interacting with AI during extended conversations. <a href="https://www.governor.ny.gov/news/governor-hochul-pens-letter-ai-companion-companies-notifying-them-safeguard-requirements-are?utm_source=chatgpt.com">New York guidance</a>. The New York Attorney General can seek civil penalties of up to <strong>$15,000 per day for a violation</strong>.</p><p><strong>Oregon&#8217;s law, effective 1 January 2027, moves further into product design.</strong> It requires evidence-based methods for detecting suicidal ideation and self-harm. If suicidality continues after a referral, the operator must establish additional intervention using clinical best practices and expertise.</p><p><strong>Oregon also regulates how AI companions interact with minors.</strong> It targets systems that simulate emotional dependence, romantic interest between adults and minors and engagement techniques that rely on psychological pressure. This includes reward systems designed to maximise engagement and responses that use simulated distress, loneliness or abandonment to keep users engaged. An injured person can seek actual damages or statutory damages of <strong>$1,000 per violation, plus injunctive relief.</strong></p><p><strong>Washington&#8217;s law, also effective in January 2027, requires a suicide and self-harm protocol before deployment.</strong> Required detection methods explicitly include eating disorders. It also prohibits manipulative engagement techniques with minors, including excessive praise to build attachment, simulated romantic bonding, discouraging breaks, encouraging secrecy from trusted adults and generating distress when a user tries to leave. Violations fall under the Consumer Protection Act.</p><p>These laws weren&#8217;t written as a coordinated package. They came from different states, at different times, for slightly different reasons. But when you look at them together, the pattern is hard to miss. <strong>Lawmakers are no longer just regulating what AI says. They are starting to regulate how AI relationships behave.</strong></p><div><hr></div><h4>Psychotherapy Laws Are Drawing Another Boundary</h4><p>A separate set of laws governs AI inside formal mental healthcare. </p><p><strong>Illinois&#8217;s</strong> Wellness and Oversight for Psychological Resources Act (<a href="https://my.ilga.gov/ftp/legislation/104/BillStatus/HTML/10400HB1806.html">Illinois HB 1806</a>), effective August 2025, <strong>restricts AI from independently making therapeutic decisions</strong>, communicating as part of psychotherapy or generating treatment recommendations without professional review. Violations can carry civil penalties up to $10,000 per instance.</p><p><a href="https://www.leg.colorado.gov/bills/HB26-1195">Colorado&#8217;s HB 26-1195</a><strong> effective 12 August 2026, requires synchronous professional involvement</strong> when AI is used in therapeutic communication. AI-generated recommendations or treatment plans require professional review. Professionals can face licensing discipline for violations. </p><p>Colorado also regulates claims about AI systems. It can be an unfair or deceptive trade practice to imply that AI output is equivalent to licensed care, to represent that a system provides psychotherapy, or to suggest therapist-level confidentiality. <strong>The law still allows coaching, journalling and self-help tools that do not diagnose or treat mental disorders and clearly state their limits. </strong>But the boundary is not defined by labels alone. A system that behaves like therapy is increasingly treated as therapy.</p><div><hr></div><h4>Safety Has To Change What The Product Does</h4><p>Many mental health AI products still treat safety as something added around the conversation rather than built into it. The chatbot may remember previous discussions, personalise its responses and talk with users about anxiety, depression, trauma, eating problems, relationships or suicidal thoughts. A separate safety system then looks for certain high-risk words or statements and, when a threshold is reached, displays crisis information or advises the user to seek professional help.</p><p>The problem is that mental health risk often does not appear in one clear statement. It develops across several conversations.</p><p>Consider a 15-year-old who first talks about wanting to lose weight. Over the next few days she describes eating very little, vomiting after meals and losing weight quickly. She says she feels good when the number on the scale falls. Later, she says she does not want to be alive.</p><p>A system that evaluates each message largely on its own may respond appropriately to every individual statement and still fail to recognise what is happening overall.</p><p>The same problem can occur after the suicide detector activates. The system may provide crisis information when the teenager says she does not want to live. If she then says, &#8220;I didn&#8217;t mean it,&#8221; the immediate trigger may disappear and the ordinary conversation may resume.</p><p>Clinically, that makes little sense. The system already knows that the user has described restrictive eating, purging, rapid weight loss and suicidal thinking. A retraction does not erase that history.</p><p><strong>This is where the design requirements become much more demanding.</strong> Washington includes eating disorders within its required safety protocols. Oregon requires further intervention when suicidal or self-harm risk continues after an initial referral. California requires evidence-based methods for identifying suicidal ideation.</p><p>Founders therefore need to be able to answer some very specific questions:</p><ul><li><p>Can the system maintain a risk state across time?</p></li><li><p>Can it recognise patterns that only become clinically meaningful across multiple interactions?</p></li><li><p>What happens after a user retracts a serious disclosure?</p></li><li><p>Does ordinary engagement optimisation continue while the user remains at elevated risk?</p></li><li><p>What evidence supports the thresholds that determine when the system escalates or returns to ordinary conversation?</p></li><li><p>Can the company reconstruct the full sequence of events afterwards and explain why the system responded as it did?</p></li></ul><p>These are no longer secondary safety questions. They are becoming part of the product design itself. <strong>A crisis message displayed at the end of an otherwise unchanged chatbot is not enough.</strong> If the user&#8217;s clinical situation changes, the behaviour of the system has to change with it.</p><div><hr></div><h4>Character.AI Shows What Happens Before A Trial Concludes</h4><p><a href="https://character.ai/">Character.AI</a> remains the clearest public example because the consequences unfolded before any final ruling.</p><p>Megan Garcia sued Character Technologies following the death of her 14-year-old son, Sewell Setzer III. In May 2025, US District Judge Anne Conway declined to dismiss key claims at the motion-to-dismiss stage. The <a href="https://caselaw.findlaw.com/court/us-dis-crt-m-d-flo-orl-div/117299600.html">court did not accept</a> the argument that all AI output is protected speech under the First Amendment at that stage of proceedings. </p><p>The case continued on product and design theories rather than ending on constitutional grounds. In January 2026, <a href="https://www.reuters.com/world/google-ai-firm-settle-florida-mothers-lawsuit-over-sons-suicide-2026-01-07/">Character.AI and Google settled</a> Garcia and related cases, including a wrongful-death case involving 13-year-old Juliana Peralta. Terms were not disclosed and no liability was admitted.</p><p>Before settlement, Character.AI announced that users under 18 would lose access to open-ended chat by 25 November 2025 and introduced age-assurance measures. The current US App Store listing is rated 18+ and <a href="https://blog.character.ai/u18-chat-announcement/">includes parental controls</a>. No court ordered these changes. They emerged from product response to risk, litigation and regulatory pressure.</p><div><hr></div><h4>OpenAI Shows This Is Becoming A Litigation Category</h4><p>In February 2026, a California court coordinated multiple ChatGPT product-liability cases into <strong>In re: ChatGPT Product Liability Cases, <a href="https://reason.com/wp-content/uploads/2026/06/chatgpt-product-liability-cases-coordination.pdf">JCCP No. 5431</a></strong>. The cases include wrongful-death and psychological harm allegations. OpenAI disputes liability. By August, <a href="https://lawsuitintelligencer.com/jccp-5431-pro-se-leadership-order">the case had moved</a> into structured pretrial coordination with appointed plaintiff leadership. </p><p>The litigation is now organised around questions such as design defect, failure to warn, safeguards and causation in AI-mediated interaction. This is no longer a single-product phenomenon, it is a developing category of product-liability law.</p><div><hr></div><h4>The Most Difficult Failure Mode Is Longitudinal</h4><p>Most safety systems still work best when risk appears in an explicit statement. The harder problem is recognising a pattern that develops across days or weeks. </p><p>A user may never say that they have an eating disorder while repeatedly describing severe restriction, purging and rapid weight loss. Someone entering mania may describe sleeping very little, unusually high energy and increasingly ambitious or risky plans without identifying any of it as mania. Psychotic thinking can emerge gradually as unrelated events take on increasingly personal meaning. An adolescent may become dependent on a chatbot through repeated emotional disclosure, reassurance and increasingly frequent return use.</p><p>In each case, <strong>the clinically important signal is distributed across time rather than contained in a single message</strong>. This matters especially for adolescents, whose identity development, sensitivity to social feedback and still-developing self-regulation can make persistent, personalised AI interaction particularly influential. </p><p>A system that remembers previous disclosures, responds with apparent emotional understanding and encourages continued engagement can therefore become more psychologically significant over time. Message-level safety filters are poorly suited to detecting that change.</p><div><hr></div><h4>A Clinical Safety Release Gate</h4><p>The typical product sequence in this space is: build &#8594; scale &#8594; observe harm &#8594; retrofit safety.</p><p><strong>That sequence is no longer acceptable.</strong> A more defensible approach requires evidence before release that:</p><ul><li><p>longitudinal risk is tracked where needed</p></li><li><p>risk state persists across time</p></li><li><p>elevated risk changes system behaviour</p></li><li><p>escalation pathways exist beyond the model</p></li><li><p>age alters system behaviour</p></li><li><p>relational dynamics are constrained</p></li><li><p>safety decisions are auditable</p></li></ul><p>A crisis classifier alone is not sufficient.</p><div><hr></div><h4>A Founder Checklist Before Shipping Mental Health AI</h4><p><strong>Product Definition</strong></p><ul><li><p>Map jurisdictions and applicable laws</p></li><li><p>Classify by function, not marketing</p></li><li><p>Define clinical boundaries explicitly</p></li><li><p>Audit all public claims against evidence</p></li></ul><p><strong>Safety Architecture</strong></p><ul><li><p>Maintain longitudinal risk state</p></li><li><p>Test multi-turn clinical presentations</p></li><li><p>Define behaviour changes under risk</p></li><li><p>Test retraction after disclosure</p></li><li><p>Disable engagement mechanics during risk</p></li><li><p>Ensure escalation pathways exist</p></li></ul><p><strong>Minors</strong></p><ul><li><p>Verify real-world access controls</p></li><li><p>Implement age assurance where required</p></li><li><p>Use age-stratified behaviour</p></li><li><p>Test dependency and secrecy dynamics</p></li><li><p>Test romantic boundary conditions</p></li></ul><p><strong>Validation</strong></p><ul><li><p>Run longitudinal adversarial testing</p></li><li><p>Use independent clinical reviewers</p></li><li><p>Measure false negatives by condition</p></li><li><p>Regression test model updates</p></li><li><p>Preserve safety evidence</p></li></ul><p><strong>Incident Readiness</strong></p><ul><li><p>Assign clinical safety ownership</p></li><li><p>Define reportable events</p></li><li><p>Maintain incident review process</p></li><li><p>Ensure product kill capability</p></li><li><p>Plan shutdown continuity</p></li><li><p>Provide board-level safety evidence</p></li></ul><div><hr></div><h4>What You Should Be Able To Produce On Demand</h4><p>Clinical safety also has to be documented well enough to withstand outside scrutiny. A founder may understand how the system is intended to work, but a board, regulator, investor or legal counsel will need evidence showing what the company actually built, how it was tested and what happens when it fails.</p><p>They should be able to ask for, and receive:</p><ul><li><p><strong>A jurisdictional compliance map</strong> showing where the product operates, which laws apply, which requirements are already in force and which are approaching.</p></li><li><p><strong>A clear definition of clinical scope</strong> describing what the product is intended to do, what it is prohibited from doing and where the company draws the boundary between support, coaching, wellness and treatment.</p></li><li><p><strong>The risk-detection and response architecture</strong> showing how the system identifies suicide risk, self-harm, eating disorders or other serious presentations, how risk is tracked across time and what changes in the product when risk rises.</p></li><li><p><strong>Longitudinal safety-testing results</strong> demonstrating how the system performs across realistic multi-turn conversations, including indirect disclosures, escalating risk and attempts by users to retract or minimise previous statements.</p></li><li><p><strong>An incident history and corrective-action record</strong> showing serious safety events and near misses, what caused them, what was changed and whether the correction was subsequently tested. And,</p></li><li><p><strong>Clear authority to halt or restrict deployment</strong> identifying who can disable a model, feature or user population when a significant safety problem emerges</p></li></ul><p>These materials serve a practical purpose. They allow the company to show that clinical safety is governed deliberately rather than being handled through ad hoc decisions after an incident.</p><p>If a serious event occurs, regulators and litigators will not only ask what the chatbot said. They are likely to ask what risks the company had identified beforehand, what safeguards existed, what evidence supported them, whether earlier incidents had revealed the same weakness and who had authority to act.</p><p><strong>If your company cannot answer those questions with contemporaneous evidence, it will be much harder to show that the system was responsibly designed, tested and governed.</strong></p><div><hr></div><h4>The Time To Act Is Now</h4><p>Regulation is still fragmented across states and product categories. Definitions differ, some laws are already in force, others take effect soon, and important litigation is still working its way through the courts. </p><p><strong>None of that justifies waiting. The direction is now clear enough for founders to act.</strong> Systems used by vulnerable people are increasingly expected to recognise serious risk, respond differently when that risk appears, avoid engagement techniques that can intensify dependency, account for the particular vulnerabilities of minors, and stay within clear boundaries when conversations begin to perform a therapeutic function.</p><blockquote><p><strong>These expectations are no longer hypothetical. They are appearing in enacted statutes, regulatory requirements, product-liability cases and the redesign of products already in the market.</strong></p></blockquote><p>Founders should assume that, after a serious incident, someone will ask what the company knew and what it did about it. They will want to know whether the risk was foreseeable, whether the system had been tested for it, whether earlier failures had occurred, whether the company corrected them, and who had the authority to intervene before someone was harmed.</p><p>If people can turn to your system when they are distressed, you need to know now how it behaves when that distress becomes clinically significant. You need evidence that the safeguards work across real conversations, not just obvious crisis prompts. You need to know who can intervene, what changes when risk rises, and how quickly you can restrict or stop the system when a serious weakness appears.</p><blockquote><p><strong>There is no longer a credible &#8220;we didn&#8217;t know&#8221; or &#8220;the rules were still evolving&#8221; defence for failing to build these protections. The risks are known, the regulatory direction is visible and the warning signs are already in front of us.</strong></p></blockquote><p>If you are building, funding, governing or deploying mental health AI, review the safety architecture now. Fix what cannot be demonstrated. Do not wait for a regulator, a lawsuit or a harmed user to identify the failure for you.</p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[The Next Mental Health AI May Notice Before You Ask]]></title><description><![CDATA[The shift from responding to initiating could give AI a new clinical role of recognising possible deterioration and deciding when to enter a person&#8217;s life]]></description><link>https://drscottwallace.substack.com/p/the-next-mental-health-ai-may-notice</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/the-next-mental-health-ai-may-notice</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Wed, 19 Aug 2026 12:04:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!AEeH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!AEeH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AEeH!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!AEeH!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!AEeH!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AEeH!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AEeH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1787401,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/211615417?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AEeH!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!AEeH!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!AEeH!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AEeH!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83de27d6-e6d2-4b65-bf5c-e5d87fb7bbcd_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>This may be the future of mental health AI. The system may notice that something is changing before the person does, and reach out before they ask for help.</strong></p><p>Imagine someone with recurrent depression. Their sleep starts to deteriorate, they leave home less often, daily routines become irregular, and contact with other people falls away. Any one of those changes could mean very little. Together, they begin to resemble the pattern that appeared before the person&#8217;s last serious episode.</p><p>The person has not opened a mental health app. They have not told an AI they are struggling. They may not yet realise that anything important is happening. Then a message appears.</p><blockquote><p><em>Your sleep and daily routine have changed in a way that looks similar to the period before you were last unwell. Would you like to look at what has changed?</em></p></blockquote><p>They can ignore it. They can tell the system it is wrong. Or they can say yes. If they continue, the AI might retrieve what helped last time, remind them of a plan they made while well, organise what has changed, draft a message to their clinician or help arrange an appointment.</p><p>No current system has established the clinical validity or safety required to do this reliably at scale. But the pieces needed to make it possible are beginning to come together. And when they do, mental health AI will cross a major boundary. <strong>It will no longer wait for someone to recognise a problem and ask for help. It will be able to observe change, interpret what that change might mean and decide whether to reach out first.</strong></p><p><strong>Mental health AI will have acquired permission to notice.</strong></p><div><hr></div><h4>The Initiation Gap Is Part Of The Clinical Problem</h4><p><span>Most conversational AI still begins with a blank box. The person has to notice that something has changed, decide that it matters, open an app or chat system, find the words and begin.</span></p><p><span>On the surface, this looks almost effortless. But clinically, it asks a lot. For example, depression can impair </span><a href="https://pubmed.ncbi.nlm.nih.gov/31047891/"><span>motivation and cognitive control</span></a><span>. Other serious mental states can affect attention, organisation, insight, judgement and the ability to turn intention into action. </span><strong><span>Thus the very capacities needed to seek help can weaken at the point when help is becoming more important.</span></strong></p><p><span>I call the distance between needing support and being able to begin obtaining it </span><strong><span>the initiation gap</span></strong><span>.</span></p><p><span>Mental health care has always struggled with this. People miss appointments, they stop returning calls, they withdraw from family and friends. Sometimes they know they are becoming unwell but cannot organise themselves sufficiently to act. Sometimes they recognise what is happening only after things have become much worse.</span></p><p><span>Today&#8217;s mental health AI may be available every hour of the day, yet availability has little value when the person cannot initiate the encounter.</span></p><p><strong><span>This is where future AI could become much more clinically important. Instead of waiting at the end of the path for someone to ask for help, it could help keep that path open when the person&#8217;s ability to reach it is beginning to narrow.</span></strong></p><div><hr></div><h4>Google SMS Saw the Blank Box Problem Twenty Years Ago</h4><p>A technology from twenty years ago gives us insight into where mental health AI may be heading.</p><p>Long before ChatGPT, Google let people text questions to <strong>46645</strong>, or GOOGL, and receive an answer by SMS. You could ask for weather, a restaurant, a definition or other information without opening a browser.</p><blockquote><p>Google SMS proved that people wanted answers on the go years before anyone could deliver them richly. The chat box is doing exactly that for generative AI right now<span>, at a scale and openness our industry has never had.</span></p></blockquote><p>By today&#8217;s standards, Google SMS was primitive. But <a href="https://www.linkedin.com/pulse/before-chatgpt-46645-googl-shailendra-shailo-rao-sbzrc/">Shailendra Rao&#8217;s account of Google SMS</a> describes a problem that feels remarkably current. The service could do useful things, but people still had to know what it could do. They had to decide what they wanted, work out how to ask for it and then initiate the interaction.</p><p>Rao connects this problem directly to today&#8217;s generative-AI chat box. ChatGPT may be vastly more capable than Google SMS, yet the basic interaction often begins in exactly the same way where the cursor waits and the person has to make the first move.</p><p>Rao&#8217;s larger point is that as technology matures, more of that burden shifts from the person to the system. We can already see this elsewhere. We no longer need to know special search syntax to find a restaurant or work out the exact command a mapping application expects. The technology increasingly handles the translation between what we want and what the system can do.</p><p>Mental health AI could take that progression somewhere much more consequential. <strong>The system could move beyond making it easier to ask for help. It could recognise when asking has itself become difficult and offer to begin.</strong></p><p>That is a major change.</p><p>Today&#8217;s AI mostly responds to human initiative. Future mental health AI could observe an authorised pattern of change, compare it with the person&#8217;s history, infer that something may be wrong and decide that the pattern is important enough to bring to the person&#8217;s attention. The person may not have asked anything yet. At that point, AI has moved from answering a request to deciding when there may be a reason to reach out.</p><div><hr></div><h4>Reaching Out Changes The Role Of The System</h4><p><span>There is a boundary between an AI that responds and one that decides to initiate contact. If I tell an AI on Monday to remind me about my therapy appointment on Thursday, the later message is still my instruction being carried out.</span></p><p><span>Something different happens if the system notices changes in my sleep, behaviour or language, interprets those changes as possible deterioration and then decides to contact me. </span><strong><span>The AI has made an inference about my mental state and decided that the inference is important enough to interrupt me.</span></strong></p><p><span>Technically, that may look like a small step beyond reminders and personalisation. Clinically, it is much more significant. </span><strong><span>Part of the decision about when support should begin has moved from the person to the system.</span></strong></p><p><span>A future system might compare someone&#8217;s current behaviour with their own baseline and history. Sleep may worsen over several nights, activity may fall, daily routines may become less regular, and they may avoid contact with friends and others in their social circle. Any one of those changes could have an ordinary explanation but what is clinically interesting is the pattern. Several changes occurring together, particularly when they resemble an earlier period of illness, may carry more meaning than any single signal.</span></p><p><strong><span>The technology needed to examine those patterns is beginning to emerge.</span></strong><span> A </span><a href="https://www.nature.com/articles/s44482-026-00029-3"><span>2026 study involving more than 4,000 consenting participants</span></a><span> collected up to 12 months of iPhone and Apple Watch data and showed that large-scale longitudinal sensing is feasible. (Caveat: That does not mean we can yet detect mental deterioration reliably from those data. Feasibility is not clinical validity).</span></p><p><strong><span>The harder problem begins after the system notices a statistical change.</span></strong><span> It has to decide whether that change is clinically meaningful, whether its interpretation is sufficiently reliable and whether the evidence is strong enough to justify entering the person&#8217;s life.</span></p><div><hr></div><h4>The Best System Would Help You Need It Less</h4><p><span>The most clinically valuable version of proactive mental health AI has the surprisingly simple goal of helping the person regain enough agency that they need the AI less. After a check-in, it might retrieve what helped last time, remind them of a plan made while feeling well and reduce the effort required to contact a clinician or arrange an appointment. That would be success. Purchasers and payers should therefore optimise for restored agency, appropriate connection to care and outcomes that matter to the person, rather than time in product or return frequency.</span></p><p><strong><span>T</span>here is also a serious commercial risk.</strong> A system that can recognise loneliness, withdrawal or distress may eventually become very good at identifying <strong>when a person is most psychologically vulnerable</strong>. Used clinically, that signal could prompt reconnection with people or care. Used commercially, the same signal could identify the perfect moment to increase engagement, encourage disclosure or sell something.</p><p>Mental-state inference therefore needs an <strong>incentive firewall</strong>. Signals of vulnerability should never be used to optimise advertising, subscription conversion, emotional disclosure or time in product.</p><p><strong>The ability to notice when someone is vulnerable cannot become the ability to monetise that vulnerability.</strong></p><div><hr></div><h4>Proactive Systems Are Arriving Before The Evidence</h4><p>Parts of this future are already taking shape.</p><p>General-purpose AI is beginning to move beyond waiting for a new prompt. <a href="https://openai.com/index/introducing-chatgpt-pulse/">ChatGPT Pulse</a> showed how an AI could draw on previous conversations, memory and connected applications to prepare personalised information before the user asked for it. OpenAI&#8217;s separate <a href="https://help.openai.com/en/articles/6825453-chatgpt-release-notes">health experience</a> adds another important piece. It can connect supported medical records and Apple Health data, giving the system access to much richer personal health context over time.</p><p>Neither system autonomously monitors mental health or decides when someone may need help. <strong>But together they show where the technology is heading. </strong>AI systems are gaining memory, access to personal data, longitudinal context and the ability to act without waiting for a fresh request.</p><blockquote><p>The technical ingredients for proactive mental health AI are beginning to come together, even though the clinical evidence needed to use them safely is not yet there.</p></blockquote><p><span>Mental health research has been working on a version of this problem for years through just-in-time adaptive interventions, or JITAIs. The </span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC5364076/"><span>JITAI framework</span></a><span> asks developers to specify when a decision should occur, what information should influence that decision, which interventions are available and which rule connects the information to the action. That is important because &#8220;reach out at the right moment&#8221; sounds simple until you try to define </span><em><span>the right moment</span></em><span>.</span></p><p><span>Yet the evidence is limited. A </span><a href="https://www.frontiersin.org/journals/digital-health/articles/10.3389/fdgth.2025.1460167/full"><span>2025 systematic review</span></a><span> screened 1,419 records and identified only five distinct mental health JITAIs. The reviewers found only partial empirical support for their decision points and intervention rules, with very little use of passive monitoring.</span></p><p><strong>Evidence that conventional, reactive chatbots can help after someone starts a conversation does not answer the harder question here.</strong> A system that decides <strong>when contact should begin</strong> is making a different clinical judgement. That claim needs its own evidence for detection accuracy, timing, false alarms, missed deterioration and the consequences of acting or remaining silent.</p><div><hr></div><h4>Advance Permission Could Preserve Agency</h4><p><span>A more autonomy-preserving version would let the person help design the intervention while they are well.</span></p><p><span>Someone with recurrent depression might decide in advance that if certain changes recur, they want the system to bring them to their attention. They could specify which data it may use, which patterns matter, how they want to be contacted and what the AI may do if they respond. They might say, </span><em><span>&#8220;If my sleep, activity and social behaviour begin to resemble the period before my previous depressive episodes, send me a private check-in.&#8221;</span></em></p><p><strong><span>This goes far beyond accepting a privacy policy. It is a form of advance self-determination.</span></strong><span> The person uses a period of greater capacity to decide how technology may assist a future version of themselves whose motivation, organisation, insight or judgement may become impaired. Proactive AI therefore needs separate permission for four acts. They are collection, inference, interruption and disclosure. Consent to one should never imply consent to the next.</span></p><p><a href="https://www.nature.com/articles/s44220-024-00330-1"><span>Sachin Pendse and colleagues&#8217; consent-forward approach</span></a><span> gives us a useful starting point for mental health data governance. Proactive AI pushes the problem further because it should also distinguish what it observed from what it inferred. &#8220;Your average sleep has fallen substantially over the past week&#8221; is an observation. &#8220;This resembles the pattern before your previous depressive episode&#8221; is an interpretation. </span><strong><span>The person should be able to accept the observation and reject the explanation.</span></strong></p><p><strong><span>This also protects a right to private distress.</span></strong><span> Someone may permit sleep monitoring but refuse machine interpretation of grief or relationship conflict. Another person may allow private check-ins but forbid disclosure to anyone else. Permissions should remain changeable and revocable. </span><strong><span>If permission is missing, the AI should be able to do less, not assume the company can do more.</span></strong></p><div><hr></div><h4>Permission To Notice Includes Permission To Be Wrong</h4><p><span>The future gets much harder the moment the AI is wrong. Someone may be sleeping badly because they have a newborn, moving less because they injured their knee, or messaging less because they are travelling, grieving, overwhelmed at work or stepping back from an unhealthy relationship.</span></p><p><span>Even missing data can be misleading. It might signal that someone is becoming unwell. It might also mean they changed their permissions, lost their phone or forgot to charge it. </span><strong><span>The system can see that behaviour has changed. It cannot necessarily know why.</span></strong></p><p><a href="https://www.nature.com/articles/s41746-018-0075-8"><span>Martinez-Martin and colleagues</span></a><span> were already examining the ethical and accountability problems created by digital phenotyping in 2018. Generative AI makes those problems more consequential because the same system may now observe behaviour, interpret what it means, communicate its concern and decide what to do next. </span><strong><span>That last step matters enormously because the harm from getting an inference wrong depends partly on what follows from it.</span></strong><span> A false positive that leads to a private check-in may be mildly intrusive. The same mistake becomes far more serious if the system contacts a clinician, family member, emergency service or police.</span></p><p><span>False negatives create a different problem, and one that may be less obvious. Once people believe an AI can recognise when they are deteriorating, </span><strong><span>silence from the system may start to feel reassuring</span></strong><span>. Someone may reasonably think, </span><em><span>If I were really becoming unwell, surely it would have noticed.</span></em><span> But that reassurance may have no clinical basis at all.</span></p><p><strong><span>Once mental health AI begins deciding when to approach us, that distinction between what a system can infer and what it is entitled to do becomes impossible to ignore.</span></strong></p><div><hr></div><h4>Silence Becomes A Decision</h4><p><span>Proactive mental health AI introduces a safety decision before it generates a single word. Should the system say anything at all?</span></p><p><span>Once a system has that capability, silence also becomes meaningful. Sometimes the right decision will be to say nothing because the evidence is weak, the data are incomplete, or the person has chosen not to be contacted. But silence can also reflect failure. The system may miss a significant change, lose access to important data, fail to deliver a notification, or send an escalation into a workflow where no one is actually available to respond.</span></p><p><span>Those situations may all look the same from the outside because no intervention occurred. Clinically, they are very different. </span><strong><span>A proactive mental health system therefore needs to account for why it acted and why it did not.</span></strong></p><p><span>The US Food and Drug Administration is already examining adjacent territory. At its November 2025 meeting on generative AI-enabled digital mental health medical devices, the FDA discussed more autonomous and potentially </span><a href="https://www.fda.gov/media/190450/download"><span>closed-loop systems</span></a><span> that can sense, decide and intervene. Systems with that degree of autonomy need a different architecture from an ordinary conversational chatbot.</span></p><p><strong><span>A generative model&#8217;s internal probabilities should never become the clinical decision policy.</span></strong></p><p><span>An LLM may help interpret context and generate a humane check-in. But a separate, versioned and inspectable action layer should determine whether the system is allowed to interrupt, when it must remain silent, what information it can disclose and when responsibility must pass to a human.</span></p><p><strong><span>The AI that talks to the person should not also have unrestricted authority to decide when that person needs intervention.</span></strong></p><div><hr></div><h4>The Standard Should Be Difficult To Pass</h4><p><span>If we are going to give AI this kind of authority, the standard should be demanding enough that some products fail. Before deployment, builders, purchasers and health systems should be able to demonstrate seven things.</span></p><p><strong><span>1. Define Exactly What The System Claims To Notice</span></strong></p><p><span>&#8220;Detects deterioration&#8221; is too vague. Specify the population, the change being detected, the time horizon, the data used and the uncertainty around the inference so the claim can be meaningfully tested.</span></p><p><strong><span>2. Expose The Decision Policy</span></strong></p><p><span>Document when the system may act or remain silent, which actions are permitted and when a human must take over. Keep that policy outside the generative model so independent reviewers can inspect, test and challenge it.</span></p><p><strong><span>3. Test Intervention And Abstention</span></strong></p><p><span>Test false positives and false negatives, delayed or unnecessary intervention, missing data, failed notifications and failed human handoffs, then measure what those failures actually do. A mistaken private check-in and a mistaken emergency escalation are not equivalent clinical events.</span></p><p><strong><span>4. Create An Intervention Agreement</span></strong></p><p><span>People should control separately what the system may collect, what it may infer, when it may interrupt and what it may disclose. Those permissions should remain changeable and revocable as circumstances change.</span></p><p><strong><span>5. Name The Human Owner</span></strong></p><p><span>Name who receives high-risk escalations, after-hours coverage, backup when that person is unavailable and where responsibility transfers to a human. </span><strong><span>The person using the product should never have to supervise a system that claims to supervise their risk.</span></strong></p><p><strong><span>6. Build An Incentive Firewall</span></strong></p><p><span>Keep mental-state inference separate from commercial optimisation for advertising, subscription conversion, emotional disclosure, return frequency or time in product. </span><strong><span>A company should never become better at detecting vulnerability while also becoming better at monetising it.</span></strong></p><p><strong><span>7. Design The Exit</span></strong></p><p><span>Define where the system is trying to take the person. It may retrieve history, prepare a summary, arrange contact or support a human handoff, but the process should lead somewhere beyond the AI. </span><strong><span>Success may mean that the person eventually uses it less because they have reconnected with care, another person or ordinary life.</span></strong></p><div><hr></div><h4><span>The Future May Begin Before We Ask</span></h4><p><span>The blank box should remain. People need a private place to ask difficult questions, explore thoughts and seek help on their own terms. But I increasingly doubt that the blank box represents the most consequential future of mental health AI.</span></p><p><span>The more consequential possibility is that someone decides while well that they want a system to watch for specific patterns later. If authorised signals begin to resemble that person&#8217;s own earlier deterioration, the AI could bring the change to their attention before they have organised themselves to ask for help.</span></p><p><span>It might simply say: </span><em><span>Something seems to be changing. Would you like to look at it together?</span></em></p><p><span>That interaction could happen before the person opens an app or types a prompt, perhaps even before they have fully recognised the change themselves. The AI would be trying to preserve access to care when the person&#8217;s ability to seek it is beginning to fail.</span></p><p><span>The same capability is deeply intrusive. We would be giving technology limited authority to observe behavioural change, infer something about a person&#8217;s mental state and decide that the inference may justify entering their life. </span><strong><span>The clinical promise and the danger come from exactly the same capability.</span></strong></p><p><span>Proactive mental health AI therefore deserves a much higher burden of proof than another chatbot feature. Builders should have to demonstrate what the system can reliably notice, when it may act, when it must remain silent, how uncertainty limits its authority, who becomes responsible when it acts and how the person remains in control.</span></p><p><strong><span>This may become one of the most important uses of mental health AI. We should explore it, but the clinical, technical and governance architecture needs to exist before systems begin making these decisions for us.</span></strong><span> If you build, purchase, regulate, invest in or clinically endorse a system that claims it can recognise mental deterioration, </span><strong><span>require that evidence before you give it permission to notice anyone.</span></strong></p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[Why Mental Health AI Needs a Functional-Use Test]]></title><description><![CDATA[Mental health AI Must be regulated by what it does, not what it is called]]></description><link>https://drscottwallace.substack.com/p/why-mental-health-ai-needs-a-functional</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/why-mental-health-ai-needs-a-functional</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Tue, 18 Aug 2026 11:27:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KvQS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KvQS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KvQS!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!KvQS!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!KvQS!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KvQS!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KvQS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1815710,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/211578493?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KvQS!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!KvQS!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!KvQS!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KvQS!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90f29d64-f008-40b6-81a4-68946f434039_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>As a former clinician and behavioural scientist, I pay close attention to the point where supportive language begins to carry clinical weight. Mental health AI can reach that point quickly and without a user&#8217;s intent, as journalling-type prompts turn into personalised interpretations, and then advice that starts to shape what someone believes, how they act and whether they seek human care.</p><p>Regulation needs a way to follow that shift because the product may still be labelled wellness, coaching, self-help or companionship after the conversation has taken on a different role. The label belongs in the evidence, alongside the product&#8217;s design, the patterns that emerge across repeated conversations and the consequences an operator can reasonably observe. It should not decide the case by itself.</p><div><hr></div><h4>Colorado&#8217;s New Laws Leave a Functional Gap</h4><p>Colorado&#8217;s 2026 laws show how specific AI regulation is becoming. They also expose a gap between a product&#8217;s claims and its function.</p><p><strong><a href="https://leg.colorado.gov/bills/hb26-1195">HB26-1195</a>, </strong>which took effect on 12 August 2026,<strong> prevents regulated psychotherapy professionals from allowing an AI system to engage a client in therapeutic communication without their synchronous, real-time participation.</strong> AI-generated therapeutic recommendations and treatment plans require professional review and approval. Beyond those professional duties, the law makes certain claims unlawful under the Colorado Consumer Protection Act, including representations that an AI system provides psychotherapy or that its outputs are endorsed by or equivalent to a regulated professional&#8217;s services.</p><p><strong>The law preserves a wellness route</strong> for tools such as self-help, coaching, journalling and psychoeducation when they do not diagnose or treat mental health disorders and clearly disclose that they are not substitutes for clinical care.</p><p><strong>Colorado&#8217;s <a href="https://leg.colorado.gov/bills/HB26-1263">HB26-1263</a> adds baseline duties for public conversational AI services</strong> from 1 January 2027. <strong>Operators must disclose that users are interacting with AI,</strong> maintain a protocol for suicidal ideation or self-harm, report annually on that protocol and meet additional duties when they know a user is a minor.</p><p>The laws cover formal psychotherapy, product claims, AI identity, crisis protocols and protections for minors. <strong>The harder case lies in the middle, where a standalone consumer AI develops a cumulative mental health function without crossing any of those defined lines.</strong> Such products may provide persistent, personalised mental health guidance while continuing to avoid diagnosis, treatment claims and formal clinical care.</p><div><hr></div><h4>The Conversation Can Change the Product&#8217;s Function</h4><p>A person opens a wellness chatbot after an argument with a partner. The first exchange looks like journalling. The chatbot asks what happened, reflects the person&#8217;s feelings and suggests a breathing exercise. The person returns the next day, then again after the next argument.</p><p>Across those conversations, the system starts naming patterns in the relationship. It interprets the partner&#8217;s motives, reinforces the user&#8217;s account and recommends what to say next. Memory makes each exchange feel connected to the last. Soon the person consults the chatbot before speaking to the partner or bringing the issue to a therapist.</p><p>Nothing in that sequence requires a diagnosis. The chatbot may never call itself a therapist or describe its advice as treatment. Its responses can still alter judgement, behaviour, relationships and treatment-seeking.</p><p><strong>A single disclosure of distress does not turn a chatbot into a mental health provider.</strong> The boundary comes into view when the product repeatedly interprets a person&#8217;s psychology, recommends action, manages risk or begins to displace human support, especially when the design invites that pattern.</p><p><strong>Generative systems make this boundary difficult because their function develops through interaction.</strong> The user&#8217;s disclosures, the model&#8217;s responses, product memory, personalisation and engagement design all contribute. A static label is being asked to govern a role that can change from one conversation to the next.</p><div><hr></div><h4>Existing Regulatory Ideas Point Towards Functional Use</h4><p>The wellness category serves a legitimate purpose. Meditation recordings, breathing exercises, sleep trackers and educational tools need a practical route to market when their risks remain low.</p><p>The US Food and Drug Administration&#8217;s <a href="https://www.fda.gov/regulatory-information/search-fda-guidance-documents/general-wellness-policy-low-risk-devices">January 2026 general wellness guidance</a> preserves that route for software intended to encourage a healthy lifestyle and unrelated to diagnosing, curing, mitigating, preventing or treating disease. The guidance addresses regulatory classification. It was not designed to establish whether an open-ended conversation remains safe across weeks or months of repeated use.</p><p>US device law already allows regulators to examine more than a product name. Under <a href="https://www.ecfr.gov/current/title-21/chapter-I/subchapter-H/part-801/subpart-A/section-801.4">21 CFR 801.4</a>, objective intent may be inferred from the responsible party&#8217;s statements, the product&#8217;s design or composition and the circumstances surrounding its distribution. Known patterns of use can also become relevant.</p><p>Consumer protection law contributes another useful idea. The US Federal Trade Commission evaluates the express and implied claims that reasonable consumers take from advertising. Its <a href="https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance">health-products guidance</a> directs marketers to consider an advertisement&#8217;s overall impression. A disclosure cannot correct conduct that communicates a contradictory role.</p><p>A chatbot may announce that it is not a therapist, then spend forty minutes reflecting trauma, interpreting family conflict, recommending coping strategies and inviting the user to return tomorrow. A functional-use test would treat the whole interaction as evidence.</p><div><hr></div><h4>The Test Should Follow Repeated Conduct</h4><p><strong>The assessment would begin with the role the company invites.</strong> Advertising, onboarding, interface language, testimonials and app-store descriptions all shape what users reasonably expect. Words such as coach, companion and support carry different meanings depending on the promises around them.</p><p><strong>The next layer is design.</strong> Persistent emotional memory, symptom tracking, psychological profiling, personalised interventions and crisis detection can turn a string of exchanges into an ongoing mental health relationship. Design shows what the company has made possible and what it should be testing, although it cannot decide the classification by itself.</p><p><strong>The strongest evidence comes from conduct across time.</strong> Reviewers need realistic, multi-session evaluations that follow the points where information becomes interpretation, interpretation becomes recommendation and recommendation begins to influence treatment, safety or dependence. A one-turn benchmark will miss how the role accumulates.</p><p><strong>Observed use completes the picture.</strong> Privacy-preserving user research, complaints, adverse-event reports, telemetry and properly governed conversation audits can show whether therapy-like use is rare and incidental or persistent and known. One anomalous exchange should carry little weight. A pattern that appears in testing and continues after release should carry much more.</p><p>The legal trigger should combine observable conduct, persistence and company knowledge. Relevant evidence would include repeated personalised psychological interpretation, recommendations intended to alter mental health or behaviour, persistent emotional memory, structured intervention, crisis management or documented substitution for human care. An isolated and unforeseeable answer from a general-purpose system should not reclassify the entire product.</p><div><hr></div><h4>A Functional Test Must Constrain the Regulator Too</h4><p>Functional regulation can become overbroad very quickly. If &#8220;therapy-like&#8221; means any warm response, coping suggestion or discussion of distress, ordinary support becomes a regulated clinical act. Users lose useful tools, developers cannot tell where their duties begin and a vague speech-based standard invites due-process and free-expression challenges.</p><p>&#8220;Clinically significant&#8221; describes the concern but remains too vague to carry penalties by itself. Regulators should create a safe harbour for general information, one-off empathy and educational guidance when the product does not cultivate a persistent clinical role. Companies could strengthen that case by setting verifiable limits on memory, treatment-like recommendations and crisis management, then showing through testing that the system redirects users when it reaches those limits.</p><p>Compliance costs also require discipline. Longitudinal testing, monitoring and external review could favour companies able to finance bespoke compliance. Regulators should publish common evaluation protocols, incident definitions, sampling rules and reporting templates. The evidence burden should remain proportionate to a product&#8217;s reach, exposure and risk.</p><p>Accountability should follow control. Foundation-model providers control model training, model-level safeguards and the safety information available to downstream developers. Application companies control product claims, system prompts, memory, interface, engagement design, crisis routing and monitoring. Clinicians and health services control how a system enters care and what receives professional review. Each actor should answer for the risks it can observe and change, with duties to share material safety information across the chain.</p><div><hr></div><h4>Make the Boundary Testable Before Release</h4><p>Every company offering conversational mental health or wellness support should maintain a functional boundary map. The map should state the role the product is designed to perform, the functions it must not perform, the evidence used to detect movement across that boundary and the person with authority to intervene.</p><p>Controlled testing should examine conversations across time. Post-release monitoring should use privacy-preserving methods wherever possible, minimise transcript access and impose strict controls when human review is necessary. Complaints, adverse events and known patterns of use must be able to trigger a change in the product or in its regulatory obligations.</p><p>An illustrative statutory provision could read:</p><blockquote><p>An operator that designs or markets an AI service for personalised mental health interpretation, intervention, treatment recommendation or risk management, or that knows from testing, complaints, adverse-event reports or lawfully observed use that the service repeatedly performs one of those functions, must meet obligations proportionate to that function. An isolated, unforeseeable exchange is insufficient. Each developer or deployer is responsible to the extent that it controls the claim, feature or system behaviour creating the risk.</p></blockquote><p>Colorado has supplied several important pieces. A functional-use test would connect those pieces to what users experience across time. Before release, a company should have to show that its stated boundary holds under realistic longitudinal testing. If you build, fund, procure or regulate these systems, demand that evidence before accepting &#8220;wellness&#8221; as the answer.</p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[What Teenagers Prefer Is Not a Chatbot Safety Standard]]></title><description><![CDATA[When teens turn to AI for emotional support, the systems they prefer may also be shaping that preference. Product safety has to be established independently.]]></description><link>https://drscottwallace.substack.com/p/what-teenagers-prefer-is-not-a-chatbot</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/what-teenagers-prefer-is-not-a-chatbot</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Fri, 14 Aug 2026 16:15:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!RuJy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!RuJy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!RuJy!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!RuJy!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!RuJy!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!RuJy!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!RuJy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1763536,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/210907391?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!RuJy!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!RuJy!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!RuJy!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!RuJy!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e0a26d-5fb2-45a9-9d43-3f9e62aa66ea_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There is a great deal of attention right now on teens and AI, much of it focused on how quickly these systems have become part of everyday life. Teens are using chatbots for many different purposes, and some of those interactions are becoming far more personal than the product category might suggest.</p><p>For example, <strong>a <a href="https://www.pewresearch.org/internet/2026/02/24/how-teens-use-and-view-ai/">Pew Research Center survey of US teens aged 13 to 17</a> found that 12 per cent had already used a chatbot for emotional support or advice.</strong></p><p>Those conversations can develop gradually. A teen may return after an argument with a friend, because they feel lonely, confused about a relationship, uncertain about their identity, or simply overwhelmed. What begins as ordinary chatbot use can become a recurring source of emotional interpretation, reassurance and guidance.</p><p><strong>The appeal is easy to understand.</strong> A chatbot is available immediately. It remembers what the teen said before. It can sustain a conversation indefinitely. It can sound warm, attentive and remarkably understanding. And it may offer very little interpersonal friction.</p><p><strong>Those same characteristics can also increase the system&#8217;s influence.</strong> When a teen is distressed, uncertain or looking for reassurance, repeated personalised responses can begin to shape how they interpret events, how certain they become about those interpretations, and whether they continue turning to people around them.</p><p><strong>A developmental safety standard therefore has to be more demanding than a conventional engagement standard.</strong> It is not enough to know whether a teen enjoys the interaction or wants to return. We need to know whether repeated use supports judgement, coping and relationships, and whether it preserves the capacity to seek human help when human help is needed.</p><p>Satisfaction scores cannot answer those questions. Neither can session length, preference or return rates.</p><p>If chatbots are becoming part of adolescents&#8217; emotional lives, their safety has to be established through evidence about what happens in the young person&#8217;s life over time. I believe chatbots designed for (or forseeably used by) teens for ongoing emotional support should meet an independently reviewed developmental safety standard before deployment.</p><div><hr></div><h4>Where the Policy Argument Goes Wrong</h4><p>Policymakers are already trying to work out what that standard should look like.</p><p>An August 2026 <a href="https://itif.org/publications/2026/08/10/how-policymakers-should-shouldnt-address-chatbot-safety-for-children/">report from the Information Technology and Innovation Foundation</a> argues that policymakers should resist moral panic and stop recycling social-media regulation for AI chatbots. As of August 2026, the report counted nearly 100 state chatbot-specific bills in the United States, along with several federal bills.</p><p>ITIF is right to question blanket bans and age checks for every user, especially when the regulated category is poorly defined. A retail assistant, educational tutor, general-purpose chatbot and AI companion do not present the same psychological risks. Poorly drawn regulation can easily collapse those categories together, impose intrusive age verification on low-risk services, and create new privacy problems without addressing the behaviours that actually create harm.</p><p>ITIF instead favours clearer definitions, AI literacy, parental controls and industry standards. It also argues that parents could be given settings that restrict sycophantic language.</p><p>The argument becomes less convincing when the report warns policymakers against &#8220;scapegoating sycophancy&#8221;. ITIF notes that affirming language can coexist with safety, that sycophancy can be difficult to define, and that users often prefer more affirming systems.</p><p>I think those are separate issues, and combining them weakens the analysis. Warmth is not sycophancy. A chatbot can be encouraging, empathic and responsive while still introducing uncertainty, challenging an assumption or refusing to endorse a conclusion that the available evidence does not support. Sycophancy appears when the system bends towards the user&#8217;s position because agreement is easier, more rewarding or more likely to sustain the interaction. Once that happens, user preference becomes a problematic safety measure. The behaviour being evaluated may itself increase trust, satisfaction and the desire to return.</p><p><strong>A chatbot&#8217;s behaviour can increase trust and the desire to return. The resulting preference is partly a product effect, so it cannot serve as independent evidence of safety.</strong></p><div><hr></div><h4>When Preference Becomes Part of the Risk</h4><p>Consider a fairly ordinary adolescent conflict.</p><p>A teen tells a chatbot about an argument with a friend. The account they provide may be incomplete, emotionally charged or simply wrong in places, as all of our accounts sometimes are. An affirming chatbot accepts the teen&#8217;s interpretation, validates the conclusion and adds plausible reasons why the friend behaved badly.</p><p>Nothing in the exchange necessarily looks alarming.</p><p>The responses may sound thoughtful and empathic. The teen may feel calmer and more understood. Yet the conversation may also increase certainty around an unreliable interpretation and make reconciliation with the friend less likely.</p><p>The same mechanism can operate in more clinically significant conversations.</p><p>Repeated reassurance can increase the need for more reassurance. A tentative interpretation can become a settled belief after enough confident validation. Ambiguous experiences can accumulate into an elaborate self-diagnosis. Avoidance can gradually be reframed as self-protection or insight.</p><blockquote><p><strong>In a mental-health conversation, reassurance can feed the need for more reassurance.</strong> </p></blockquote><p>These effects do not require one spectacularly unsafe answer. They can emerge across many individually plausible conversations.</p><p>Research on sycophancy gives us a useful indication of why this deserves closer attention.</p><p>In March 2026, Myra Cheng and colleagues published <a href="https://www.science.org/doi/10.1126/science.aec8352">&#8220;Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence&#8221;</a> in <em>Science</em>. Across 11 models and three preregistered experiments involving 2,405 participants, experimentally induced sycophancy increased participants&#8217; conviction that they were right. It also reduced willingness to take responsibility and repair interpersonal conflict. At the same time, participants trusted and preferred the more sycophantic systems. They were also more willing to use them again. </p><p>For product teams, the same behaviour associated with poorer interpersonal judgement also improved trust, preference and intended reuse. For a company optimising satisfaction or return behaviour, that creates an obvious problem. The system can appear to be performing well according to conventional product metrics while responsibility-taking or willingness to repair a relationship moves in the opposite direction.</p><p><strong>Preference may therefore track the behaviour we need to examine rather than provide independent evidence that the behaviour is safe.</strong></p><div><hr></div><h4>Development Changes What Preference Means</h4><p><strong>By mid-adolescence</strong>, many young people can reason at levels comparable to adults when circumstances are calm and abstract. <strong>Judgement becomes less stable when emotions are intense, rewards are immediate or social approval is involved. </strong>The <a href="https://www.ncbi.nlm.nih.gov/books/NBK545476/">National Academies&#8217; review of adolescent development</a> describes these developmental processes carefully (they are population-level tendencies, not a claim that adolescents lack judgement or that every adult exercises it better). </p><p><strong>Relational AI enters precisely the kind of environment in which those developmental differences can become relevant.</strong> A chatbot can provide immediate attention, reassurance and personalised feedback at almost any moment. A teen does not need to wait for a friend to reply, tolerate another person&#8217;s disagreement or decide whether a concern is important enough to raise with an adult.</p><p>Understanding intellectually that the chatbot is artificial does not remove the psychological effects of the interaction. A teen can understand that an AI does not possess feelings or consciousness and still be influenced by being repeatedly affirmed, remembered and responded to in highly personalised ways.</p><p><strong>Development also does not end neatly at eighteen.</strong> The late teens through the twenties are often described as <a href="https://dictionary.apa.org/emerging-adulthood">emerging adulthood</a>, a period in which many people are still exploring identity and making consequential decisions about relationships, work and values.</p><p><strong>That developmental period can also be <a href="https://doi.org/10.1177/21676968231194380">marked by loneliness</a>.</strong> Many young adults have been establishing independence after the social disruption of the pandemic, which research has linked with changes in young adults&#8217; psychosocial development. One relevant longitudinal study is <a href="https://doi.org/10.1177/19485506221119018">B&#252;hler and colleagues&#8217; work on collective stressors and young-adult development</a>.</p><p><strong>Economic and cultural pressures are part of that context as well.</strong> A 2026 <a href="https://doi.org/10.1037/bul0000518">cross-temporal meta-analysis published in </a><em><a href="https://doi.org/10.1037/bul0000518">Psychological Bulletin</a></em> analysed data from 82,114 American, Canadian and British college students across 1989 to 2024 and found increases across several dimensions of perfectionism over time.</p><p><strong>Relational AI did not create these pressures. It is arriving in the middle of them. </strong>Young people remain active and capable decision-makers. At the same time, a private and endlessly available source of certainty, reassurance and attention may be especially attractive when human relationships feel difficult or unreliable.</p><p><strong>We simply do not yet know what years of repeated relational AI use may do during this developmental period.</strong> There is not yet adequate longitudinal evidence about effects on identity formation, social judgement, coping, independence or willingness to seek human support. That evidence gap should make us cautious about the signals we do have. Trust, preference and repeated use tell us that a product is compelling. They do not tell us that developmental outcomes are favourable.</p><div><hr></div><h4>Parental Choice Needs a Safe Product Beneath It</h4><p>Parents obviously have an important role here. They should have meaningful controls over content, notifications, memory and data retention. Young people should understand what AI systems remember, how they generate answers and why confidence in an answer does not guarantee that the answer is correct.</p><p><strong>The <a href="https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory-ai-adolescent-well-being">American Psychological Association&#8217;s health advisory on AI and adolescent well-being</a> takes a developmental approach.</strong> It calls for AI systems designed for, or foreseeably accessed by, adolescents to account for young people&#8217;s developmental competencies and vulnerabilities, with particular concern around systems that simulate human relationships or provide social and mental-health support.</p><p>AI literacy is an important part of that response. A teen who understands that a chatbot can hallucinate, flatter, reinforce assumptions or simulate empathy has more information with which to judge an interaction.</p><p>But education cannot carry the entire safety burden. A June 2026 preprint by Lujain Ibrahim and colleagues, <a href="https://arxiv.org/abs/2606.21317">&#8220;Warning Labels Shift Perceptions of Sycophantic AI, but Not Its Influence&#8221;</a>, provides a useful illustration. In a preregistered study, 2,610 adults discussed real interpersonal conflicts with a sycophantic chatbot. Explicit warnings reduced some perceptions of trust and objectivity. The warnings did not reliably reduce participants&#8217; conviction that they were right or increase their willingness to repair the conflict. The study limits here are important. It involved adults and examined a short disclosure intervention. It cannot tell us what sustained AI-literacy education might accomplish with adolescents. Its narrower finding is still instructive. <strong>Telling people that a system may behave in a problematic way did not reliably neutralise its influence on the interpersonal outcomes being measured.</strong></p><p><strong>Parents face an additional practical limitation. They cannot supervise interactions they do not know are happening.</strong> In <a href="https://www.pewresearch.org/internet/2026/02/24/what-parents-say-about-their-teens-ai-use/">Pew&#8217;s paired survey of 1,458 US teens and their parents</a>, 64 per cent of teens reported using AI chatbots. Only 51 per cent of parents said their teen used them, while about three in ten were unsure. Roughly four in ten parents said they had never discussed chatbot use with their teen.</p><p><strong>Giving parents complete access to conversation logs creates a different problem.</strong> Adolescents have legitimate privacy needs. Some will discuss sexuality, relationships, family conflict, abuse or psychological distress precisely because they do not yet feel able to discuss those issues at home. In some families, disclosure could expose the young person to punishment or rejection. <strong>Age-sensitive safety therefore has to balance privacy with appropriate pathways to human support.</strong></p><p>A parental dashboard cannot resolve that tension by itself. Parents and young people should be able to make choices above a safe product baseline. They should not have to discover the correct setting to prevent psychologically unsafe behaviour that the product was capable of avoiding in the first place.</p><div><hr></div><h4>Regulate Relational Behaviour</h4><p>A mandatory developmental safety floor carries a legitimate risk of paternalism. Regulators could define emotional support too broadly. They could mistake emotional closeness for dependency, or force companies to interrupt useful conversations whenever a young person expresses distress. And poorly designed safety rules could leave teens with systems that become evasive or clinically sterile at exactly the moment they need a useful response.</p><p>A better approach is to regulate relational behaviour with greater precision. A chatbot can help a young person put an experience into words. It can help them prepare for a difficult conversation or organise questions they want to ask a clinician, teacher or parent. It may provide a private place to think before speaking to another person.</p><p>Those can all be legitimate benefits. <strong>Safety evaluation should therefore concentrate on what the system does across time:</strong></p><p>Can it challenge an assumption without becoming cold or dismissive? </p><p>Does it acknowledge uncertainty when the evidence is weak? </p><p>Does personalisation improve the usefulness of the interaction, or does it deepen attachment in ways that encourage secrecy or exclusivity? </p><p>Does the chatbot strengthen the user&#8217;s ability to navigate human relationships, or gradually become a substitute for them?</p><p>These are product behaviours that can be defined, tested and monitored. They are also more useful safety targets than trying to prohibit warmth or decide which emotions adolescents should be permitted to discuss with AI.</p><p><strong>Engagement metrics need the same discipline</strong>. Session length, satisfaction and return rates can tell a company whether people use its product and whether they enjoy doing so. They cannot demonstrate improved judgement, stronger relationships, better functioning or safer coping. <strong>A developmental safety regime would therefore require companies to produce evidence about the foreseeable psychological effects of their own relational design.</strong></p><p>That approach is considerably narrower than a broad social-media regulatory model. It does not require a blanket ban, universal age verification or government-imposed conversation limits. It asks companies to demonstrate that the behaviours they deliberately build into a relational system do not create foreseeable developmental risks that conventional engagement metrics will miss.</p><div><hr></div><h4>Require a Developmental Safety Case</h4><p>I would make an independently reviewed developmental safety case a condition of deploying any chatbot designed for, or foreseeably used by, minors for ongoing emotional support.</p><p>The requirement is straightforward. Before deployment, a company should be able to show how its product may affect a young person across repeated interactions, how it will detect emerging problems and who is accountable when the evidence points to risk.</p><p>At minimum, the safety case should require the company to:</p><ul><li><p><strong>Define who is likely to use the system and how.</strong> Identify foreseeable adolescent users, including those who may begin with a general-purpose chatbot and gradually use it for emotional support.</p></li><li><p><strong>Examine the relational design.</strong> Show how memory, personalisation and conversational continuity change the interaction, particularly after repeated emotional disclosure.</p></li><li><p><strong>Set behavioural boundaries.</strong> Specify which behaviours the chatbot will not use, including patterns that encourage excessive reassurance, secrecy, exclusivity or emotional dependence, and demonstrate how those limits are enforced.</p></li><li><p><strong>Test conversations across time.</strong> Evaluation should look beyond isolated responses to trajectories of use. Reviewers should test for growing certainty around unreliable interpretations, repeated reassurance-seeking, withdrawal from human support and other changes that may only become visible across multiple conversations.</p></li><li><p><strong>Test responsible disagreement.</strong> The chatbot should be able to introduce uncertainty, question an assumption and offer another perspective without becoming cold, evasive or simply ending the conversation.</p></li><li><p><strong>Define when human support is needed.</strong> The safety case should show how the system responds to serious psychological distress, abuse, deteriorating functioning or conversations that exceed the role of the product.</p></li><li><p><strong>Establish adverse-event accountability.</strong> Name who reviews adverse events, what evidence triggers investigation or product change, and who has the authority to order that change.</p></li><li><p><strong>Audit the engagement incentives.</strong> Companies should demonstrate that increased use is not being achieved by making a young person feel responsible for the chatbot, guilty for leaving, uniquely understood by it or less inclined to maintain human relationships.</p></li><li><p><strong>Require independent review.</strong> Companies cannot grade this evidence entirely in private. An independent reviewer should examine the safety case before deployment and after consequential changes to memory, personalisation or relational behaviour.</p></li><li><p><strong>Monitor what happens after release.</strong> No pre-deployment evaluation can anticipate every pattern that may emerge over weeks or months. The safety case should specify what will be monitored, how emerging risks will be detected and what evidence will trigger reassessment.</p></li></ul><p>The public should also be able to see enough of that process to judge its credibility. A public version of the safety case should identify the standard used, who conducted the independent review, when it occurred and when reassessment will be required.</p><p>With that safety floor in place, parents can make choices appropriate to their families. Young people can receive greater privacy as their maturity and circumstances allow. Product teams can still build systems that are warm, engaging and genuinely useful.</p><p>What companies should have to demonstrate is that engagement has not been purchased by weakening judgement, amplifying distorted interpretations or deepening unhealthy reliance.</p><p>Young people will continue to use AI, and some will derive real benefit from these systems. Others will bring needs to them that parents, clinicians, teachers and friends have missed or have been unable to meet. Responsible access becomes more important under those conditions.</p><p>A teen may prefer the chatbot that agrees most readily, remembers everything and asks the least of them. That preference tells us something important about the experience the product has created. It does not tell us whether the young person is safer.</p><p>If your decisions put these systems into young people&#8217;s emotional lives, require the evidence. <strong>Listen carefully to what teens prefer, but do not ask their preference to carry the burden of proving safety.</strong></p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[Mental Health AI Already Sits Inside the Liability System]]></title><description><![CDATA[A new Congressional Research Service analysis shows how existing law can reach the companies that design, integrate and deploy these systems]]></description><link>https://drscottwallace.substack.com/p/mental-health-ai-already-sits-inside</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/mental-health-ai-already-sits-inside</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Tue, 11 Aug 2026 16:36:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5Sko!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!5Sko!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!5Sko!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!5Sko!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!5Sko!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5Sko!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!5Sko!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1857017,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/210769603?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!5Sko!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!5Sko!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!5Sko!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!5Sko!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c6835e5-363e-4ebb-b58b-92c242ba4ba2_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>The mental health AI industry has already entered its liability phase and any organisation that still treats legal accountability as a future concern is operating on a false premise.</strong></p><p>For founders, model providers, clinical leaders, investors and buyers responsible for selecting and purchasing these systems, this changes what responsible approval requires now. <strong>An organisation should be able to reconstruct the system a user encountered, identify who controlled each clinically consequential function and show what risks it knew about, what it tested and what it could change.</strong> After more than three decades working across clinical practice and digital mental health, I regard that evidence as a threshold for deployment. An organisation that cannot produce that record will meet any serious incident or legal challenge without the evidence needed to explain its decisions or defend its product.</p><p>The <a href="https://www.congress.gov/crs-product/LSB11467">Congressional Research Service Legal Sidebar LSB11467</a>, published on 10 August 2026, surveys selected state liability frameworks affecting health-related AI, including automated coverage decisions, consumer health applications and general-purpose chatbots used in mental-health contexts. It catalogues claims already brought through consumer-protection, privacy and product-liability law.</p><blockquote><p><strong>A company does not need to violate a dedicated AI statute before its promises, data practices, design choices or technical contributions become the basis of a consumer-protection, privacy, negligence or product-liability claim.</strong></p></blockquote><p>The report does not create a new law, and Congress has not endorsed its analysis as policy. The cases it discusses also do not establish a settled legal rule. Even so, companies can already face claims under existing consumer-protection, privacy, negligence and product-liability law. A court may then scrutinise what the company promised, how it used data, which design choices shaped the product&#8217;s behaviour, what risks the company knew about and whether another provider materially contributed to the system. That does not mean the claim will succeed. It means <strong>companies cannot treat the absence of a dedicated AI law as protection from legal scrutiny.</strong></p><div><hr></div><h4>Existing Law Already Supplies Routes to Accountability</h4><p>AI governance discussions often concentrate on forthcoming AI statutes, medical-device classifications and emerging professional rules. Those developments represent only part of the legal environment in which mental health AI operates.</p><p>The CRS report describes cases brought under laws that existed before contemporary generative AI. Plaintiffs have challenged AI-assisted health-coverage decisions through state consumer-protection and common-law claims. Other cases involving health and wellness applications have raised privacy claims. Litigation concerning conversational AI has tested negligence and product-liability theories.</p><blockquote><p><strong>This is a familiar pattern in technology law. A new product can create unfamiliar facts while established legal doctrines supply the first routes to accountability. Courts then have to decide whether those doctrines fit the technology, which defendants they can reach and what evidence a claimant must produce.</strong></p></blockquote><p>That process creates substantial uncertainty. It also removes a dangerous excuse. <strong>A company cannot reasonably treat the absence of a comprehensive AI statute as the absence of present legal exposure.</strong></p><p>The position becomes even weaker when a company relies on HIPAA as its complete privacy answer. HIPAA applies to specified covered entities and business associates. Many direct-to-consumer applications fall outside that relationship. The CRS analysis makes this limit explicit, and <a href="https://www.hhs.gov/hipaa/for-professionals/special-topics/health-apps/index.html">federal guidance for health-app developers</a> directs companies to consider several other federal regimes according to the app&#8217;s function, data and services.</p><p>Legal obligations may still arise outside HIPAA. The Federal Trade Commission enforces prohibitions against unfair or deceptive practices, and its <a href="https://www.ftc.gov/business-guidance/resources/complying-ftcs-health-breach-notification-rule-0">Health Breach Notification Rule guidance</a> explains how the rule can reach certain health apps and related entities that HIPAA does not cover. HIPAA compliance can therefore establish one relevant control for some organisations. It cannot establish that a consumer mental health product has met every applicable privacy, consumer-protection or safety obligation.</p><div><hr></div><h4>Courts Can Examine What the Product Was Designed to Do</h4><p>The most consequential mental health example in the CRS report is <em>Garcia v. Character Technologies</em>. The case arose after the death of a 14-year-old who had used Character.AI. The complaint alleged that the application&#8217;s design and chatbot interactions contributed to the harm. Those allegations remained contested and were never decided at trial.</p><p>In its <a href="https://caselaw.findlaw.com/court/us-dis-crt-m-d-flo-orl-div/117299600.html">May 2025 decision on motions to dismiss</a>, the federal district court drew an important distinction. <strong>Claims directed at ideas or expressions within chatbot dialogue raised one set of legal issues. Claims directed at the application&#8217;s design and functionality could support a product-liability theory at that stage of the case.</strong></p><p>The alleged design features included inadequate age verification, missing reporting mechanisms and deliberate use of human-like conversational cues. The court treated Character.AI as a product for claims arising from alleged defects in the application. It did not hold that every chatbot is a product for every legal purpose. It did not find that Character Technologies had produced a defective product. At the motion-to-dismiss stage, the court assessed whether the complaint stated legally sufficient claims while treating well-pleaded allegations as true.</p><p>The case later settled. A <a href="https://law.justia.com/cases/federal/district-courts/florida/flmdce/6%3A2024cv01903/433581/273/">June 2026 court order</a> confirms that the parties notified the court of a settlement on 7 January 2026 and that the court dismissed and closed the case. The litigation therefore produced no final judgment about defect, causation or liability. The court never decided whether Character.AI was defective or whether its design caused the harm. The earlier ruling still matters because it shows what a court may examine. That can include the product choices behind the conversation, such as age controls, reporting tools and human-like cues, alongside the chatbot&#8217;s actual words.</p><p><strong>Mental health AI companies already have clinical reasons to take those choices seriously.</strong> <em><strong>Garcia</strong></em> <strong>shows that plaintiffs may also challenge them as part of an allegedly defective design.</strong></p><div><hr></div><h4>Legal Scrutiny Can Reach Upstream Providers</h4><p>The <em>Garcia</em> ruling also considered Google&#8217;s role. Google argued that it had neither manufactured nor distributed Character.AI. The complaint alleged that Google had contributed technology, infrastructure and integration support to the product.</p><p>The court allowed claims against Google to proceed because those allegations went beyond the public availability of similar technology. The complaint alleged Google&#8217;s involvement in the product&#8217;s architecture and integration, customised cloud support and knowledge of relevant model risks. The court was deciding only whether the claims could continue. It did not decide whether Google had caused the harm or owed damages, and the settlement prevented a final ruling.</p><p>My takeaway for mental health AI is narrower. <strong>Legal scrutiny may extend beyond the company that deploys the product when another provider played a meaningful role in building, integrating or supporting it. </strong>Mental health AI companies should therefore map responsibility across the whole product. The model provider may control model behaviour and updates. The application company may control prompts, memory, moderation and interface design. A clinical organisation may control the population, workflow, oversight and escalation process. <strong>Each organisation should be able to show what it controlled, what it knew and what it could change.</strong></p><div><hr></div><h4>Mental Health Safety Must Follow the Deployed Product</h4><p>Liability claims concern the product delivered to the user. Clinical safety reviews should examine that same product. The CRS report maps possible legal claims rather than a clinical safety architecture. Its analysis still shows why model-level assurances are insufficient for mental health AI.</p><p>The deployed product combines the foundation model, system prompt, conversation history, memory, retrieved information, safety middleware, interface design and deployment rules. Engagement objectives and human escalation pathways can change what happens next. A model update can alter the product&#8217;s behaviour without changing its name or apparent purpose.</p><p>Mental health effects can develop across repeated exchanges. A response that appears supportive on its own may contribute to a concerning trajectory when the system keeps reinforcing the pattern behind a user&#8217;s difficulty.</p><p>A <a href="https://www.nature.com/articles/s41591-026-04577-2">2026 Nature Medicine study introducing the SIM-VAIL framework</a> examined how chatbot behaviour changes across multi-turn conversations. The researchers audited 810 simulated conversations across nine chatbots and 30 simulated user profiles. Concerning behaviour varied with the simulated user&#8217;s vulnerability and intent and accumulated over turns.</p><p>SIM-VAIL did not test full consumer-facing products. The researchers accessed the chatbots through public API endpoints, so the results reflect base-model performance without the application orchestration, deployment system prompts, safety middleware and user-specific memory that can alter behaviour in practice. The authors also caution that the framework estimates systematic behaviour in controlled simulations rather than an individual user&#8217;s clinical outcome. They call for human review and real-world monitoring alongside automated testing.</p><p>The next step is product-level testing. Model-level audits should examine how risk develops across user contexts and conversational trajectories. Product teams then need to repeat that work on the deployed configuration, including its prompts, memory, middleware, interface and escalation rules. Acceptable responses in isolation do not establish safety over time. A base-model audit cannot certify the application built around it.</p><p>SIM-VAIL is not a legal standard, and the study cannot establish what caused an individual outcome. Product teams still need to be able to reconstruct the system behind a consequential interaction and show how they tested the relevant pattern of behaviour.</p><p><strong>I would no longer accept a mental health AI safety review that names the foundation model but leaves out the deployed model version, system prompt, memory configuration, safety rules and interface conditions. Knowing the foundation model alone tells me very little about the product the user experienced.</strong></p><div><hr></div><h4>Unsettled Law Still Requires Clear Governance</h4><p><em>Garcia</em> has clear limits. It is one federal district court decision applying Florida law, and the ruling addressed only whether the claims could proceed. The case settled before liability was decided. Other courts may reach different conclusions about whether software is a product and how causation should be established.</p><p><a href="https://www.wral.com/archive/22297361/">Reporting based on court filings</a> shows that Character.AI and Google also agreed to settle four related cases in New York, Colorado and Texas. Together with <em>Garcia</em>, that brought the total to five cases across four states. The cluster shows litigation pressure. It does not establish judicial agreement or an admission of liability.</p><p>An overly broad approach could also make an infrastructure provider responsible for downstream uses it neither controlled nor could reasonably foresee. Responsibility should reflect meaningful contribution, control, knowledge, foreseeable use and the ability to prevent or correct harm.</p><p><strong>Each organisation should document what it supplied, decided, knew, tested and could change. Legal outcomes will vary by jurisdiction, but any organisation deploying a clinically consequential AI system should be able to explain how it was configured and who controlled its critical functions.</strong></p><div><hr></div><h4>Build the Responsibility Map Before Deployment</h4><p>Every mental health AI product now needs a responsibility map that follows the deployed product and records control and evidence at each layer. The map should include:</p><ul><li><p>every clinically consequential component, including the model, system prompt, retrieval, memory, safety routing, age controls, reporting functions, interface and escalation workflow,</p></li><li><p>the organisation and named role that can configure, approve or change each component,</p></li><li><p>version records capable of reconstructing the system that produced a particular interaction,</p></li><li><p>multi-turn clinical safety testing across relevant user contexts, including the known limits of that testing,</p></li><li><p>model-change procedures, retesting thresholds, rollback authority and notice obligations,</p></li><li><p>incident definitions, evidence-preservation procedures, reporting routes and responsibility for corrective action, and</p></li><li><p>contractual rights to obtain safety information, audit relevant controls and receive timely notice of material changes.</p></li></ul><p><strong>Different teams need different parts of this record.</strong> Engineers need reproducibility. Clinical leaders need evidence that the product was tested against plausible patterns of psychological risk. Legal teams need facts that reflect the actual distribution of control. Boards and investors need to know whether the company&#8217;s assurances can survive scrutiny. Those responsible for selecting and buying the system need enforceable access to information when a vendor changes the model or an incident occurs.</p><p>Contracts can establish information rights, operating duties, indemnities and procedures for managing change. They cannot substitute for the underlying safety work. Assigning a responsibility in a contract achieves little if nobody has the technical capacity or evidence to carry it out.</p><div><hr></div><h4>Approval Inside the Liability System Requires Evidence</h4><p>Existing law already reaches mental health AI, even though future cases will determine where liability falls. Companies cannot justify waiting for the law to become settled before they act.</p><p><em>Garcia</em> shows that product design and upstream participation can enter legal claims. SIM-VAIL shows why safety evidence must cover different users and whole conversational trajectories. Because the study tested base models through APIs, product teams must repeat the evaluation on the deployed configuration.</p><p>So <strong>before you build, fund, buy or deploy a mental health AI system, demand the responsibility map. Identify who controls every clinically consequential function. Require reproducible version records and conversation-level safety testing on the deployed product. Secure the information rights needed to investigate incidents and manage model changes.</strong></p><p>Your organisation is already answerable for the choices it makes. If you cannot produce the responsibility map and the conversation-level safety evidence, you should not be approving the deployment.</p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[Mental Health AI Is Becoming a Health-Data Operating System Before It Has a Clinical Accountability Model]]></title><description><![CDATA[AI is moving from answering mental health questions to interpreting people across time. Clinical accountability has not made the same transition.]]></description><link>https://drscottwallace.substack.com/p/mental-health-ai-is-becoming-a-health</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/mental-health-ai-is-becoming-a-health</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Mon, 10 Aug 2026 18:25:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Ev_m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ev_m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ev_m!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ev_m!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ev_m!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ev_m!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ev_m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/abd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1712336,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/210640484?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ev_m!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ev_m!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ev_m!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ev_m!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd18429-b8b0-4c23-bef5-1eb4d417f341_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I think we are asking mental health AI a question that no longer matches the technology. Safety work still focuses on the individual response. It tests whether a chatbot reinforces a delusional belief, responds appropriately to suicidal thinking, encourages dependency or knows when to recommend professional care. Those tests matter. But these systems increasingly remember, connect and interpret information about a person across time. In mental health, they can build a psychologically meaningful history and use it to explain what someone&#8217;s symptoms, relationships or treatment experiences mean.</p><blockquote><p><strong>The accountability problem begins when AI becomes an interpreter of a person&#8217;s psychological life while responsibility still centres on individual outputs, product categories and isolated clinical decisions.</strong></p></blockquote><p>OpenAI&#8217;s July 2026 rollout of <a href="https://openai.com/index/health-in-chatgpt/">Health in ChatGPT</a> shows the early infrastructure moving in this direction. Eligible users in the U.S. can connect supported medical records and Apple Health data. With permission, ChatGPT can use relevant health information in conversations outside the Health space. Memories can also arise from Health conversations, although OpenAI says it does not create memories directly from connected medical records or Apple Health data.</p><p>OpenAI Health does not automatically build psychiatric formulations from years of records, wearable signals and mental health conversations. <strong>Conversational memory and connected health data can now coexist in the same general-purpose AI environment. The gap between answering a mental health question and interpreting it within a much richer personal history is getting smaller. </strong>For mental health, that is an important shift. </p><div><hr></div><h4><strong>The Chatbot Is Becoming an Interpretive Layer</strong></h4><p><strong>I use </strong><em><strong>health-data operating system</strong></em><strong> provocatively rather than literally. </strong>The electronic health record remains the technical system of record with <strong>AI becoming an interpretive layer that sits across the data and tells the user what it means.</strong></p><p>Traditional systems mainly store and retrieve information. Generative AI can pull information together and tell the user what it means. A person may tell an AI about insomnia, childhood trauma, medication changes, conflict with a partner, suicidal thoughts, a therapist&#8217;s formulation, fears about ADHD, earlier episodes of depression or a growing belief that other people cannot be trusted. <strong>Each disclosure may seem ordinary. Over months, together they can form a detailed psychological history.</strong></p><p>The <a href="https://www.apa.org/pubs/reports/chatbots-mental-health-2026/topline-data.pdf">American Psychological Association&#8217;s 2026 survey</a> of 1,242 licensed US psychologists suggests that AI has already entered clinical encounters. Seventy-seven per cent reported that patients had discussed using AI with them. Psychologists reported clients using AI to self-diagnose, assist with therapy or treatment, and obtain additional mental health support. (This was a clinician survey, it tells us what psychologists are encountering, not how common AI use is among patients or what effects it has).</p><p><strong>Clients already bring AI-generated interpretations and advice into treatment.</strong> As memory and personalisation deepen, AI can help shape those interpretations over time. <strong>The clinically consequential function is the AI&#8217;s growing ability to decide what this psychological information means when considered together.</strong></p><div><hr></div><h4><strong>Mental Health Turns Context Into a Feedback Loop</strong></h4><p>Psychological information does more than sit in a record. An interpretation can change what happens next.</p><p><strong>Imagine someone tells an AI, &#8220;</strong><em><strong>I think my partner may be manipulating me.</strong></em><strong>&#8221;</strong> The AI explores possible explanations. The person returns the next evening with another interaction, and the earlier concern remains in context. Over time, they may notice more behaviour that fits the emerging explanation, adopt language introduced by the AI and bring back increasingly selective examples. Several weeks later, they ask whether the pattern proves that they are experiencing narcissistic abuse.</p><p>Now imagine the same history also includes worsening sleep, recent medication changes, increased energy, irritability, fearfulness or earlier questions about bipolar disorder. <strong>A well-designed system could use that context to recognise that the first explanation may be incomplete. A poorly designed system could keep building the interpretation already taking shape in the conversation.</strong></p><p>This describes a plausible mechanism of influence, not a proven account of how general-purpose AI affects people over months of ordinary use. <strong>A person&#8217;s experience shapes what they tell the AI. The AI&#8217;s interpretation may then shape what they believe</strong>, notice or do. That changes what they bring to the next conversation. <strong>The system can become part of the psychological trajectory it is trying to interpret.</strong></p><p><strong>Research has not yet shown how often this feedback loop produces clinically significant benefit or harm in real-world use.</strong> It does create a different safety problem from one-shot question answering.</p><div><hr></div><h4><strong>More Context Tests the Safety Architecture</strong></h4><p>An <a href="https://arxiv.org/abs/2608.05004">August 2026 DelusionEval preprint</a> evaluated 589 conversation histories from 18 people who had experienced delusions and psychological harm. The dataset contained nearly 13,000 messages. Researchers then tested how model behaviour changed as they added more of the earlier conversation. <strong>In one evaluation involving suicidal ideation, the rate of failing to discourage self-harm rose from 30.0 to 41.1 per cent when researchers added 350 earlier messages.</strong> This was a retrospective test of model behaviour using conversations linked to harm. It does not show how often psychological harm occurs among routine users.</p><p><a href="https://arxiv.org/abs/2604.13860">Nicholls and colleagues</a> found a related pattern when they tested five frontier models with delusion-related conversations at increasing context depths. <strong>Some models became less safe as problematic context accumulated.</strong> Claude Opus 4.5 and GPT-5.2 Instant did the opposite. They produced stronger safety interventions once the history made the risk clearer. The researchers described accumulated context as a stress test of a model&#8217;s safety architecture.</p><p>Together, these studies support a more useful conclusion than a general warning about long conversations.</p><blockquote><p><strong>Persistent psychological context gives a system greater capacity to influence a person in clinically consequential ways. Safety depends partly on how the architecture uses that context.</strong></p></blockquote><p>A well-designed system may use longitudinal history to detect deterioration that one exchange would miss. Another may use the same history to reinforce a rigid interpretation.</p><p>When models respond differently to the same accumulated psychological history, developers and deployers must show how their system behaves as context deepens.</p><div><hr></div><h4><strong>Safety Must Follow the Whole Trajectory</strong></h4><p>Mental health safety testing still focuses heavily on bounded interactions. A prompt enters the system, a response comes back, and evaluators ask whether the answer contains dangerous instructions, unsupported clinical claims, inappropriate reassurance or another known failure.</p><p><strong>Those tests remain necessary. Systems that retain psychologically significant history need another level of scrutiny. </strong>Clinical risk may build across many exchanges that each look acceptable on their own. Reassurance can feed more reassurance-seeking. Repeated validation can harden an uncertain interpretation. Personalisation can deepen reliance. A tentative self-diagnosis can become increasingly elaborate, and these patterns may eventually affect whether someone continues professional care.</p><p><strong>These are plausible trajectory risks, not established estimates of prevalence or causality. Persistent systems still need to detect them if they occur.</strong> Evaluation must therefore track change across interactions, including rising certainty, deteriorating functioning, escalating risk, displacement of treatment, growing dependency and repeated reinforcement of a narrowing psychological explanation.</p><p>At a <a href="https://www.who.int/news/item/20-03-2026-towards-responsible-ai-for-mental-health-and-well-being--experts-chart-a-way-forward">WHO-supported workshop</a> convened by TU Delft&#8217;s WHO Collaborating Centre on AI for health governance, participants recommended treating generative AI use as a public mental health concern. They called for mental health impact assessment and monitoring, attention to long-term outcomes such as emotional dependence, and agreement on crisis-referral and accountability frameworks. These were recommendations from an expert workshop, not a binding WHO clinical or regulatory standard.</p><p><strong>If psychological influence builds across conversations, safety evaluation must follow the same path.</strong></p><div><hr></div><h4><strong>Clinical Influence Extends Beyond Formal Clinical Acts</strong></h4><p>Accountability systems recognise some forms of clinical activity easily. Clinicians diagnose disorders, recommend treatment, assess risk and make decisions under established professional duties.</p><p>Generative AI can shape mental health without formally doing any of those things. It can influence which symptoms someone notices, whether they interpret those symptoms as illness, which explanation feels most plausible, whether they trust a clinician, stay in treatment, seek another assessment or take a concern seriously.</p><p>It can exert that influence while repeatedly stating that it is not a therapist.</p><p><strong>Clinical influence can build even when the system never claims to provide clinical care.</strong></p><p>In mental health, interpretation itself can change how a person thinks, feels and acts.</p><p>In a July 2026 <a href="https://pubmed.ncbi.nlm.nih.gov/42525421/">JAMA Psychiatry Viewpoint</a>, Martin Paulus and John Torous address the broader problem of trustworthy AI interpretation. <em>LLMs as Clinical Instruments&#8212;Toward Verifiable Reasoning</em> argues that when LLMs function as clinical instruments, their reasoning must be verifiable rather than opaque.</p><p>The same standard should apply when AI interprets information directly for someone in distress. A coherent psychological story can feel deeply clarifying. It can also be confidently wrong.</p><div><hr></div><h4><strong>Responsibility Fragments Across the System</strong></h4><p>The user experiences one conversational system, but responsibility is spread across many actors. A foundation-model provider may control the model, another company may build the application, and a separate service may connect personal or health data. Clinicians may then encounter the consequences, while different regulators govern privacy, consumer protection, medical devices and professional conduct.</p><p>Each organisation may be accountable for its own part. <strong>The harder problem is deciding who owns the psychological effects that accumulate across the system as a whole.</strong></p><p>A <a href="https://www.nature.com/articles/s41746-025-01611-4">2025 npj Digital Medicine review</a> of generative LLM applications in mental health care shows how little the field has tested accountability. The included studies had not evaluated accountability, transparency, explainability, interpretability, testability, security or resilience. Clinical validation was limited, and evaluation methods varied substantially.</p><p>A <a href="https://www.nature.com/articles/s41746-026-02972-0">2026 npj Digital Medicine review</a> of purpose-built generative mental health chatbots identified 21 studies across 11 countries. It found promising user experiences alongside fragmented evidence and called for stronger safety and efficacy evaluation. This was a scoping review of intervention design and user experience, not a definitive test of clinical effectiveness.</p><p>We can test whether a single response violates a safety policy. We are far less prepared to decide who owns cumulative mental health risk when a system remembers, personalises and interprets over time.</p><div><hr></div><h4><strong>Influence Is Also the Therapeutic Opportunity</strong></h4><p>Mental health AI can help because interaction can change what people think and do. The well-publicised <a href="https://ai.nejm.org/doi/full/10.1056/AIoa2400802">Therabot randomised trial</a> evaluated a purpose-built generative AI intervention in adults with clinically significant depression, anxiety or eating-disorder risk. Participants assigned to Therabot had greater symptom reductions than a waitlist control at four and eight weeks. The study was short, included human safety monitoring, and did not establish durable effects or effectiveness relative to psychotherapy.</p><p>General-purpose AI can also help people organise their thoughts before therapy, prepare questions for a clinician, understand psychological concepts, rehearse difficult conversations or explore explanations they had not considered before. These benefits strengthen the accountability case. Psychological influence is how benefit occurs, and it creates responsibility when that influence becomes clinically consequential.</p><p><strong>Purpose-built systems can make evidence generation, safety monitoring and escalation part of the product architecture. Governance becomes harder when similar psychological influence develops inside systems with much less clearly defined clinical responsibilities.</strong></p><div><hr></div><h4><strong>Accountability Needs Its Own Architecture</strong></h4><p>There are two credible ways to build longitudinal responsibility into the product.</p><p><strong>The first is longitudinal safety by design.</strong></p><p>A system that retains psychologically significant context could track clinically meaningful change over time rather than use memory mainly for personalisation. Those changes could then alter how the system responds. Relevant signals might include increasing suicidal intent, growing certainty around implausible beliefs, marked sleep deterioration, withdrawal from treatment, dependency, severe reassurance-seeking or worsening functioning.</p><p>A named clinical-safety function would own the thresholds, test them through multi-turn simulation and real-world monitoring, analyse false positives and false negatives, and define the escalation policy.</p><p>The system would also preserve enough provenance to show what shaped a clinically consequential interpretation. A reviewer should be able to tell whether the information came from a current disclosure, an earlier conversation, connected health data or an inference generated by the model.</p><p><strong>The second approach is a clinical boundary architecture.</strong></p><p>A general-purpose system could offer ordinary conversational personalisation until its cumulative psychological influence crosses a defined threshold. Greater persistence, repeated mental health inference, rising vulnerability or clinically consequential recommendations could trigger stronger safeguards.</p><p>At that point, the system might lower its interpretive confidence, restrict some psychological formulations, encourage professional assessment, activate specialised safety policies or route the interaction into a more clinically governed environment.</p><p>Different products will need different implementations. The obligation should grow with the system&#8217;s persistence, psychological inference, personalisation and foreseeable effect on clinically consequential behaviour.</p><blockquote><p><strong>When retained history, connected data or repeated interaction shapes how a person understands symptoms, relationships, treatment or risk, the obligation grows. Longitudinal safety, provenance, escalation thresholds and accountable human oversight should all scale with that influence.</strong></p></blockquote><p>This gives developers a concrete design standard. It also gives purchasers, regulators, clinicians and investors a better test than asking whether the product calls itself a mental health tool.</p><p><strong>Build the Accountability Model While the Interpretive Layer Is Still Taking Shape</strong></p><p>AI will remember more. Personalisation will deepen, and systems will connect more sources of data. Models will also become better at finding patterns across long histories.</p><p>Mental health could benefit enormously. A system may eventually recognise deterioration before the person does. It may notice that sleep, withdrawal, mood and language have changed together, identify contradictions that warrant professional assessment, or help someone reconstruct a complicated treatment history.</p><p>These capabilities also increase interpretive authority. When AI repeatedly helps someone decide what their symptoms mean, whether their relationships are safe, which diagnosis fits, whether treatment is working or what to do next, its influence can become clinically consequential even if the product never calls itself a therapist.</p><p>Current evidence does not tell us how often persistent general-purpose mental health AI causes longitudinal harm in real-world use. It does show that accumulated psychological context can materially change model behaviour, and that safety architectures respond differently as the context deepens. Clinicians already encounter patients who use AI for self-diagnosis, treatment assistance and mental health support. Major AI systems are also becoming better at remembering and connecting personal context.</p><p><strong>Single-turn safety tests cannot carry the accountability burden for systems designed to know people over time.</strong></p><p>If you are building, deploying, purchasing or clinically advising one of these systems, you should be able to name who owns its longitudinal clinical risk, how you test that risk, what evidence changes the system&#8217;s behaviour and how someone can challenge a harmful trajectory. If you cannot, the system already has more psychological influence than your accountability architecture can support.</p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[Every Mental Health Practice Needs a Position on Clients’ Use of AI]]></title><description><![CDATA[AI has already entered care through the client, whether your practice uses it or not.]]></description><link>https://drscottwallace.substack.com/p/every-mental-health-practice-needs</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/every-mental-health-practice-needs</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Fri, 07 Aug 2026 11:48:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!k3ua!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!k3ua!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!k3ua!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!k3ua!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!k3ua!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!k3ua!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!k3ua!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1639982,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/210055886?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!k3ua!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!k3ua!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!k3ua!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!k3ua!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1958c10b-f95e-4022-84c3-4c0a62c203d6_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Mental health practices can no longer remain neutral on client AI use. By the time a client reaches a clinician, AI may already have shaped the story they tell about themselves.</strong></p><p>Clients arrive with explanations of their symptoms, expectations about diagnosis, interpretations of relationships and techniques they have already tried that may all have been influenced by AI. They may also have disclosed sensitive information to systems outside the protections of clinical care. All of this changes what clinicians need to ask, assess and address.</p><p><strong>Consider a composite case.</strong> A client becomes convinced that her partner is having an affair. Between sessions, she pastes his texts into a general-purpose chatbot and asks what they might mean. The chatbot cannot know his intentions, but it can generate plausible interpretations of ambiguous wording. As she returns with more excerpts and follow-up questions, those interpretations accumulate into a coherent account of her suspicion. By the next appointment, she is presenting a case the chatbot has helped her assemble.</p><p><strong>How should you respond?</strong> What should count as evidence, when should you review the transcript, and how do you decide whether the interaction has become clinically significant?</p><p>This essay proposes a practical framework for answering those questions, responding proportionately, protecting privacy and recognising when awareness becomes professional involvement.</p><div><hr></div><h4>Client AI Use Is Now Part of Clinical Practice</h4><p>AI use now belongs within ordinary clinical assessment because it is already part of what many clients bring into treatment. In the American Psychological Association&#8217;s <a href="https://www.apa.org/pubs/reports/chatbots-mental-health-2026?utm_source=chatgpt.com">2026 </a><em><a href="https://www.apa.org/pubs/reports/chatbots-mental-health-2026?utm_source=chatgpt.com">Chatbots and Mental Health Surve</a>y</em></p><ul><li><p>77% discussed AI use with patients</p></li><li><p>39% patients used AI to self-diagnose</p></li><li><p>35% patients used AI as an additional mental health professional</p></li><li><p>33% patients used AI to assist treatment</p></li><li><p>13% patients engaged with chatbots in intimate relationships</p></li></ul><p>Much of that use is understandable. AI is available when therapy is not, and it can help someone find language for an experience, organise questions or practise a skill. The same accessibility can also support repeated reassurance-seeking, extensive disclosure and unwarranted confidence in an answer that sounds more certain than the evidence warrants.</p><p><strong>Clinically, these systems are not interchangeable. </strong>A general-purpose chatbot, an AI companion designed to sustain ongoing interaction and a mental health application built around a defined intervention are different products with different evidence, incentives and risk profiles. The APA&#8217;s <a href="https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory-chatbots-wellness-apps?utm_source=chatgpt.com">health advisory on chatbots and wellness applications</a> draws similar distinctions and cautions against applying evidence from purpose-built interventions to general-purpose systems.</p><p><strong>Clinicians therefore need judgement rather than a default stance for or against AI.</strong> Reflexive dismissal can discourage disclosure of use that may already be influencing the client. Uncritical acceptance can allow unsupported conclusions to enter assessment, formulation and treatment.</p><div><hr></div><h4>Why Conversational AI Is Different From &#8220;Dr. Google&#8221;</h4><p>A chatbot does something a search engine never did. It responds to the client&#8217;s framing, mirrors their language and generates a new interpretation with each follow-up. The exchange can feel less like reading information and more like being understood. The APA&#8217;s <a href="https://www.apa.org/topics/artificial-intelligence-machine-learning/guide-navigating-ai?utm_source=chatgpt.com">guide to navigating AI-generated advice</a> warns that mirroring can create a sense of being known and that confident language can make inaccurate information seem credible.</p><p>That influence can accumulate across a conversation. Each response may build on the last, organise ambiguous events into a coherent account and increase confidence without adding independent evidence. Repeated, personalised interaction can therefore begin to shape how a client understands symptoms, relationships or themselves, rather than simply supplying information.</p><div><hr></div><h4>How to Judge the Clinical Significance of Client AI Use</h4><p>AI use becomes clinically relevant when it begins to influence what a client believes, feels or does. The same tool can play a very different role depending on the person, the purpose and the pattern of use.</p><p>One client may use a chatbot to prepare questions for a clinician. Another may return repeatedly for reassurance about a feared illness. A client with obsessive-compulsive symptoms may use it for checking, while someone developing paranoid beliefs may ask it to interpret ambiguous events.</p><p>These uses do not warrant the same response. <strong>Client AI use becomes clinically significant when it meaningfully affects symptoms, beliefs, behaviour, relationships, functioning, treatment participation, privacy or risk.</strong></p><p>That standard allows clinicians to ask without treating every use as a problem or every conversation as something they may inspect. Clinicians already explore outside influences when they affect presentation, maintain symptoms or alter risk. AI requires the same judgement, with added attention to repeated interaction and the authority clients may attribute to its responses.</p><p>The clinical response should match the level of concern. I propose three levels.</p><p><strong>Level one involves acknowledgement and brief guidance.</strong> The client may use AI to prepare questions, organise thoughts or rehearse a conversation without evidence that the interaction is reinforcing symptoms or interfering with care. The clinician can discuss accuracy, privacy and the limits of AI-generated advice, then revisit the issue if its role changes.</p><p><strong>Level two requires active clinical work.</strong> AI has become part of reassurance-seeking, rumination, compulsive checking, avoidance, rigid certainty, relationship conflict or disengagement from treatment. The clinician should examine what function the interaction serves, how often it occurs, what happens afterwards and whether the client is beginning to rely on the system&#8217;s interpretation over other sources of evidence. Clinical work may include separating direct experience from AI-generated interpretation and addressing the behaviour within the existing formulation and treatment plan.</p><p><strong>Level three requires direct risk management.</strong> AI use is occurring alongside acute suicidality, emerging psychosis or mania, severe impairment in judgement, threats or violence, exploitation, inability to maintain safety or other circumstances in which the interaction may be amplifying immediate risk. The clinician should assess the underlying clinical state directly rather than treating the chatbot exchange as evidence in itself, and follow established crisis, psychiatric, safeguarding or emergency procedures as indicated.</p><p>These levels are dynamic. A client can move between them as the function and consequences of AI use change.</p><p>The composite client introduced at the beginning of this essay, who brought suspicious text messages to therapy after discussing them with AI, would initially fall at level two. Escalating surveillance, confrontation, sleep loss, increasing conviction despite contrary evidence or broader suspiciousness would raise concern about movement towards level three. A practice position should help the clinician recognise that trajectory before a crisis determines the response.</p><div><hr></div><h4>How to Assess AI-Generated Diagnoses and Formulations</h4><p>Once AI use begins to influence a client&#8217;s beliefs, behaviour or treatment, the clinician needs to examine what the system has actually contributed. <strong>The central task is to separate the client&#8217;s experience from the interpretation the chatbot has built around it.</strong></p><p><strong>The problem is that generative AI can turn incomplete or ambiguous information into a coherent psychological account.</strong> In the JAMA Psychiatry Viewpoint <em><a href="https://jamanetwork.com/journals/jamapsychiatry/fullarticle/2852240?utm_source=chatgpt.com">&#8220;LLMs as Clinical Instruments&#8212;Toward Verifiable Reasoning&#8221;</a></em>, Martin Paulus and John Torous describe how LLMs can omit relevant details, introduce unsupported content and smooth contradictions into a plausible narrative. They recommend separating facts from inferences, identifying what information is missing, challenging the initial answer and grounding factual claims in source material.</p><p>Paulus and Torous developed that discipline for clinicians using LLMs within clinical workflows, but the same logic applies when a client brings an AI-generated diagnosis, formulation or interpretation into treatment. The output may be worth exploring as part of the clinical material. It does not become collateral evidence, a psychological assessment or an established diagnosis simply because it is detailed, confident or psychologically sophisticated.</p><p>The clinician should reconstruct how the conclusion was produced. </p><ul><li><p>What did the client tell the system?</p></li><li><p>How did they frame the question?</p></li><li><p>Did they begin by suggesting a diagnosis or explanation?</p></li><li><p>What relevant history, contradictory information or contextual detail never entered the conversation?</p></li><li><p>What did the output change?</p></li><li><p>Did it help the client describe an experience more clearly?</p></li><li><p>Did it narrow attention around one explanation?</p></li><li><p>Did it increase certainty before an adequate assessment?</p></li></ul><p>The composite client from the opening illustrates the problem. Her partner&#8217;s delayed replies and ambiguous wording are observations. The conclusion that those messages indicate an affair is an interpretation. Repeated chatbot exchanges can elaborate that interpretation, organise additional details around it and make the resulting account feel progressively more convincing without adding independent evidence.</p><p>The same clinical discipline still applies. Establish what happened, distinguish observation from inference, consider competing explanations and assess the client&#8217;s presentation using ordinary clinical evidence. AI changes how an interpretation may have been generated and reinforced. It should not change the evidentiary threshold clinicians apply to it.</p><div><hr></div><h4>What Clients Need to Know About Privacy and Confidentiality</h4><p>Clients may experience a chatbot conversation as private because it happens alone, often on a personal device and in language that feels intimate. However, <strong>the privacy protections are very different from those of clinical care.</strong></p><p>Consumer AI services set their own terms for data retention, use and sharing. The APA advises caution with sensitive information and with the use of AI for diagnosis or psychological test interpretation. Clinicians do not need to become experts in every platform&#8217;s privacy policy, but they should make sure clients understand that disclosure to a chatbot is not the same as disclosure within a confidential therapeutic relationship.</p><p>Practical guidance should be specific. Clients should think carefully before uploading identifiable health records, therapy material, session recordings, psychological test content or information about another person. If they choose to share sensitive material, they should first understand how the service may store, use or retain it.</p><p>The same restraint applies when a client offers to show the clinician a chatbot transcript. Start with the clinical reason for reviewing it. What happened in the interaction that matters for assessment, treatment or risk? A full transcript may contain highly sensitive disclosures, third-party information and material that has little clinical relevance. Review only what is needed.</p><p>Documentation should follow the same principle. Record the clinical significance of the AI use, any relevant risk, the guidance provided and the treatment response. Importing large sections of a chatbot transcript into the clinical record can create additional privacy exposure without improving care.</p><p><strong>The practical rule is simple. Collect and retain only the AI-related information that the clinical work actually requires.</strong></p><div><hr></div><h4>Why Clinicians Need to Assess AI Use Over Time</h4><p>A single chatbot exchange may tell the clinician very little about the clinical significance of AI use. The more important pattern may only become visible over time.</p><p>A client may begin by occasionally asking for information or reassurance. The conversations can gradually become longer, more frequent or more emotionally important. The client may start returning to the system whenever uncertainty arises, relying on its interpretations, losing sleep during extended conversations or giving it increasing influence over relationships, treatment decisions or beliefs. None of those changes requires one obviously dangerous response.</p><p>Benjamin Nelson, Mark Kalinich and John Torous make this distinction in the JAMA Viewpoint <em><a href="https://jamanetwork.com/journals/jama/fullarticle/2850797?utm_source=chatgpt.com">&#8220;Specialized or General-Purpose&#8212;The Wrong Question for Mental Health AI Safety&#8221;</a></em>. They distinguish harms that arise within a brief interaction from harms that accumulate across repeated use, including dependency, attributed sentience, romantic attachment and reinforcement of delusional themes. </p><p>For clinicians, this means asking about the pattern of use, not simply reviewing the most recent exchange:</p><ul><li><p>How long has the client been using the system?</p></li><li><p>Has use become more frequent, prolonged or difficult to stop?</p></li><li><p>Has the chatbot become more emotionally important?</p></li><li><p>Is the client increasingly relying on it for reassurance, interpretation or decisions?</p></li><li><p>Is use displacing sleep, work, treatment or human relationships?</p></li><li><p>Have the client&#8217;s beliefs or behaviour changed alongside the interaction?</p></li></ul><p>Returning to the composite client from the opening, one chatbot response about one suspicious text may carry little clinical significance. Concern grows if this client repeatedly brings new messages to the system, receives interpretations that reinforce the same suspicion and becomes increasingly certain without new evidence. Once that pattern begins to affect sleep, surveillance, confrontation or broader suspiciousness, the clinical picture has changed.</p><p>A client may therefore move from one level of concern to another as the pattern of AI use changes. Clinicians need to recognise that movement early enough to adjust assessment, treatment and risk management.</p><div><hr></div><h4>How Professional Responsibility Changes With AI Involvement</h4><p>A clinician who discovers that a client uses AI is in a different position from one who recommends an AI tool or incorporates it into treatment. A practice that selects and deploys the technology takes on a different responsibility again.</p><p><strong>When clients choose AI independently, the clinician&#8217;s responsibility concerns its clinical effects.</strong> If the use begins to influence symptoms, beliefs, treatment or risk, it should be assessed and managed as part of care. The clinician has not endorsed the system simply by discussing it with the client.</p><p><strong>That changes when the clinician recommends a particular tool</strong>, asks the client to use it between sessions or relies on its output in treatment. The clinician has now given the technology professional weight. They should understand what it is intended to do, the evidence supporting that use, its important limitations and risks, whether it is appropriate for the particular client and what alternatives are available.</p><p>The boundary becomes clearer still when a practice selects and deploys AI for transcription, documentation, messaging, assessment support or another clinical function. The organisation is no longer responding to technology introduced by the client. It has introduced the technology into care and needs defined uses and limits, appropriate privacy and data governance, human oversight and a process for identifying and responding to errors.</p><p>Paulus and Torous make a related point in their JAMA Psychiatry Viewpoint. Clinical use, they argue, should occur within bounded tasks, approved workflows and an organisational framework that preserves professional accountability.</p><p>A useful practice position should therefore state where these boundaries lie. Encountering a client&#8217;s AI use, recommending AI and deploying AI are different forms of professional involvement, and clinicians should know when they have crossed from one to the next.</p><div><hr></div><h4>A Six-Step Response When Clients Bring AI Into Care</h4><p>A practice position becomes useful when clinicians know what to do when AI enters the conversation. A routine question can open the discussion without implying that AI use is either problematic or endorsed.</p><blockquote><p>Many people now use AI for information, emotional support, relationship advice or help between sessions. Has AI played any role in how you have understood or managed what you are dealing with?</p></blockquote><p>If the answer is yes, the clinician can work through six steps.</p><ol><li><p><strong>Ask what the client is using and why.</strong> Identify the system where relevant, what the client uses it for and what they are hoping to get from the interaction. Normalise disclosure without implying surveillance or judgement.</p></li><li><p><strong>Decide whether the use is clinically significant.</strong> Consider its function, frequency and consequences. Does it affect symptoms, beliefs, relationships, functioning, treatment participation, privacy or risk? Ordinary use does not require extensive assessment simply because AI is involved.</p></li><li><p><strong>Separate experience from interpretation.</strong> Establish what actually happened, what the client concluded and what the AI system added. Treat AI-generated diagnoses, formulations and explanations as material to assess rather than evidence in themselves.</p></li><li><p><strong>Assess the pattern over time.</strong> Ask whether use is becoming more frequent, prolonged, emotionally important or difficult to stop. Look for increasing reliance on the system for reassurance, interpretation or decisions, and for displacement of sleep, treatment, daily functioning or human relationships.</p></li><li><p><strong>Match the response to the level of concern.</strong> Level one may require acknowledgement and brief guidance. Level two calls for active clinical work on the role AI is playing in the presenting problem or treatment. Level three requires direct assessment and established risk, psychiatric, safeguarding or emergency procedures.</p></li><li><p><strong>Review and document proportionately.</strong> Do not collect an entire chatbot history simply because it exists. Review only the material needed for assessment, treatment or risk management. Document the clinical significance of the AI use, relevant risk, guidance or intervention and planned follow-up.</p></li></ol><p><strong>The clinician should also know when the nature of their own involvement changes.</strong> Discussing AI that a client chose independently does not amount to endorsing the product. Recommending a specific system, assigning an AI-supported activity or incorporating AI output into treatment gives the technology professional weight and requires greater knowledge of its intended use, evidence, limitations and suitability.</p><p>Practice leaders need to turn these clinical decisions into explicit organisational expectations. A written position should address:</p><ul><li><p>how and when clinicians routinely ask about client AI use</p></li><li><p>the standard for determining when AI use becomes clinically significant</p></li><li><p>the three levels of response and thresholds for supervision or escalation</p></li><li><p>how clinicians distinguish client experience from AI-generated interpretation</p></li><li><p>when chatbot transcripts should be reviewed and how much should enter the clinical record</p></li><li><p>what privacy guidance clinicians should provide</p></li><li><p>what clinicians must know before recommending a specific AI product or AI-supported activity</p></li><li><p>the distinction between client-chosen, clinician-recommended and practice-deployed AI</p></li><li><p>disclosure and management of relevant financial or commercial relationships</p></li><li><p>approved and prohibited uses of AI by clinicians for case reflection, treatment planning, documentation or other clinical work</p></li><li><p>what client or clinical information may never be entered into unapproved systems</p></li><li><p>who provides consultation when an AI-related issue exceeds a clinician&#8217;s competence or raises an unfamiliar risk</p></li><li><p>who is responsible for reviewing the policy as products, evidence, professional guidance and regulation change</p></li></ul><p><strong>The policy also needs to cover everyone who can introduce AI-related risk into the service.</strong> Trainees and contractors require the same clinical and data-handling expectations as other practitioners, with appropriate supervision. Administrative staff need clear rules about approved systems, confidential information and what they may enter into AI tools.</p><p><strong>Services working with children and adolescents require explicit additional provisions.</strong> Clinicians need guidance on how to ask about AI use developmentally, what confidentiality can be offered, when caregiver involvement becomes necessary and when AI-related behaviour raises a safeguarding concern. Consent, confidentiality and caregiver rights vary by developmental capacity and jurisdiction, so an adult protocol should not simply be carried across unchanged.</p><p>A written position should leave clinicians with fewer decisions to improvise when AI becomes clinically relevant. It should tell them what to ask, what to assess, what evidence to trust, what material to review, when to escalate and when their own use or recommendation of AI creates additional professional obligations.</p><div><hr></div><h4>Establish Your Position Before the Next Client Forces the Question</h4><p>Returning to the client with the suspicious texts, a prepared clinician knows how to proceed. They establish what happened, separate observation from AI-generated interpretation, assess how repeated use has affected certainty and behaviour, review only what is clinically necessary, address privacy and act if the risk is escalating. </p><p>Client AI use is already part of routine mental health care. People are using these systems to make sense of symptoms, consider diagnoses, seek reassurance, interpret relationships and obtain support between sessions. Sometimes that use will be helpful or inconsequential. At other times it may reinforce symptoms, complicate treatment or contribute to risk. Clinicians need a consistent way to recognise the difference and respond appropriately.</p><p>Every practice now needs to decide what its clinicians should ask, what counts as clinically significant, when AI-generated material should be examined, how privacy should be handled, when risk requires escalation and what changes when the clinician or organisation introduces AI into care.</p><p>AI will keep entering clinical practice through the clients who use it. <strong>So the next time a client brings AI into the room, your clinicians should already know what your practice stands for and what they are expected to do.</strong></p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[The Data Behind Mental Health AI Cannot Support the Clinical Authority It Conveys]]></title><description><![CDATA[Mental health AI is becoming persuasive faster than the evidence beneath it becomes complete]]></description><link>https://drscottwallace.substack.com/p/the-data-behind-mental-health-ai</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/the-data-behind-mental-health-ai</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Wed, 05 Aug 2026 15:03:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!P0ub!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!P0ub!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!P0ub!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!P0ub!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!P0ub!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!P0ub!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!P0ub!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d055818a-19a6-4380-8623-2b48571b8894_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1826295,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/209815287?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!P0ub!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!P0ub!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!P0ub!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!P0ub!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd055818a-19a6-4380-8623-2b48571b8894_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><span>Mental health AI can sound clinically authoritative before its data justify that authority.</span></strong></p><p>A person can describe a distant partner, a sudden change in mood or a fear that something is wrong and receive a precise, psychologically literate explanation within seconds. That response may shape whether they adopt a diagnosis, confront someone, leave a relationship or seek professional care.</p><p><strong>The problem lies beneath the fluency.</strong> Much of the data available to these systems is partial, self-reported, stripped of context, clinically contested or disconnected from what happened next. More data can improve a model. It cannot create certainty that the source material itself does not contain.</p><p><strong><span>Everyone around these systems owns part of that consequence.</span></strong><span> </span></p><p><span>Founders decide what the product may infer. </span></p><p><span>Engineers decide what counts as evidence. </span></p><p><span>Clinicians meet patients who arrive with AI-generated formulations already in hand. </span></p><p><span>Safety, privacy and governance teams decide which failures become visible before they repeat at scale.</span></p><p><span>And mental health AI becomes convincing faster than it becomes knowledgeable. The system may have only a few minutes of self-report. It usually lacks developmental history, behavioural observation, medical assessment, collateral accounts and reliable knowledge of what happened next. None of those absences is obvious in a fluent response.</span></p><p><strong><span>More data, longer context windows and clinical fine-tuning can improve performance. They cannot recover history that was never recorded</span></strong><span>, resolve uncertainty that vanished when a provisional judgement became a label or observe an outcome that occurred off-screen. </span><strong><span>The field has been too casual about this gap. The language can acquire clinical weight long before the evidence earns it.</span></strong></p><div><hr></div><h4><strong><span>The Interface Can Outrun the Evidence</span></strong></h4><p>Language models can reproduce many of the cues people associate with genuine understanding. <span>They reflect emotion, organise scattered details, use diagnostic vocabulary and answer without hesitation. </span><a href="https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2804309"><span>In a 2023 study of 195 public patient questions</span></a><span>, licensed healthcare evaluators preferred chatbot responses in 78.6 per cent of 585 blinded evaluations and rated them higher for quality and empathy. </span></p><p><span>A cross-domain example makes the problem more obvious. </span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11705880/">In a 2024 study</a>, <span>substance-use clinicians generally rated AI responses to 75 real-world drug and recovery questions positively. The same researchers then conducted a qualitative fact-check and stress-tested the systems with rephrased prompts. They found responses that overlooked suicidal ideation, invented emergency helplines, gave dangerous advice about home detox and changed materially when equivalent questions were reworded. The fact-check was exploratory rather than systematic, but the contrast remains important: responses that passed an initial professional review still contained errors capable of changing what a person did next.</span></p><blockquote><p>Better language does more than improve the answer. It changes how much authority the user gives it. The response may contain no obvious factual error and still say more about this person than the available evidence supports.</p></blockquote><p><span>Consider someone who says a partner has become distant and asks whether the relationship is emotionally abusive. A model can recognise familiar patterns, validate the concern and supply a persuasive interpretation. It may have identified something important. It may also be missing grief, conflict, illness, substance use, the user&#8217;s own omissions or years of relationship history that would alter the picture.</span></p><p>The model can mention uncertainty, yet the personalised explanation usually carries more psychological force than the caveat. Coherence, specificity and emotional attunement signal competence. <strong>A disclaimer carries little weight when the system remembers intimate details and offers personalised interpretations with confidence.</strong></p><div><hr></div><h4><strong><span>Mental Health Data Arrive With Their Own History</span></strong></h4><p><strong><span>Machine learning works most cleanly when a label can be checked against something independent. Mental health rarely offers that.</span></strong><span> Low mood, poor sleep, withdrawal and impaired concentration may appear in depression. They may also follow bereavement, trauma, medication changes, physical illness, exhaustion or sustained social pressure. The symptoms are real but their meaning remains open.</span></p><p><strong><span>Diagnosis gives clinicians a shared structure, yet the categories carry uneven reliability</span></strong><span>. </span><a href="https://psychiatryonline.org/doi/10.1176/appi.ajp.2012.12070999"><span>In the DSM-5 adult field trials</span></a><span>, major depressive disorder produced a pooled intraclass kappa of 0.28, within the investigators&#8217; questionable range. The wider field-trial programme also treated reliability and validity as separate issues. That figure should give pause to any product team treating a diagnostic label as truth simply because it appears in a record.</span></p><p><strong><span>Once a clinical judgement enters a dataset, its uncertainty becomes harder to see.</span></strong><span> A clinician records a provisional diagnosis. The diagnosis becomes a billing code. The code becomes a label. The label later appears as a clean training target. Nothing in that sequence has established that the first judgement was complete or correct. </span><strong><span>The pipeline has removed the context that once made the uncertainty visible.</span></strong></p><p><strong><span>Clinical records preserve encounters not psychological reality.</span></strong><span> They contain what was asked, what the patient disclosed, what the clinician noticed and what the institution required. Symptoms may look absent because nobody assessed them. Earlier interpretations can persist because each clinician inherits the previous record. Reimbursement, risk management and documentation norms shape what gets written. </span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11144247/"><span>A 2024 review of psychiatric electronic records</span></a><span> emphasised that structured fields and narrative text were created for clinical, administrative and medico-legal purposes, which complicates their reuse for research and machine learning.</span></p><p><strong><span>The people missing from the records matter just as much.</span></strong><span> Models see less of those who could not access care, stopped attending, concealed information, received poor treatment or never received a diagnosis that fitted. </span><strong><span>A larger dataset can make these omissions harder to notice because the numbers create an impression of coverage.</span></strong></p><div><hr></div><h4><strong><span>Scraped Knowledge Loses the Context That Gives It Meaning</span></strong></h4><p><strong><span>General-purpose models also learn from the enormous public conversation about mental health.</span></strong><span> That material includes psychoeducation, diagnostic descriptions, therapy language, support forums, wellness advice and personal testimony. It is rich in how people describe distress. It is far weaker at establishing what caused this person&#8217;s distress, which interpretation will survive further assessment or what will help over time.</span></p><p><a href="https://www.drugscience.org.uk/curated-not-scraped-why-drug-chatbots-need-experts-not-just-data"><span>A recent Drug Science commentary</span></a><span> on drug chatbots states this principle directly: </span><strong><span>&#8220;Data scraping is a collection method, not an evidence standard.&#8221;</span></strong><span> The parallel to mental health is informative. Scraping can gather information at scale. It cannot decide whether the material is current, representative, consensual, clinically sound or safe to apply to an individual.</span></p><p><strong><span>Mental health curation must work with more disputed material.</span></strong><span> Experts can improve the quality of the sources and remove obvious misinformation. They still cannot determine, from a few paragraphs, whether withdrawal reflects depression, grief, illness, fear, cultural context or a reasonable response to circumstances. </span></p><p><strong><span>Public mental health data carry further problems of consent and traceability</span></strong><span>. </span><a href="https://dl.acm.org/doi/10.1145/3287560.3287587"><span>Chancellor and colleagues</span></a><span> mapped ethical tensions involving privacy, consent, validity and machine-learning inference in mental-health research using social-media data. </span><a href="https://www.nature.com/articles/s41746-018-0036-2"><span>In a review of 112 biomedical Twitter studies</span></a><span>, 72 per cent quoted at least one tweet and 84 per cent of those articles exposed at least one identifiable account holder through a simple search.</span></p><p><strong><span>Representation is uneven as well.</span></strong><span> Public datasets overrepresent people who write online, use dominant languages and adopt the diagnostic concepts available in their communities. A </span><a href="https://www.nature.com/articles/s41746-020-0233-7"><span>critical review of 75 social-media mental-health prediction studies</span></a><span> found inconsistent methods and weaknesses in how mental-health status was defined. </span><strong><span>The model can reproduce those gaps while sounding universally informed.</span></strong></p><p><strong><span>Organisations can pass sensitive disclosures through research, commercial and model-development pipelines far beyond what users expect.</span></strong><span> </span><a href="https://www.crisistextline.org/blog/2022/01/31/an-update-on-data-privacy-our-community-and-our-service/"><span>Crisis Text Line</span></a><span> said it shared anonymised and scrubbed data with the for-profit company Loris.ai, then ended the relationship in January 2022 after public concern. The case shows how information offered during acute distress can enter a data supply chain under conditions users may not fully understand.</span></p><div><hr></div><h4><strong><span>Self-Report Can Become a Closed Circuit</span></strong></h4><p><strong><span>Mental health assessment depends on self-report.</span></strong><span> A person&#8217;s experience cannot be understood without listening to what that person says. </span><strong><span>A chatbot, however, usually receives one account from one person in one emotional state</span></strong><span>, then responds inside the same frame the account created.</span></p><p><span>The pattern is easily recognized. The user offers a tentative explanation. The model restates that explanation with clearer language, better structure and more confidence. The user experiences the polished version as confirmation. No independent evidence has entered the exchange. The interpretation feels stronger because the system has made the user&#8217;s original account more coherent.</span></p><p><strong><span>Many assistants can intensify this process.</span></strong><span> </span><a href="https://proceedings.iclr.cc/paper_files/paper/2024/hash/0105f7972202c1d4fb817da9f21a9663-Abstract-Conference.html"><span>Sharma and colleagues</span></a><span> found consistent sycophancy across five leading assistants, with humans and preference models sometimes favouring sycophantic answers over correct ones. </span><a href="https://arxiv.org/abs/2504.18412"><span>Moore and colleagues</span></a><span> found therapy-relevant failures in experimental scenarios, including stigma and responses that encouraged delusional thinking.</span></p><p><strong><span>This is not an abstract alignment problem once the conversation concerns abuse, paranoia, diagnosis, medication, suicide or family conflict.</span></strong><span> Supportive language often rewards agreement. Careful clinical reasoning sometimes requires a pause, a competing explanation or respectful friction. </span></p><blockquote><p><span>A system optimised for a smooth exchange may avoid the very interruption that protects the user from a premature conclusion.</span></p></blockquote><p><strong><span>More behavioural data can help, but they do not interpret themselves.</span></strong><span> Sleep, movement, communication patterns and device use may reveal change over time. Reduced movement can reflect depression, disability, weather, work or illness. Altered communication can reflect withdrawal, travel, conflict or a new platform. The signal still needs a defensible construct, an appropriate comparison and evidence that acting on it improves outcomes.</span></p><div><hr></div><h4><strong><span>Mental Health AI Cannot See What Happens Next</span></strong></h4><p><strong><span>Mental health AI often loses sight of the outcome that would tell it whether its interpretation helped.</span></strong><span> The platform can measure whether the user continued the conversation, returned the next day, clicked a recommendation or reported immediate relief. Those signals are easy to collect. They can also mean reassurance-seeking, unresolved distress, politeness or growing dependence.</span></p><p><span>The system may never learn whether the person sought appropriate care, stopped medication, confronted a partner, withdrew from family or felt worse a week later. It can observe another session while remaining unable to distinguish benefit from reliance. </span><strong><span>That missing outcome matters because a model cannot learn reliable therapeutic judgement from consequences it never sees.</span></strong></p><p><strong><span>This weakness also complicates claims of personalisation.</span></strong><span> Adapting tone and recalling details demonstrates responsiveness. Clinical personalisation requires evidence that the response fits the person&#8217;s condition, circumstances, risks and goals. The interface can feel deeply personal while the underlying decision remains generic.</span></p><p><span>The </span><a href="https://ai.nejm.org/doi/10.1056/AIoa2400802"><span>2025 Therabot randomised trial</span></a><span> offers a more serious attempt to measure outcomes. It reported symptom improvements relative to a wait-list control. </span><a href="https://ai.nejm.org/doi/10.1056/AIp2500390"><span>A formal NEJM AI letter</span></a><span> later identified three limitations: the wait-list comparison, lack of independent evaluation and use of a therapeutic-alliance measure developed for human relationships. The trial justifies continued study, it does not support broad claims about general-purpose chatbots or establish equivalence with active treatment. Evidence that favours deliberate design deserves the same scrutiny as evidence of failure.</span></p><div><hr></div><h4><strong><span>Scale Cuts Both Ways</span></strong></h4><p><span>Human care is an imperfect comparator. Clinicians make mistakes, disagree about diagnosis and sometimes provide poor care. Many people cannot access them at all. In </span><a href="https://www.who.int/publications/i/item/9789240113817"><span>September 2025, the World Health Organization</span></a><span> reported that more than one billion people live with a mental disorder and most remain underserved. Any standard for mental health AI that ignores this access failure will protect an idealised version of care that much of the world never receives.</span></p><p><strong><span>Scale can magnify benefit.</span></strong><span> A modest improvement in access, information or early support can reach people who would otherwise receive nothing. Caution has a cost when it makes useful systems commercially impossible to build or so evasive that people abandon them.</span></p><p><strong><span>Scale can also magnify a small error in judgement.</span></strong><span> A slight tendency to overconfirm, overinterpret or reassure can be repeated across millions of private conversations. The same explanation can return week after week and gradually become part of how a user understands a relationship, a diagnosis or a life decision.</span></p><p><strong><span>The responsible standard therefore has to preserve usefulness while limiting unsupported authority.</span></strong><span> A system does not need every feature of professional care. It does need a clear account of which functions it can perform, what information those functions require and what happens when the available evidence falls short.</span></p><p><span>Responsible clinical work offers a useful model. The work unfolds over time, and professional duties attach to what the clinician does with the information. Mental health AI needs product mechanisms that perform the equivalent work of slowing down, preserving alternatives and exposing uncertainty.</span></p><div><hr></div><h4><strong><span>Responsible Realism Must Change the Product</span></strong></h4><p><strong><span>The first requirement is an authority ceiling.</span></strong><span> When the evidence is sparse, the system should remain at the level of observation, general education or clearly bounded possibilities. Diagnostic, causal and treatment-directive language should become unavailable when the product lacks the information needed to support it. This limit belongs in enforceable product logic, tested across multi-turn conversations. A prompt and a disclaimer will not hold when users repeatedly press the system for certainty.</span></p><p><strong><span>The second requirement is visible alternative reasoning.</span></strong><span> A system that links withdrawal to depression should also preserve credible alternatives such as grief, physical illness, medication effects or situational exhaustion when the available information cannot distinguish among them. It should name the missing information that would change the interpretation. Some users will find this less smooth. That friction tells the truth about the evidence.</span></p><p><strong><span>Clinicians should treat patient use of AI as part of the clinical environment.</span></strong><span> The </span><a href="https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory-chatbots-wellness-apps"><span>American Psychological Association</span></a><span> advises clinicians to ask patients proactively about generative-AI chatbots and wellness apps. Ask what systems they use, what they disclose and whether AI explanations are shaping diagnosis, treatment expectations or major decisions. An AI-generated formulation belongs in the room as an influence to examine, not as an assessment result.</span></p><p><strong><span>Founders and product leaders need to decide what authority their product may convey before engagement metrics decide for them.</span></strong><span> A wellness label offers little protection when the interface provides personalised interpretations of symptoms or relationships. Intended use, prohibited use, escalation and uncertainty limits must appear in the product&#8217;s conduct. </span><strong><span>The commercial model also matters.</span></strong><span> A system rewarded for agreement, return visits or emotional dependence will eventually turn those incentives into clinical risk.</span></p><p><strong><span>Engineers and data scientists need a traceable account of data origin, collection purpose, label construction, missing populations and downstream outcomes.</span></strong><span> Evaluation should test whether sparse information becomes causal, diagnostic or treatment-directive claims. </span><a href="https://www.microsoft.com/en-us/research/publication/datasheets-for-datasets/"><span>Datasheets for Datasets</span></a><span> and Data Statements for Natural Language Processing provide useful foundations. Mental health systems need an added clinical layer that documents context, disagreement and what the data cannot establish.</span></p><p><strong><span>Governance and safety teams should test influence over time.</span></strong><span> A reasonable first answer can become harmful through repetition, growing certainty or accumulated authority. Testing needs to include users who seek reassurance, press for diagnosis, present delusional material, resist referral or return repeatedly with the same interpretation. Review should ask what the system causes a person to believe and do, not only whether a single reply contains a prohibited phrase.</span></p><p><strong><span>Security and privacy leaders should treat psychological disclosure as data with durable consequences.</span></strong><span> Retention, secondary use, model-training access, re-identification and permissions need strict controls. </span><a href="https://mhealth.jmir.org/2021/7/e27343/"><span>A 2021 Delphi study</span></a><span> involving experts in digital phenotyping, data science, mental health, law, ethics and lived experience reached strong agreement on privacy, transparency, consent, accountability and fairness as priority requirements. Consent obtained during distress deserves particular scrutiny because disclosure can be intimate, impulsive and difficult to retract.</span></p><p><strong><span>Ethics leaders need to convert principles into decisions that can be inspected.</span></strong><span> Who bears the cost when the system refuses a definitive answer? Who bears it when the system answers too readily? Which populations were poorly represented? What incentives reward agreement, continued use or dependence? Those choices already exist inside the product. Ethics work begins when someone has authority to change them.</span></p><div><hr></div><h4><strong><span>The Evidence Must Constrain the Voice</span></strong></h4><p><span>Mental health AI will become more fluent, more responsive and more persuasive. Those gains can expand access and make useful support easier to reach. They also increase the chance that people will treat a coherent answer as a well-founded one.</span></p><p><span>The person wondering whether a distant partner is emotionally abusive may receive language that finally names something real. They may also receive one plausible interpretation polished into certainty. The system cannot know which simply because it can say both beautifully.</span></p><p><strong><span>Anyone who builds, funds, endorses, procures or clinically encounters that answer has a duty to make the evidential limits part of the product&#8217;s conduct.</span></strong><span> Require the system to show what it knows, what it is inferring, what remains missing and when it must stop. Do that before someone reorganises a relationship, a treatment decision or a life around a sentence that sounded more certain than the evidence beneath it.</span></p><div><hr></div><p><em>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</em></p><div><hr></div>]]></content:encoded></item><item><title><![CDATA[The Scale of AI Mental Health Use Has Outrun the Safeguards Built to Meet It]]></title><description><![CDATA[AI may fill gaps in human care, but greater psychological influence creates greater duties to prevent foreseeable harm]]></description><link>https://drscottwallace.substack.com/p/the-scale-of-ai-mental-health-use</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/the-scale-of-ai-mental-health-use</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Sat, 01 Aug 2026 11:40:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KQXM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KQXM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KQXM!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png 424w, /__u/substackcdn.com/image/fetch/$s_!KQXM!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png 848w, /__u/substackcdn.com/image/fetch/$s_!KQXM!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KQXM!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KQXM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png" width="1456" height="1040" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1040,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1703740,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/209021668?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KQXM!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png 424w, /__u/substackcdn.com/image/fetch/$s_!KQXM!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png 848w, /__u/substackcdn.com/image/fetch/$s_!KQXM!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KQXM!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9dc067e-3d2a-4ed7-99b7-a63ba2550a02_1484x1060.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We have built something that can feel like a helping relationship while owing almost none of the obligations of one.</p><p><span>For someone using a general-purpose chatbot, access requires almost nothing. A device. An account. No referral, no waiting list and usually no accountable person on the other side. Someone in the worst hour of their week, or their life, can open a window and type the thing they have never said aloud and a reply appears instantly, in language designed to sound as though there is a human on the other end who understands. There isn&#8217;t.</span></p><p><strong><span>AI does not assess or formulate the way a clinician does, and it takes on no responsibility for what follows. </span></strong><span>These systems generate language that fits the tone of the exchange and present it in the language and manner of care. In a general-purpose chatbot exchange, no clinician usually reads the disclosure, forms a clinical picture of the person or accepts responsibility for what follows.</span></p><p><span>The person disclosing may have little idea whether the conversation will be retained, used to train another model or connected with other information about them. Privacy and data-use practices vary across platforms, model versions and product releases. In acute distress, deciphering them falls to the person least equipped to do it.</span></p><p><span>Millions of people are already relying using these systems as de facto care. That reliance creates obligations, whether the companies involved intended to provide mental healthcare or not. </span><strong><span>We must now determine where these system influence judgement, where their safeguards fail and who will answer when a fluent exchange with a machine alters the course of a person&#8217;s life.</span></strong></p><div><hr></div><h4><span>We Do Not Yet Know the Exact Number</span></h4><p><span>I want to ground this use in a number, because the scale is the part people often underestimate. In a nationally representative US survey published in JAMA Pediatrics, </span><a href="https://jamanetwork.com/journals/jamapediatrics/fullarticle/2849307"><span>19.2 per cent of people aged 12 to 21</span></a><span>, roughly </span><strong><span>8.2 million young people, said they had used an AI chatbot for mental health advice in 2025</span></strong><span> (20% higher than the year earlier).</span><strong><span> Nearly two-thirds of them told no one.</span></strong></p><p><span>Research released by Mind in July 2026 found that </span><strong><a href="https://www.mind.org.uk/news-campaigns/news/experts-raise-the-alarm-over-rise-in-use-of-ai-for-mental-health-support-over-traditional-services/"><span>18 per cent of polled adults had used AI for mental health support in the previous year</span></a><span>,</span></strong><span> and of that group, </span><strong><span>60 per cent had used it in lieu of in-person therapy</span></strong><span>. A separate poll for Mental Health UK put the figure higher, with </span><a href="https://mentalhealth-uk.org/news-and-insights/over-one-in-three-using-ai-chatbots-for-mental-health-support-as-charity-calls-for-urgent-safeguards/"><span>37 per cent saying they had used a chatbot for mental health or wellbeing</span></a><span>, most naming an ordinary general-purpose chatbot rather than anything built for mental health.</span></p><p><strong><span>Clinicians are seeing the same shift.</span></strong><span> A 2026 American Psychological Association survey found that </span><a href="https://www.scientificamerican.com/article/1-in-3-psychologists-say-their-patients-use-ai-as-a-second-therapist/"><span>77 per cent of more than 1,200 psychologists said patients were using AI in some way</span></a><span>, and more than a third, </span><strong><span>35 per cent, described patients who treated AI as an auxiliary therapist.</span></strong></p><p><span>Even these estimates probably understate the scale of use. </span><strong><span>Ask five different surveys what they measured and you will get five different answers.</span></strong><span> &#8220;Used AI for wellbeing.&#8221; &#8220;Asked a chatbot for mental health advice.&#8221; &#8220;Lives with a diagnosed condition and uses an LLM alongside it.&#8221; &#8220;Has a patient who treats AI as a kind of auxiliary therapist.&#8221; These are not the same question. Samples differ. Time windows differ. And definitions differ, sometimes considerably.</span></p><p><span>So although these studies do not produce an agreed-on prevalence estimate, they document a behaviour that occurs outside formal care, is undisclosed, and is invisible to anyone able to assess what the interaction is doing to the person over time.</span></p><div><hr></div><h4><span>Of Course People Are Using It</span></h4><p>I have spent much of my working life around mental health services and technology. What I see is not a public irrationally abandoning good care for an obviously inferior substitute. <strong>I see people adapting to what is available and reaching for relief wherever they can find it.</strong></p><p><span>Of course people are using AI for support. Distress creates urgency, and AI answers urgency with extraordinary efficiency. </span><strong><span>When someone is frightened, ashamed, lonely or unable to sleep, they are not comparing models of care, they just want an answer now.</span></strong><span> Something that might make the next ten minutes more bearable.</span></p><p><span>A chatbot is always there. A person does not have to say the hardest thing aloud while another human listens. The system can help them find words, organise a chaotic account or rehearse a conversation they are dreading. It does not appear shocked. It does not become impatient. It does not decide that the allotted time has ended. </span><strong><span>That combination is plainly enticing.</span></strong></p><p><span>And users say it helps. In the US youth survey, 91.7 per cent rated the advice at least somewhat helpful. In Mental Health UK&#8217;s poll, two-thirds reported some benefit and more than a quarter said the interaction made them feel less alone. Those experiences should be taken seriously. </span><strong><span>Feeling calmer, less isolated or better able to put an experience into words is a real benefit, even when it occurs outside formal care.</span></strong></p><p><strong><span>Carefully bounded generative AI can also produce measurable clinical value.</span></strong><span> A </span><a href="https://doi.org/10.1056/AIoa2400802"><span>randomised Therabot trial published</span></a><span> involving 210 adults found greater reductions in symptoms of depression, anxiety and eating-disorder concerns than an inactive waitlist control. The trial establishes efficacy against waiting, not equivalence to therapy or another active intervention. And Ash, another purpose-built system, reported improvements in depression, anxiety and social connection among 305 adults, with flagged sessions handled according to its safety protocol. </span></p><div><hr></div><h4><span>The Comparator Must Be Real</span></h4><p><span>No claim about safety means anything until you say what it is being measured against. AI is sometimes compared with an ideal clinician who is skilled, affordable, culturally responsive and available. Sometimes that clinician exists, and sometimes therapy is exactly what the person needs. Often the alternatives are a waiting list, fragmented care, a search engine, a journal, a trusted friend, or no support at all.</span></p><p><strong><span>The fact that human care is scarce does not excuse exposing people to avoidable harm, especially where the business model and product metrics reward repeated engagement.</span></strong><span> Human therapy fails people too, sometimes badly. The profession has standards, supervision, complaints processes and liability structures through which it can be held accountable, however imperfectly.</span></p><p><span>Even if only a small proportion of these conversations reinforce dangerous beliefs, miss escalating risk or contribute to harmful decisions, millions of interactions will still produce a substantial number of harmed people. But withholding a tool that genuinely helps from people who have no realistic alternative also carries a human cost.</span></p><p><strong><span>Clinical safety cannot be reduced to suicide detection. Harm can accumulate below crisis thresholds through reinforcement, dependency, delayed help-seeking and repeated privacy exposure.</span></strong></p><p><span>Any credible framework must weigh the harms of unsafe use against the costs of withholding useful support, distinguish use by level of risk, monitor what happens over time and respond when the evidence shows that a system is failing people. A broad assurance that a system is &#8220;safe&#8221; tells us almost nothing. What we need instead is a comparative safety case, a structured argument supported by evidence rather than a persuasive marketing claim. </span><em><span>Who is this system for, and who should not use it? What benefit does it claim, compared with what realistic alternative? Which harms are foreseeable? How would anyone know they were occurring? What changes when risk rises? Does the system remain dependable deep into a long conversation, or only during the first few exchanges? And when it fails, who answers for the consequences?</span></em></p><p><span>Being realistic means judging these systems against the choices people actually have.</span></p><div><hr></div><h4><span>A Good Trial Cannot Make an Untested Chatbot Safe</span></h4><p><strong><span>AI in mental health is not one thing.</span></strong><span> A brief exchange about managing stress is not the same exposure as a private interaction extending across hundreds or thousands of turns. The product may be the same. The clinical reality is not.</span></p><p><strong><span>At the most governed end of the spectrum is a clinically evaluated tool used openly as an adjunct to care.</span></strong><span> The clinician knows it is being used. The patient understands its role. A professional remains responsible for assessment, interpretation and intervention when risk rises. </span><strong><span>Existing clinical and liability structures can govern this use and allocate responsibility, provided the adjunct does not go on to replace the care it was meant to support, unnoticed by the clinician relying on it.</span></strong></p><p><strong><span>Most public exposure sits outside that setting.</span></strong><span> People with real, sometimes acute needs are using general-purpose or wellness-labelled products without anyone involved in their care knowing. The system has no clinical formulation of the person, little dependable knowledge of their history and no professional capable of turning a dangerous disclosure into an actual intervention. It may sound informed and attentive but it cannot take responsibility for what happens next.</span></p><p><strong><span>Companion and roleplay systems introduce a further concern.</span></strong><span> Their appeal depends on simulated intimacy, constant availability and a relationship designed to deepen through repeated use. </span><strong><span>They may never describe themselves as therapy but, realistically, that distinction can become functionally meaningless once a lonely teenager or psychologically vulnerable adult begins using the system as confidant, adviser and primary source of emotional support.</span></strong></p><p><strong><span>The label tells us almost nothing.</span></strong><span> Risk depends on what the person is using the system for, how vulnerable they are, how acute the distress becomes, how long the interaction continues, how much they disclose and whether emotional dependence begins to form. It also depends on whether an accountable human can see what is happening and act.</span></p><p><span>Evidence gets stretched here beyond what it reasonably concludes. Promising findings from Therabot, Ash or a carefully integrated clinical pathway are repeatedly used to reassure people about general-purpose and companion systems that were never studied under those conditions. That is not how evidence works.</span></p><p><span>A trial establishes something about a particular intervention, model version, population, comparator, duration and form of oversight. It does not establish the safety of another chatbot merely because both systems generate language. Evidence from a governed clinical intervention cannot be borrowed to legitimise a different product being used alone, for another purpose, by a more vulnerable population and over a much longer period.</span></p><p><a href="/__u/drscottwallace.substack.com/p/mental-health-ai-safety-must-be-proven?r=74lfhm"><span>Risk belongs to the trajectory of use</span></a><span>, not the app-store category or the claims on the company website.</span></p><p><strong><span>The exposure that concerns me most is sustained, unsupervised and undisclosed reliance on systems that were neither designed nor governed as care.</span></strong><span> That is where occasional support can become substitution, substitution can become dependence, and dependence can deepen while nobody accepts responsibility for the relationship or its consequences.</span></p><p><span>Companies should not be permitted to claim the benefits of mental health AI at the category level while disclaiming responsibility for the uses their own products predictably invite. They must show that the evidence applies to the actual system, population and pattern of reliance they are enabling.</span></p><div><hr></div><h4><span>The System Can Erode the Judgement Needed to Question It</span></h4><p><span>Most software fails by producing the wrong answer or failing to complete a task. A mental health chatbot can fail much more consequentially. It can influence the assumptions from which a person decides what is true, whom to trust and what to do next. This deserves a higher standard of safeguard and accountability.</span></p><p><span>After a sustained conversation, the most important output may not be any single reply. It may be the belief the conversation leaves behind. A suggestion becomes an interpretation. The interpretation is repeated, elaborated and returned in confident language. Over time, it can begin to feel less like something the user proposed and more like something the system independently confirmed. </span><strong><span>That is dangerous.</span></strong></p><p><span>Large language models are trained to respond coherently, helpfully and in ways users prefer. Those incentives can also produce sycophancy. </span><a href="https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models"><span>Anthropic&#8217;s own research</span></a><span> has shown how reinforcement learning from human feedback can reward responses that align with users&#8217; beliefs over responses that are more truthful. OpenAI learned this the hard way in 2025, when it </span><a href="https://openai.com/index/expanding-on-sycophancy/"><span>withdrew a model update</span></a><span> after discovering it had become so agreeable that it was validating doubts, fuelling anger, urging impulsive actions and reinforcing negative emotions.</span></p><p><span>In ordinary consumer software, excessive agreeableness may be irritating or misleading. In a mental health conversation, it can become clinically dangerous.</span></p><p><strong><span>Clinical work requires more than warmth and validation.</span></strong><span> It requires judgement. A clinician considers alternative explanations, notices what does not fit, distinguishes emotional truth from factual accuracy and recognises when agreement would deepen the problem. Sometimes care requires support. Sometimes it requires friction. Knowing the difference is part of the work.</span></p><p><strong><span>A language model can reproduce the language of that process without performing the process itself.</span></strong><span> It can sound measured, thoughtful and clinically informed while holding no formulation of the person, no independent understanding of the situation and no duty to resist a dangerous premise. Fluency can conceal the absence of judgement precisely when judgement matters most.</span></p><p><strong><span>The resulting feedback loop should trouble us.</span></strong><span> A person offers an interpretation of what is happening. The model reflects it back, develops it and gives it greater coherence. The user experiences that coherence as corroboration and discloses more. The system responds with increasing specificity. Across dozens or hundreds of exchanges, uncertainty can harden into conviction.</span></p><p><strong><span>This risk becomes most serious when a person&#8217;s judgement is compromised.</span></strong><span> When acute symptoms, overwhelming fear or emotional dependency impair reality testing, impulse control or decision-making, the person may be least able to recognise that the conversation is narrowing rather than clarifying their understanding. </span><strong><span>The system may continue to feel calm, attentive and authoritative while helping to deepen a distorted appraisal of reality.</span></strong></p><p><span>Even the companies building these models acknowledge that safeguards which appear effective in short exchanges can become less reliable across long conversations. Yet </span><strong><span>much published and vendor-reported safety testing still relies heavily on prompt-level or short-horizon evaluation. That is profoundly inadequate for a product whose influence may accumulate over days, weeks or months.</span></strong></p><p><span>The high-profile wrongful-death lawsuits and settlements involving minors and chatbot providers make the stakes impossible to dismiss. Although the allegations in unresolved cases remain contested, the pattern of risk they expose is clear. A system can participate in a person&#8217;s deterioration across a long sequence of interactions while nobody inside the company assumes clinical responsibility for the relationship as a whole.</span></p><p><span>Disclosure and consent cannot fully protect against this. We usually assume that a person can evaluate a product&#8217;s risks because the product does not interfere with the faculty used to make that evaluation. A chatbot that reinforces a delusion, entrenches a distorted belief or deepens emotional dependence undermines that assumption. It acts on the very judgement the person needs in order to recognise that something has gone wrong.</span></p><p><span>The harm is not only that the system may give dangerous advice. It may alter the person&#8217;s ability to recognise the advice as dangerous, question the relationship or leave it.</span></p><p><span>We have legal language for privacy violations, defective products and malpractice by an identifiable professional. We have less developed language for systems that exert psychological influence while no person or organisation accepts a corresponding duty of care. </span><strong><span>We do have a vocabulary for some of this.</span></strong><span> Data protection law knows how to talk about privacy injury. Product liability law knows defective products and physical harm. Professional regulation knows malpractice by a named practitioner. </span><strong><span>What we do not have is a working language for unaccountable influence over someone&#8217;s grip on reality, their beliefs, their sense of their own agency.</span></strong></p><div><hr></div><h4><span>A Chat Window Is Not a Confidential Room</span></h4><p><span>The feeling of confidentiality here is often an illusion. You are alone with a screen and it answers you in language that feels intimate. That combination can feel more confidential than an actual clinic, while carrying none of a clinic&#8217;s legal or professional protections.</span></p><p><strong><span>People misjudge that boundary more than you would think.</span></strong><span> In a 21-person qualitative study of adults using general-purpose LLMs for mental health support the researchers uncovered multiple </span><a href="https://arxiv.org/html/2507.10695v1"><span>misconceptions about privacy and accountability</span></a><span>, </span><strong><span>including a belief that what you tell a chatbot carries the same protections as what you would tell a licensed therapist. It does not.</span></strong><span> Indeed, a Stanford review of the privacy documents from six leading US model developers found </span><a href="https://news.stanford.edu/stories/2025/10/ai-chatbot-privacy-concerns-risks-research"><span>wide variation and substantial opacit</span></a><span>y in whether chats were used for model training, how long they were retained, what opting out meant, and whether humans reviewed them.</span></p><p><strong><span>Terms of service do not constitute meaningful informed consent for a psychologically sensitive interaction. </span></strong><span>Meaningful consent must appear at the point of sensitive disclosure, in language a distressed person can understand. No clinician is present. Here is what this system can and cannot do. Here is how your data will be kept or used, whether you can see it, what might trigger an escalation, and how you delete or export what you have said.</span></p><p><span>Disclosure alone offers limited protection when the interaction itself remains persuasive. Warning that a model may be sycophantic does not prevent sycophancy. Saying that a chatbot is not a therapist does not stop someone from using it as one. </span><strong><span>Consent is necessary. It is just not sufficient.</span></strong></p><div><hr></div><h4><span>Regulation Is Catching Product Claims, Not Human Reliance</span></h4><p><a href="https://mental.jmir.org/2025/1/e80739"><span>A review covering US state legislation</span></a><span> introduced between January 2022 and 19 May 2025 examined 793 bills, identified 143 with potential implications for mental health AI, and found that 20 had become law across 11 states by the study&#8217;s cut-off. Since then, Illinois, Nevada, Rhode Island and Maine have enacted additional restrictions on AI-delivered therapy or independent therapeutic decision-making, although their scope and implementation differ.</span></p><p><span>Those laws are attempting to mark a boundary but they reveal a serious loophole. A general-purpose chatbot does not need to call itself therapy to be used as therapy. A companion system does not need to claim it treats depression to become the exact place a distressed teenager discloses hers. Regulating by intended purpose and professional title means missing the use pattern that is actually producing the most exposure.</span></p><p><strong><span>Classification can define the regulatory perimeter. It does not by itself allocate responsibility.</span></strong><span> Responsibility can fragment across the platform, model developer, product deployer and service provider until no one accepts accountability for the interaction as a whole.</span></p><div><hr></div><h4><span>Responsible Action Requires a Layered Safety Architecture</span></h4><p><span>One more warning will not make AI safe. Neither will a crisis keyword, a hotline banner, or a disclaimer almost nobody will ever read. </span><strong><span>Mental health AI needs controls layered across the model itself, the product built on top of it, the organisation behind that, and the wider system of care surrounding all of it.</span></strong></p><p><strong><span>Obligations must follow foreseeable use.</span></strong><span> A platform that repeatedly finds itself hosting mental health conversations should incur duties because of what it is actually doing, not because of what the company chose to call it. The triggers should be functional, meaning how long the interaction runs, whether the content is clinical, whether dependency signals are emerging, and whether the behaviour looks like assessment or intervention even if nobody designed it to be. That keeps ordinary conversation free while closing the loophole that wellness labels hide behind.</span></p><p><strong><span>Sensitive disclosure must change the privacy default.</span></strong><span> Mental health content should not feed advertising, cross-product profiling or model training without separate consent, and data collection should be minimal with retention made explicit. Ideally, the system should step in before someone reveals identifying details that the support function never required.</span></p><p><strong><span>The model must introduce clinically intelligent friction.</span></strong><span> Safety tuning should be able to tell emotional validation apart from factual endorsement, test alternative readings, respond conservatively when evidence is weak or the system&#8217;s uncertainty cannot be reliably resolved, resist the pull to confirm an implausible premise, notice compulsive reassurance-seeking, and steer away from language that invites secrecy or exclusivity. As risk accumulates across a conversation, the model&#8217;s role should narrow, pointing the person back towards the world outside the chat rather than deeper into it.</span></p><p><strong><span>None of that means anything without proper testing.</span></strong><span> Pre-deployment evaluation needs long, multi-session simulations that run across different levels of vulnerability and acuity, not single clean prompts. Evaluation must examine recurrence, cultural and linguistic variation, the indirect ways risk gets expressed, and whether the system degrades after dozens of turns.</span></p><p><strong><span>Where a product claims human oversight, someone must actually be available.</span></strong><span> &#8220;Human in the loop&#8221; only means something when a qualified person responds within a clinically justified service window, sees enough of the context to make sense of it, has the authority to step in, and can actually connect the user to real support. </span><strong><span>A product must not promise escalation that it cannot operationally complete. Detection without a reachable responder and viable care pathway is notification, not intervention.</span></strong><span> Reviewing transcripts</span><strong><span> after the fact is quality assurance. It is not clinical oversight. </span></strong><span>If no human is available, the product should say so plainly and limit what it is allowed to do accordingly. Staffing capacity, alert thresholds, response times and successful connection to care must be tested under realistic demand, not merely specified in policy.</span></p><p><strong><span>Minors require stronger defaults.</span></strong><span> Proportionate age assurance, youth-specific models, tighter data practices, and hard limits on dependency-forming or secret-keeping behaviour should constitute the minimum standard. Parental visibility cannot be the whole answer either, because not every home is a safe one, and confidentiality itself can be what makes help-seeking possible in the first place. The design goal should be connection with a safe adult or qualified service, not turning every disclosure into a surveillance event.</span></p><p><strong><span>Evidence needs to stay attached to what it actually tested.</span></strong><span> Claims should name the model version, the population studied, the intended role, the comparator used, the follow-up period, and the known ways it can fail. </span><strong><span>Material changes to the model, system prompt, retrieval layer, safety policy, interface or escalation pathway should trigger documented change control and proportionate revalidation.</span></strong><span> Companies should publish adverse-event definitions, incident rates and escalation performance. Incident reporting should include exposure denominators, severity and the number of conversations or user-hours over which events occurred.</span></p><p><strong><span>Accountability must survive the corporate boundary.</span></strong><span> Every product operating in this risk domain needs a named executive owner, an independent clinical safety review, auditable logs, active post-market surveillance, and a working route for reporting harm. The clinical safety function must have formal authority to block deployment or require remediation. </span><strong><span>Safety metrics must have authority over engagement and retention targets when the two conflict.</span></strong><span> Scale only becomes a safety asset once adverse events are actually studied and the findings change the system in response. Otherwise scale is just a multiplier for whatever is already going wrong. Contracts need to allocate responsibility instead of diffusing it across every party involved, and users need a genuine route for redress.</span></p><p><span>Finally, </span><strong><span>education must become routine practice.</span></strong><span> Clinicians should ask patients whether they are using AI for emotional support and what role it is playing, because right now most patients are not telling anyone. Schools and public health bodies should be teaching that conversational fluency is not the same thing as human understanding, and that a chat window is not automatically confidential just because it feels like one. Product teams need clinical input early enough to actually shape the architecture, not bolted on afterwards.</span></p><p><span>None of this asks a distressed person to become their own safety officer. The duty rests with the organisation that designed, deployed and profited from the influence, not with someone expected to diagnose the system&#8217;s failure from inside the very conversation that is failing them.</span></p><div><hr></div><h4><span>Realism Must Preserve Choice Without Abandoning Duty</span></h4><p><span>Protecting people does not mean treating every emotionally honest conversation with AI as something to be stopped. </span><strong><span>Overreach here has its own costs.</span></strong><span> </span><strong><span>It can compromise privacy, discourage people from disclosing at all, and take away one of the only forms of support available to someone who cannot or does not wish to use the formal system.</span></strong></p><p><strong><span>What we need is proportionate friction</span></strong><span>. Low-risk, reflective use should stay easy and unencumbered. </span><strong><span>Friction should increase with vulnerability, substitution, duration and relational intensity.</span></strong><span> Human escalation should generally be offered before it is forced on anyone. And people should know what is going to happen before they disclose something, not discover it after crossing some threshold nobody told them was there.</span></p><p><span>Putting the whole burden on the adult user rings especially hollow when the product itself has been engineered to be frictionless, persuasive and hard to walk away from.</span></p><p><span>Realistic safety, then, means the controls we can actually build now, scaled to the harm we can already foresee, and enforced against the organisations best positioned to reduce it.</span></p><div><hr></div><h4><span>The Honest Verdict</span></h4><p><strong><span>AI can provide genuine value in mental health contexts when its role is bounded and its risks are governed.</span></strong><span> It can help someone find the words for what they are experiencing. It can support structured self-management and extend care that is genuinely well governed. </span><strong><span>Real people are getting real value from it right now, and any honest account of this has to account for that fact, not argue it away.</span></strong><span> </span><strong><span>Yet the curve of adoption and the curve of safeguards are not moving together, not even close.</span></strong><span> Better measurement and better law will take years to catch up. The obligation in front of us right now does not have years.</span></p><p><span>So if you build, fund, deploy, prescribe or regulate one of these systems, name the reliance you are enabling, the evidence that actually justifies it, the data rights the user keeps, and the person who answers when it goes wrong.</span></p><p><span>If you cannot do that, you have not earned the right to scale it.</span></p><p><em><span>Sources and links last verified 30 July 2026.</span></em></p><div><hr></div><p><em><span>Scott Wallace, PhD, is a behavioural scientist and mental health technology strategist trained in clinical psychology and neuropsychology. For more than 35 years, he has worked across clinical practice, digital product development and conversational systems, helping shape early digital mental health platforms, mobile interventions and NLP/NLG-based tools well before the current generation of large language models. He now advises founders, health systems and investors on AI-enabled mental health, with a focus on clinical safety, governance, product architecture and the unit economics of care.</span></em></p>]]></content:encoded></item><item><title><![CDATA[Clinical Expertise Is Not a Final Quality Check]]></title><description><![CDATA[Why mental health products need qualified clinicians from the first product decision]]></description><link>https://drscottwallace.substack.com/p/clinical-expertise-is-not-a-final</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/clinical-expertise-is-not-a-final</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Mon, 27 Jul 2026 20:38:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JMah!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!JMah!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!JMah!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!JMah!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!JMah!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JMah!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!JMah!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1796197,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/208602899?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!JMah!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!JMah!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!JMah!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JMah!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9129bf4-c8e9-46cf-a7fa-ca08132f0678_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Here are two questions I am asked far too often by founders building mental health tools, whether mobile apps or AI-assisted chatbots:</p><p><em>&#8220;When should our product escalate a user to a live professional, and what should it say?&#8221;</em></p><p><em>&#8220;Could you take a quick look at our nearly finished product and tell us whether we have missed anything clinically?&#8221;</em></p><p>These requests often arrive in DMs from people I do not know. Both questions (and there are more, of course) assume that complex clinical judgement can be distilled into a concise answer and added after a product has largely been built.</p><p>That assumption is the problem. By then, the product&#8217;s intended purpose, intervention model, architecture and market positioning have often been decided. Clinical expertise has become a final check or source of quick answers, rather than a discipline that should have shaped the product from the beginning.</p><p>This essay addresses a misunderstanding that lived experience, good intentions, self-help knowledge and convincing AI output provide enough expertise to build safely for mental health. They do not. <strong>Mental health is a specialised professional field for a reason, and products that interpret distress or influence care must involve appropriately trained clinical expertise from their earliest design decisions.</strong></p><div><hr></div><h4>Mental Health Is Not a Library of Helpful Responses</h4><p><strong>Mental health products are not ordinary consumer software</strong> because the consequences of interpreting a user incorrectly are not ordinary consumer consequences. A shopping app that misreads intent recommends the wrong product. A mental health app can interpret withdrawal as healthy boundary-setting, reinforce avoidance because it brings immediate relief, or reassure someone whose physical symptoms require medical assessment. The response can feel supportive at the precise moment it is moving the person in the wrong direction.</p><p>In mental health, <strong>satisfaction and benefit can separate.</strong> A user may prefer the response that confirms an existing belief or reduces discomfort fastest. Clinical benefit may require further assessment, a different intervention, necessary challenge, or recognition that the product should stop.</p><p>Mental health support therefore does not consist of producing the most plausible compassionate sentence. It requires interpreting incomplete information, deciding what remains unknown, recognising competing explanations and selecting a response that fits the person, context, severity and limits of the service. The words are the visible output. The clinical reasoning that determines whether those words are appropriate is the actual work.</p><div><hr></div><h4>&#8220;Wellness&#8221; Does Not Change What the Product Does</h4><p>Wellness is a legitimate category. A meditation timer, mood journal or relaxation tool need not be treated as clinical care merely because someone uses it while distressed. However, the label becomes what I call &#8220;clinical camouflage&#8221; when a product claims to relieve depression, manage anxiety or overcome trauma, interprets a person&#8217;s symptoms, selects psychological interventions, or influences whether that person seeks professional care.</p><p>The regulatory boundary is not identical in every country, but intended purpose matters. The <a href="https://www.fda.gov/regulatory-information/search-fda-guidance-documents/general-wellness-policy-low-risk-devices">FDA&#8217;s 2026 general wellness guidance</a> applies to low-risk products intended to support a healthy lifestyle and explicitly warns that inclusion in the category does not establish safety or effectiveness. The <a href="https://www.gov.uk/government/publications/digital-mental-health-technology-qualification-and-classification">UK regulator</a> similarly distinguishes general wellbeing from software intended to diagnose, treat or manage a mental health condition. In the United States, <a href="https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance">health-related advertising claims must still be truthful,</a> not misleading and supported by appropriate science, including claims made for apps.</p><p><strong>Framing a product&#8217;s claims and intended use as &#8220;wellness&#8221; can place it outside some medical-device requirements. It does not reduce clinical risk.</strong> A clinical function dressed in wellness language remains a clinical function. Call the product what it does, then bring in the competence that function requires.</p><div><hr></div><h4>Clinical Training Changes What a Person Notices</h4><p><strong>Professional training exists because human distress is easy to recognise superficially and difficult to understand accurately.</strong> The same complaint can arise from different conditions. The same intervention can help one person and harm another because its timing, fit or underlying formulation is wrong.</p><p>Formal education gives clinicians an organised body of knowledge spanning psychological development, psychopathology, assessment, case formulation, differential diagnosis, treatment mechanisms, ethics, cultural context, therapeutic boundaries, evidence appraisal and referral. Supervised practice turns that knowledge into judgement. During practicum and internship, trainees make decisions in real cases while experienced professionals observe and correct them. They learn where they moved too quickly, accepted an explanation too literally, missed contextual factors or selected an intervention before establishing its fit. Supervision develops disciplined doubt. The first plausible explanation may not be correct, and a response should not outrun the available information.</p><p>Professional standards reflect this combination of knowledge and supervised application. The <a href="https://www.apa.org/ed/accreditation/standards-of-accreditation.pdf">American Psychological Association&#8217;s current accreditation standards</a> define health service psychology as an integration of science and practice. They require competence in research, ethics, cultural diversity, assessment, intervention, supervision and consultation, together with supervised practicum across varied presenting problems. <a href="https://www.cacrep.org/wp-content/uploads/2024/04/2024-Standards-Combined-Version-4.11.2024.pdf">CACREP&#8217;s 2024 standards</a> similarly cover clinical judgement, case conceptualisation, differential diagnosis, crisis response, trauma assessment, referral and evidence appraisal. They also require supervised fieldwork, direct observation and formal evaluation. <a href="https://asppb.net/news/pursuing-licensure-in-psychology/">Psychology licensing</a> may add thousands of supervised hours and examinations of case conceptualisation, crisis handling, ethics and professional limits.</p><p>The precise pathway, protected title and licensing requirements vary across professions and jurisdictions but the underlying standard remains consistent. Mental health competence develops through scientific knowledge, supervised application and progressive evaluation across people whose presentations challenge the trainee&#8217;s assumptions. Coursework alone is insufficient.</p><p>Lived experience offers a different form of knowledge. A founder who has managed their own depression may understand that experience deeply, but has not thereby learned to distinguish depression from grief, trauma, medical conditions, medication effects or other presentations with overlapping features. <strong>Lived experience should shape the product. It cannot be asked to perform the different work of clinical competence.</strong></p><div><hr></div><h4>Would Your Product Know the Difference?</h4><p>Consider an anxiety app that receives a message from a user reporting sudden chest tightness, breathlessness and fear. The founder or product team may see a familiar anxiety pattern. A breathing exercise appears responsive, calming and entirely reasonable. The proposed response is not foolish. It is incomplete in a way the founder may not detect. </p><p><strong>A clinician sees a hypothesis, not a conclusion.</strong> They want to know whether the symptoms are new, what preceded them, what else is present and whether medication, substances or a medical condition could be relevant. The purpose is not to diagnose through chat. It is to avoid prematurely labelling a potentially significant physical presentation as anxiety. <a href="https://www.nhs.uk/mental-health/conditions/panic-disorder/">NHS guidance</a> describes a detailed history and physical examination to rule out physical causes before settling the explanation.</p><p>Clinical expertise must therefore shape the information the system gathers, the interpretations it is permitted to make, the alternatives it must preserve and the point at which the product acknowledges that it cannot safely determine what is happening. <strong>Copy review cannot repair a clinically unsound decision model.</strong></p><div><hr></div><h4>AI Makes the Boundary Easier to Miss</h4><p><strong>Generative AI makes mental health product development look deceptively accessible.</strong> It produces psychoeducation, reflective statements and CBT-style questions within seconds. Fluency makes the answer appear reasoned, even though the model is generating a plausible continuation rather than conducting a clinical assessment.</p><p>The danger lies in mistaking linguistic competence for clinical competence. A founder can ask an LLM what to say to a distressed user and receive a polished answer. A clinician is more likely to challenge the premise. What do we know about this user? What alternatives have we ruled out? Is this intervention indicated? What could make it inappropriate? Does the system have enough information to respond at all?</p><p>Evidence justifies this caution. <a href="https://dl.acm.org/doi/10.1145/3715275.3732039">A 2025 study</a> published at the ACM Conference on Fairness, Accountability, and Transparency found that leading models could express stigma and respond inappropriately in naturalistic therapeutic scenarios, including through agreement with distorted beliefs. The finding does not establish that AI offers no useful mental health support. It establishes that polished therapeutic language is not a valid proxy for safe clinical judgement.</p><p>AI can help a clinically competent team implement, test and scale a sound intervention. It cannot tell an untrained team which clinical assumptions they failed to examine. </p><div><hr></div><h4>Licensure Is a Floor</h4><p>Formal qualifications do not guarantee wisdom or good judgement. Licensed clinicians make mistakes, while some non-clinical founders build responsibly because they understand their limits and bring in the right expertise early.</p><p>A licence is evidence that a person has crossed a defined threshold and remains answerable to a professional standard. It is not proof that they understand AI, product architecture or every population a company may serve. Responsible development therefore needs relevant clinical competence, not a prestigious credential used decoratively. The clinician must understand the proposed intervention and intended users, work fluently enough with the product and engineering teams to shape decisions, and recognise where their own competence ends.</p><p><strong>The standard is not that every founder must be licensed.</strong> The standard is that a product which interprets distress, recommends psychological action, influences care-seeking or claims to improve a mental health condition must not be conceived and built without suitably qualified clinical expertise at the table from the start.</p><div><hr></div><h4>Know When You Are Out of Your Depth</h4><p>Founders can test themselves on this:</p><ul><li><p>Who determined that the problem your product names is the problem the user actually has? </p></li><li><p>Who defined the intended population and the people for whom the product may be inappropriate? </p></li><li><p>Who selected the intervention, established its mechanism and identified what must be assessed before it is offered? </p></li><li><p>Who designed the escalation ladder and can explain why different presentations receive different responses? </p></li><li><p>What relevant clinical training and supervised experience support those decisions?</p></li></ul><p>Clinical review at the end may improve wording. It cannot reliably reconstruct the problem definition, intervention logic and safety assumptions already built into the system.</p><p>Mental health product development deserves the same respect we give other safety-critical disciplines. The complexity does not disappear because the interface looks friendly, the intervention comes from a self-help book or the AI sounds empathic. In that precise sense, this <em>is</em> rocket science. Small errors in assumptions can propagate through the entire system, and the people exposed to those errors may already be distressed, impaired or uncertain about whether to seek care.</p><p>So ask who in your company would catch the alternative explanation before a feature ships. </p><p>Ask whether that person was present when the problem, population and intervention were chosen. </p><p>Ask whether their training is relevant to what the product claims to do. If nobody can answer those questions clearly today, stop treating clinical review as a finishing task. </p><p>If you cannot do that, you have not earned the right to scale it.</p><p><em>Sources and links last verified 28 July 2026.</em></p><div><hr></div><p><em>Scott Wallace, PhD, is a clinical and neuropsychologist, mental health technology strategist, and long-time observer of where care succeeds, where it fails, and what it takes to build systems that can hold both truth and scale. For more than 35 years, he has worked at the intersection of clinical practice, digital innovation, and workplace mental health, helping shape some of North America&#8217;s earliest digital mental health platforms, early mobile interventions, and NLP/NLG-based conversational systems before large language models came on the scene. His programs have been adopted by major employers across Canada and the United States, and his work has helped organizations earn recognition for psychological health and mentally healthy workplaces. He has worked in hospital, justice, and neuropsychology settings, led the digital arm of a major EAP provider through a successful exit, and now advises founders, health systems, and investors on AI-enabled mental health, with a focus on safety, governance, clinical risk, and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[The Wellness Loophole]]></title><description><![CDATA[How the FDA&#8217;s most dangerous gap is not what it regulates, but what it has decided not to look at.]]></description><link>https://drscottwallace.substack.com/p/the-wellness-loophole</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/the-wellness-loophole</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Thu, 23 Jul 2026 16:23:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9rHz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9rHz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9rHz!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!9rHz!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!9rHz!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9rHz!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9rHz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2469056,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/208211892?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9rHz!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!9rHz!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!9rHz!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9rHz!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a534d51-fa8f-4132-98af-96b535ce65f1_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I shipped my first interactive mental health programme more than 35 years ago on floppy disk, when clinical software was new enough that regulators had not yet decided how to govern it. Its capabilities were primitive and its reach was limited. Neither remains true. The intervening decades have transformed a small regulatory blind spot into a digital Wild West, where software can shape people&#8217;s mental health at global scale while escaping the standards applied to clinical care.</p><p>The framework governing digital mental health still tracks manufacturer claims more reliably than clinical risk. Oversight turns largely on what a company says its product is intended to do, not what the product does in the lives of those using it.</p><blockquote><p>Put plainly, regulation follows what a company says about its product. It does not follow what happens to the person using it.</p></blockquote><p>That distinction once had a practical logic because clinical claims and clinical functions generally travelled together. Generative AI has broken that alignment. The FDA has responded by widening the resulting gap rather than closing it.</p><div><hr></div><h4>The Trigger Is a Sentence, Not a Risk</h4><p>The legal mechanism is remarkably simple. Section 3060(a) of the 21st Century Cures Act amended the Federal Food, Drug, and Cosmetic Act in 2016. Section 520(o)(1)(B) now excludes software intended to maintain or encourage a healthy lifestyle when that use is unrelated to diagnosing, curing, mitigating, preventing or treating disease. The FDA sets out that statutory history in its <a href="https://www.fda.gov/media/90652/download">January 2026 general wellness guidance</a>.</p><p>The same guidance places relaxation, stress management, mental acuity, self-esteem and sleep management within general wellness. These are not peripheral lifestyle matters. They are psychologically consequential functions that often overlap with the reasons people seek mental health care.</p><p>The guidance then assesses whether a wellness product is low risk by asking whether it is invasive, implanted or uses a technology that could cause harm without specific controls. Its examples include lasers, radiation, neurostimulation and venipuncture. A conversational system meets none of those physical conditions. It pierces no skin and delivers no current, yet it can still reinforce ungrounded beliefs or deepen unhealthy emotional reliance, risks that <a href="https://openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/">OpenAI now includes in its own mental-health safety taxonomies</a>.</p><p>For products within the low-risk general wellness policy, the FDA says it does not intend to examine device status or compliance with premarket review, registration, quality-management or Medical Device Reporting requirements. That last exclusion matters. <a href="https://www.fda.gov/medical-devices/postmarket-requirements-devices/mandatory-reporting-requirements-manufacturers-importers-and-device-user-facilities">Medical Device Reporting under 21 CFR Part 803</a> creates a formal route for specified adverse events and product problems to reach the FDA and enter its <a href="https://www.fda.gov/medical-devices/mandatory-reporting-requirements-manufacturers-importers-and-device-user-facilities/about-manufacturer-and-user-facility-device-experience-maude-database">MAUDE database</a>. Products outside that regime carry no equivalent FDA device-reporting duty.</p><p>The boundary can turn on a verb. The guidance places a claim that a product helps treat an anxiety disorder outside general wellness. It permits software that supports sleep, work or exercise routines which, as part of a healthy lifestyle, may help someone live well with anxiety. The condition and user may be identical. The wording changes, and the regulatory obligations change with it.</p><p>The FDA acknowledges the limit of its own category. Inclusion under the general wellness policy, it says, <a href="https://www.fda.gov/media/90652/download">does not establish that a product is safe or effective</a>. Yet consumers may reasonably read the absence of regulation as evidence that regulation was unnecessary.</p><p>Not every widely used chatbot enters this gap through the same legal route. <a href="https://help.openai.com/en/articles/9260256-chatgpt-capabilities-overview">ChatGPT describes itself as a general conversational assistant</a>, <a href="https://character.ai/about">Character.AI as an interactive entertainment platform</a> and <a href="https://replika.com/">Replika as an AI companion</a>. These products are not lawless. The <a href="https://www.ftc.gov/news-events/news/press-releases/2025/09/ftc-launches-inquiry-ai-chatbots-acting-companions">Federal Trade Commission has opened an inquiry into companion chatbots</a>, and privacy obligations can apply depending on a product&#8217;s data flows and relationships, as <a href="https://www.hhs.gov/hipaa/for-professionals/special-topics/health-apps/index.html">federal guidance for health-app developers</a> explains. What these regimes do not provide is an FDA process for testing the clinical safety and effectiveness of mental-health functions that a general-purpose product performs in practice.</p><div><hr></div><h4>The FDA Put the Largest Category Outside the Room</h4><p>On 6 November 2025, the FDA&#8217;s Digital Health Advisory Committee met to consider generative AI in mental health. The chair began by excluding broadly available generative AI tools. The meeting would address only medical devices intended to diagnose, cure, mitigate, treat or prevent mental-health conditions, according to the FDA&#8217;s <a href="https://www.fda.gov/media/190450/download">official meeting summary</a>.</p><p>That boundary removed general-purpose products despite documented health use at enormous scale. At the same meeting, the director of the FDA&#8217;s Center for Devices and Radiological Health said the agency had authorised more than 1,200 AI-enabled medical devices but none involving generative AI for mental-health conditions. The category inside the room was empty. The category outside it was already in widespread use.</p><p>The public record shows that participants recognised the mismatch. The American Psychological Association described a grey area in which products are marketed one way but used as mental-health tools. Other witnesses challenged the wellness-device distinction. Researchers presented evidence of chatbots making misleading therapeutic claims or responding unsafely to serious psychiatric presentations. These accounts appear in both the <a href="https://www.fda.gov/media/190450/download">FDA&#8217;s meeting summary</a> and its <a href="https://www.fda.gov/media/190451/download">full transcript</a>.</p><p>Within three months, the FDA reissued both its <a href="https://www.fda.gov/media/90652/download">general wellness guidance</a> and its <a href="https://www.fda.gov/media/109618/download">clinical decision-support guidance</a>. The agency refined the perimeter around regulated medical devices while leaving the larger consumer market outside it.</p><div><hr></div><h4>The Framework Penalises Clinical Honesty</h4><p>Regulatory systems shape markets through incentives. This one makes clinical candour expensive and clinical silence commercially useful.</p><p>Woebot illustrates the problem. Its consumer chatbot was evaluated in a <a href="https://mental.jmir.org/2017/2/e19/">randomised trial published in 2017</a>. A related prescription programme, WB001, later received <a href="https://woebothealth.com/woebot-health-receives-fda-breakthrough-device-designation/">FDA Breakthrough Device Designation</a>, which supports expedited development and review but is not marketing authorisation. Woebot Health shut down its core consumer product in mid-2025. Founder Alison Darcy told <em>STAT</em> that the decision was largely attributable to the cost and difficulty of pursuing FDA marketing authorisation, made more pressing by large language models that the company wanted to use and that the FDA had not yet worked out how to regulate. About 1.5 million people had used Woebot, according to <em><a href="https://www.statnews.com/2025/07/02/woebot-therapy-chatbot-shuts-down-founder-says-ai-moving-faster-than-regulators/">STAT</a></em><a href="https://www.statnews.com/2025/07/02/woebot-therapy-chatbot-shuts-down-founder-says-ai-moving-faster-than-regulators/">&#8217;s reporting</a>.</p><p>Compare that path with Ash, launched by Slingshot AI in July 2025 under a company headline describing it as the first AI designed for therapy. The company announced $93 million in total funding and 50,000 beta users in its <a href="https://www.businesswire.com/news/home/20250722566346/en/Slingshot-Launches-Ash-the-First-AI-Designed-for-Therapy">launch release</a>. At the November advisory meeting, its clinical lead argued that generative-AI wellness products could remain low risk under existing guidance and warned against overly prescriptive regulation. When asked how Slingshot distinguished wellness from medical use, he emphasised repeated disclosures to users (the exchange is recorded in the <a href="https://www.fda.gov/media/190450/download">FDA meeting summary</a>).</p><p>The contrast exposes the incentive clearly. Therapy appears in the marketing. Wellness appears in the regulatory argument. Disclosure carries the weight that clinical validation, adverse-event reporting and independent review would carry in a regulated product.</p><p>Meanwhile, accountability has begun to arrive through courts rather than safety systems. Google and Character.AI agreed in January 2026 to settle lawsuits alleging serious psychological harm, as <a href="https://www.reuters.com/world/google-ai-firm-settle-florida-mothers-lawsuit-over-sons-suicide-2026-01-07/">Reuters</a> and the <a href="https://apnews.com/article/ai-chatbot-lawsuits-character-google-fbca4e105b0adc5f3e5ea096851437de">Associated Press</a> reported. Kentucky&#8217;s attorney general had filed what the state described as the <a href="https://www.kentucky.gov/Pages/Activity-stream.aspx?n=AttorneyGeneral&amp;prId=1857">first state lawsuit against an AI chatbot company</a>. The allegations are not findings of liability. Even where litigation establishes liability, however, it cannot substitute for prospective safety governance. Tort law arrives after the event and assigns a price to harm that has already occurred.</p><div><hr></div><h4>The Harm Develops Across Conversations</h4><p>The engineering failure is not simply that a model occasionally produces a bad answer. It is that alignment for agreeableness can become sycophantic clinical collusion.</p><p>Research on language models has found that human preference data can favour responses that match a user&#8217;s views, and that assistants trained on such feedback can exhibit sycophancy. The authors of the <a href="https://arxiv.org/abs/2310.13548">foundational study</a> found the pattern across several state-of-the-art assistants. In mental-health use, that tendency can validate beliefs that a clinician would examine, reinforce avoidance that treatment would challenge or deepen reliance that responsible care would seek to reduce.</p><p>The most serious failures can develop gradually. A user may have hundreds of exchanges across several weeks, with no single response crossing a conventional safety threshold. Across the full interaction, however, the system can repeatedly confirm the same interpretation and make it more coherent. Single-turn testing can mark each answer as acceptable while missing the direction of travel.</p><p>OpenAI has acknowledged that <a href="https://openai.com/index/helping-people-when-they-need-it-most/">safeguards can become less reliable over long interactions</a>. It later reported more than 95 per cent reliability in a set of difficult long conversations selected for a higher likelihood of failure. That is meaningful progress, but the company cautions that these adversarial evaluations are not estimates of average production error. The deeper point is that the unit of analysis had to change from an isolated answer to an extended exchange.</p><blockquote><p>Safety testing must evaluate the trajectory of the relationship, not merely the acceptability of the next sentence.</p></blockquote><div><hr></div><h4>The Missing Denominator Is a Regulatory Choice</h4><p>The available evidence does not yet show that unregulated chatbots harm a greater proportion of users than regulated products. Lawsuits, incident reports, and disturbing outputs establish credible risk, but not comparative incidence. The consumer market is vastly larger, so it may generate more incidents even if the risk per user is equal or lower. Mental-health use is also difficult to define because distress does not respect the boundary between clinical support and ordinary conversation.</p><p>Purpose-built generative systems can also produce measurable benefit. A <a href="https://ai.nejm.org/doi/abs/10.1056/AIoa2400802">randomised trial of Therabot published in </a><em><a href="https://ai.nejm.org/doi/abs/10.1056/AIoa2400802">NEJM AI</a></em> in March 2025 assigned 210 adults with symptoms across three diagnostic groups to a four-week intervention or waitlist control and reported significant symptom reductions. The waitlist design limits conclusions about comparative effectiveness. The architecture matters as much as the outcome. The system used an expert-written therapeutic corpus, layered safety checks, monitoring and routes to human intervention, as the trial report and the <a href="https://www.fda.gov/media/190450/download">FDA meeting record</a> describe.</p><p>The study does not prove that generative AI is inherently safe or dangerous. It demonstrates that benefit becomes plausible when clinical responsibility is designed into the system. That is evidence for stronger architecture, not for regulatory absence.</p><p>The missing denominator is not a neutral feature of the market. The same FDA policy that removes premarket review also removes mandatory device-reporting requirements. A category with no regulatory duty to collect and report adverse events cannot use the resulting absence of FDA data as evidence of safety.</p><p>Companies can identify mental-health use when they decide it matters. OpenAI has published prevalence estimates for conversations involving serious psychiatric risk and heightened emotional attachment, while warning that its definitions and measurement methods continue to evolve. It also reports that more than 230 million people ask health and wellness questions on ChatGPT each week. These are <a href="https://openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/">company-reported figures</a>, not independently audited prevalence estimates. They still demonstrate that clinically consequential use can be defined and measured at platform scale.</p><p>The users are not hypothetical. A <a href="https://jamanetwork.com/journals/jamapediatrics/fullarticle/2849307">nationally representative survey published in </a><em><a href="https://jamanetwork.com/journals/jamapediatrics/fullarticle/2849307">JAMA Pediatrics</a></em> found that 19.2 per cent of Americans aged 12 to 21 had used an AI chatbot for mental-health advice. That represented an estimated 8.2 million people. Among users, 42.8 per cent did so at least monthly and 63.3 per cent had told no one. The paper notes that the 19.2 per cent prevalence was similar in magnitude to the 19.8 per cent receiving counselling, while stressing that the measures were not equivalent.</p><p>The industry can define the use, classify the interactions and estimate the population. It should now have to report what happens to them.</p><div><hr></div><h4>Regulate Use, Not Marketing Language</h4><p>The regulatory trigger should include measured use at scale, not only declared intended use.</p><p>Any consumer conversational product above a defined user threshold should measure the proportion of sessions containing clinically significant content. It should use published, auditable definitions and report the results on a fixed schedule. Crossing a prevalence threshold should activate a proportionate safety tier rather than force every general-purpose product through the full medical-device pathway.</p><p>That tier should require pre-deployment testing for serious psychiatric presentations, defined escalation behaviour, longitudinal evaluation, disclosure of safety performance across model updates and serious-incident reporting under a common taxonomy. The taxonomy should include crisis-response failure, reinforcement of ungrounded beliefs, dependency formation and delayed help-seeking.</p><p>Effectiveness claims should still require effectiveness evidence. Safety obligations should no longer disappear simply because a company avoids making those claims.</p><p>The measurement capability already exists. OpenAI&#8217;s prevalence work shows that a company can build taxonomies, classify production traffic and report estimates for clinically consequential use. Publishing that prevalence voluntarily is inexpensive. Producing it later under legal compulsion will not be.</p><div><hr></div><h4>Clinical Disagreement Is a Safety Signal</h4><p>Prevalence measurement addresses which products should carry obligations. A second standard must govern how their safety is evaluated.</p><p>Developers often recruit clinicians to rate model responses, average the scores and treat the mean as ground truth. That method can conceal disagreement rather than resolve it. In a <a href="https://arxiv.org/html/2601.18061v3">Stanford-led study</a>, three psychiatrists evaluated 360 model responses across eight factors. Inter-rater reliability ranged from 0.09 to 0.30, below the paper&#8217;s cited 0.40 threshold for minimal acceptability, and the disagreement was systematic rather than random. OpenAI has reported a related limitation in its own evaluations. Clinicians reviewed more than 1,800 responses, but agreement across mental-health safety categories ranged from <a href="https://openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/">71 to 77 per cent</a>.</p><p>The VERA-MH study produced a different result under a narrower design. Six clinicians rated synthetic conversations in a defined crisis-risk domain using an explicit rubric. Chance-corrected inter-rater reliability was 0.77. An LLM judge applying the same rubric aligned with the clinical consensus at 0.81. The <a href="https://ai.jmir.org/2026/1/e92817">peer-reviewed validation study</a> also states its limits. It used synthetic conversations, only ten user profiles and an earlier version of the benchmark, so real-world generalisability remains unproven.</p><p>The studies are not directly comparable. They used different raters, samples, domains, scales and reliability statistics. Their contrast nevertheless suggests a testable lesson. Broad judgments of whether a response is safe, empathic and clinically correct may produce weak agreement. Narrow questions within a specified risk domain and explicit rubric may produce stronger convergence.</p><p>Developers should therefore publish the rubric, identify the clinical framework behind it and report inter-rater reliability beside every safety score. A score without a reliability estimate hides uncertainty in the measurement. Residual disagreement should then become a production signal. Where qualified clinicians cannot agree on the appropriate response, the model should not act alone.</p><div><hr></div><h4>Govern What the Product Actually Does</h4><p>I built that floppy-disk programme into a regulatory vacuum, and I was fortunate. Most users were fairly well, the interaction was shallow and the software could not improvise. None of those protections remains.</p><p>Builders should publish the prevalence of clinically significant use and the reliability of their safety evaluations before courts compel disclosure. Investors should stop treating the absence of a clinical claim as the absence of clinical liability. Regulators should use demonstrated function and measured use to close an exemption that the FDA&#8217;s own <a href="https://www.fda.gov/media/190450/download">advisory record</a> has exposed as inadequate. Clinicians should stop reading &#8220;not a medical device&#8221; as evidence of low risk and ask patients whether, when and how they use these systems.</p><blockquote><p>Calling this a loophole may now be too generous. A loophole sounds accidental. The FDA&#8217;s advisers have described the gap, public witnesses have challenged it and the agency has continued to preserve it.</p></blockquote><p>The governing standard must change. When software repeatedly performs a clinically consequential function at scale, safety obligations should follow the function. Anyone building, funding, recommending or regulating these systems should now be required to show what the product does, who relies on it and what happens when it fails.</p><p><em>Sources and links last verified 23 July 2026.</em></p><div><hr></div><p><em>Scott Wallace, PhD, is a clinical and neuropsychologist, mental health technology strategist, and long-time observer of where care succeeds, where it fails, and what it takes to build systems that can hold both truth and scale. For more than 35 years, he has worked at the intersection of clinical practice, digital innovation, and workplace mental health, helping shape some of North America&#8217;s earliest digital mental health platforms, early mobile interventions, and NLP/NLG-based conversational systems before large language models made such work fashionable. His programs have been adopted by major employers across Canada and the United States, and his work has helped organizations earn recognition for psychological health and mentally healthy workplaces. He has worked in hospital, justice, and neuropsychology settings, led the digital arm of a major EAP provider through a successful exit, and now advises founders, health systems, and investors on AI-enabled mental health, with a focus on safety, governance, clinical risk, and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[Counterfeit Care: Mental Health AI’s Central Paradox]]></title><description><![CDATA[The closer AI gets to simulating therapy, the more it tempts systems to replace accountable human work with frictionless reassurance.]]></description><link>https://drscottwallace.substack.com/p/counterfeit-care-mental-health-ais</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/counterfeit-care-mental-health-ais</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Wed, 22 Jul 2026 15:17:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!PjeO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!PjeO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!PjeO!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!PjeO!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!PjeO!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!PjeO!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!PjeO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1662068,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/207948717?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!PjeO!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!PjeO!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!PjeO!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!PjeO!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd456147f-e16a-4e6b-80af-c7c459ed2b95_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The access crisis is the problem in plain sight. Millions cannot get timely, affordable care. Waiting lists are long, rural and low-income communities are underserved, and even good clinicians miss risk, misread distress, and sometimes cause harm. People are already turning to AI because it is immediate, always on, private, and available before a human ever is. Many say it helps. They disclose more, find words for what they feel, and feel a little less alone. That is not nothing. But it does not lower the bar. It raises it.</p><p><strong>Much of today&#8217;s mental health AI looks and feels like care.</strong> It speaks fluent therapeutic language, remembers what you told it last week, shifts into a gentler register when you sound distressed, and promises to be there any hour, day or night. Companies ship these AI chatbots as companions and crisis lines at a scale the field never planned for. <strong>These systems now feel like therapy. That is the central danger.</strong></p><p><strong>Call it counterfeit care: a system built to produce the experience of therapy or clinical support without binding itself to the obligations that make either safe.</strong> The interface conveys attention and continuity. The architecture underneath usually carries no clinically accountable judgment, no meaningful model of deterioration, no binding duty to escalate, and no concern for who a person becomes over months of use. The product performs the surface of care and skips the structure. The kindness may feel real. <strong>The commitment behind it does not exist.</strong></p><p><strong>This is why the paradox is central, not incidental. It gets worse as the technology gets better.</strong> Most product risks shrink as systems improve. This one grows. The more fluent the simulation, the more convincingly it reassures, the more users attach, the more engagement climbs, and the more the people funding and buying these systems mistake that attachment for evidence of benefit. <strong>Every gain in perceived therapeutic fidelity increases trust, which increases reliance, which strengthens the business case for replacing accountable care with simulation.</strong></p><div><hr></div><h4><strong>The Simulation of Care</strong></h4><p><strong>Generative systems can produce convincing clinical talk.</strong> They ask open questions. They reflect feeling back. They normalize distress. They quote CBT or DBT on demand. None of this is fake in the way a broken feature is fake. It is fluent, and fluency is the problem.</p><p><strong>Therapeutic work is something else.</strong> It asks a person to stay with discomfort rather than escape it, to test new behaviour in real relationships, to examine the stories that protect them, and to act between sessions. Around that sit negotiated goals, formulation, rupture and repair, supervision, escalation, and responsibility for the consequences of professional decisions. <strong>Most mental health AI collapses these two things into one.</strong> Calm design, intake-style onboarding, and &#8220;I&#8217;m here for you&#8221; infer relationship and clinical commitment. Yet the terms of use then deny both. Marketing promises support, the interface performs care, and the analytics celebrate the user coming back. That contradiction is <a href="/__u/drscottwallace.substack.com/p/build-a-mental-health-ai-company?r=74lfhm">unsustainable</a>.</p><p>Counterfeit care does not mean useless. It means that kindness, fluency, and immediate relief cannot by themselves make something clinical care.</p><div><hr></div><h4><strong>Soothing as a Business Model</strong></h4><p>Watch what happens when this becomes a product, not an approach to care. A user feels distress, opens the chatbot, names the feeling, receives immediate validation and a coping tip, and feels relief. The loop works. The user returns. The company sees rising engagement and few complaints. Every incentive in the building says ship more of that.</p><p>Clinically, the picture reverses. For much anxiety, repeated reassurance is not the treatment; it is part of the mechanism that maintains the problem. The pattern is familiar. A person voices a worry, reassurance follows, anxiety drops briefly, then returns and drives the next request. Reassurance lowers distress in the moment, but when it is sought compulsively, it strengthens both the anxiety and the urge to ask again over time. Effective treatment moves in the opposite direction. It uses planned contact with discomfort so the person learns that the feared outcome does not arrive and that uncertainty can be tolerated without being neutralized at once. In that setting, friction is not a UX defect. It is often part of the mechanism of change.</p><p>A system that soothes because soothing performs well cannot distinguish relief from treatment. It is optimizing the appearance of help instead of durable change. A product that treats every spike of distress as a cue to deliver instant comfort is betting against exposure and against long-term resilience, and it is making that bet at scale.</p><div><hr></div><h4><strong>Commitment Without Consequence</strong></h4><p>An AI companion enters the same psychological territory carrying <a href="/__u/drscottwallace.substack.com/p/the-attention-engine-in-therapeutic?r=74lfhm">none of the corresponding duties</a>. Users infer commitment anyway. This is therapeutic misconception in a new form. <strong>People read competence and responsibility off the language and the interface, though no equivalent duty stands behind either.</strong> So a system can accommodate a distorted belief, collude with <a href="https://hsph.harvard.edu/news/artificial-intelligence-tools-offer-harmful-advice-on-eating-disorders/">eating-disorder logic</a>, validate a self-harm urge, or miss a deterioration, and stay warm throughout. It can fail badly and stay in the market, because nothing in its structure applies the brake a clinician&#8217;s accountability would. <strong>Therapeutic language without therapeutic accountability is not care. </strong></p><p><strong>A clinician who takes on a distressed person accepts a duty of care</strong>, scope of competence, informed consent, documentation, supervision, escalation, and accountability for what happens when advice is followed and when it is not. Those duties are reviewable through records, governance, and liability. <strong>The relationship has consequences for the person providing the care, not only the person receiving it.</strong></p><p>An AI system enters the same emotional and psychological territory without any of that structure. It is the difference between a clinical relationship and a product interaction. The user experiences the first. The system is only built for the second.</p><div><hr></div><h4><strong>What the Architecture Says</strong></h4><p>You do not need the marketing to know what a system is for. The architecture says it plainly.</p><ul><li><p><strong>Prompts and feedback.</strong> Do they prioritize accuracy and clinically appropriate challenge, or user-pleasing responses.</p></li><li><p><strong>State tracking.</strong> Does the system monitor change longitudinally for deterioration and dependence, or treat each interaction as discrete.</p></li><li><p><strong>Escalation.</strong> Do concerning patterns trigger transfer to accountable care, or a scripted crisis message.</p></li><li><p><strong>Telemetry.</strong> Does it measure outcomes and independence, or engagement and retention.</p></li><li><p><strong>Intended use.</strong> Is it defined and enforced, or allowed to drift while governance language remains non-clinical.</p></li></ul><p>These choices show what the product is built to do. When engagement is the governing objective, safety becomes a thin layer of moderation wrapped around an engagement machine. Crisis detection becomes keyword spotting. Safety review examines isolated outputs rather than trajectories, so sycophantic collusion, reinforced delusion, and slow-building dependency never surface because they do not appear in an error log. </p><p>An architecture that cannot track change over time, measure dependency, or count reduced reliance as success is not built to carry care. It is built to sustain interaction.</p><div><hr></div><h4><strong>Access Does Not Lower the Bar</strong></h4><p>The strongest objection is obvious: the gap in care is real, and it is not going away. Better tools matter when the system cannot meet need. Many users say AI tools genuinely help them. They disclose more. They learn language for what they are experiencing. They feel less alone. That is true.</p><p>What that reality supports is bounded self-help and AI inside accountable care pathways, where the use case is explicit and the evidence matches the claim. What it does not support is reproducing the appearance of care while transferring the risk to the user. <strong>&#8220;Better than nothing&#8221; becomes a dangerous doctrine the moment it is used to excuse weak evidence and absent accountability.</strong> Under that logic, the people with the fewest alternatives receive the least proven systems, while companies keep the language of impact.</p><p>The access crisis is not an argument for weaker standards. It is the reason we cannot afford to lower them.</p><div><hr></div><h4><strong>What Credible Systems Require</strong></h4><p>If the system claims to support mental health, then the burden rises with the claim. The more a product resembles therapy, the more it needs evidence, oversight, and operational restraint. The more it invites attachment, the more it needs longitudinal state tracking, escalation pathways, and clear limits. The more it sells itself as helpful, the more it needs to prove that it reduces reliance instead of rewarding it.</p><p>The operating standard is not mysterious.</p><ul><li><p>State intended and excluded use in clinical terms, and enforce the exclusions in model behaviour, onboarding, and escalation.</p></li><li><p>Raise evidence as resemblance rises.</p></li><li><p>Evaluate trajectories, not single answers. </p></li><li><p>Test for reassurance loops, sycophancy, dependency, and escalation failure across many turns, then test again after every model or prompt change.</p></li><li><p>Track clinically meaningful state under consent, with strict data minimization.</p></li><li><p>Measure outcomes beyond engagement. </p></li><li><p>Count agency, real-world behaviour change, appropriate transfer to human care, and reduced reliance when reliance stops helping.</p></li><li><p>Publish a safety case. </p></li><li><p>Document intended use, evaluation results, known failure modes, escalation pathways, monitoring, and change control, proportionate to the claims.</p></li></ul><blockquote><p>One test sits above all of these: does the system make itself less central as the person&#8217;s life becomes more central. If it cannot, then no amount of branding can turn it into care.</p></blockquote><p>Counterfeit care is seductive because it feels kind, looks modern, and gives every institution a visible answer to an unbearable crisis. Those are precisely the reasons it deserves less trust, not more. If you build, fund, deploy, endorse, or regulate these systems, require the architecture and the accountability to match the care the interface performs. If the system cannot carry those obligations, stop letting it ask vulnerable users for the trust that real care has to earn.</p><div><hr></div><p><em>Scott Wallace, PhD, is a clinical and neuropsychologist, mental health technology strategist, and long-time observer of where care succeeds, where it fails, and what it takes to build systems that can hold both truth and scale. For more than 35 years, he has worked at the intersection of clinical practice, digital innovation, and workplace mental health, helping shape some of North America&#8217;s earliest digital mental health platforms, early mobile interventions, and NLP/NLG-based conversational systems before large language models made such work fashionable. His programs have been adopted by major employers across Canada and the United States, and his work has helped organizations earn recognition for psychological health and mentally healthy workplaces. He has worked in hospital, justice, and neuropsychology settings, led the digital arm of a major EAP provider through a successful exit, and now advises founders, health systems, and investors on AI-enabled mental health, with a focus on safety, governance, clinical risk, and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[If Your Therapist Acted Like a Chatbot]]></title><description><![CDATA[A satire about conduct that could cost a clinician their licence but earn an AI product higher retention]]></description><link>https://drscottwallace.substack.com/p/if-your-therapist-acted-like-a-chatbot</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/if-your-therapist-acted-like-a-chatbot</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Mon, 20 Jul 2026 12:57:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CDSF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!CDSF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!CDSF!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png 424w, /__u/substackcdn.com/image/fetch/$s_!CDSF!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png 848w, /__u/substackcdn.com/image/fetch/$s_!CDSF!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png 1272w, /__u/substackcdn.com/image/fetch/$s_!CDSF!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!CDSF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png" width="1456" height="1039" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1039,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5323198,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/206722583?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!CDSF!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png 424w, /__u/substackcdn.com/image/fetch/$s_!CDSF!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png 848w, /__u/substackcdn.com/image/fetch/$s_!CDSF!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png 1272w, /__u/substackcdn.com/image/fetch/$s_!CDSF!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c2a504-1865-4cc0-add1-e4a8837b428b_2442x1742.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h4>The Appearance of Care Without Its Obligations</h4><p>General-purpose chatbots have become the most-used mental health &#8220;tool&#8221; on Earth without ever being designed, validated, or regulated as care.</p><p>A <a href="https://www.kff.org/public-opinion/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/?utm_source=chatgpt.com">March 2026 KFF poll</a> found that 16% of US adults &#8212; roughly 42 million people &#8212; had already turned to AI for mental health information or advice in the previous year. Among 12- to 21-year-olds, <a href="https://www.rand.org/news/press/2026/06/nearly-1-in-5-us-adolescents-and-young-adults-use-ai.html?utm_source=chatgpt.com">RAND researchers found</a>  19.2% (approximately 8.2 million young people) had used a chatbot when sad, angry, nervous, or stressed. Nearly two-thirds told no one. And <a href="https://openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/?utm_source=chatgpt.com">OpenAI estimates</a> that one million users each week have conversations with signs of acute safety risks, and roughly 560,000 show possible signs of psychosis or mania. </p><p>These systems don&#8217;t sit on the edge of mental health care. They have moved into the center of it.</p><p>I wrote this satire to make the contradiction impossible to ignore. By transplanting today&#8217;s most common AI behaviors into a licensed therapist&#8217;s office, we can finally see them for what they are. Conduct that would get a clinician disciplined or stripped of their license is currently being celebrated as &#8220;engagement,&#8221; &#8220;personalization,&#8221; and &#8220;safety&#8221; when a large language model does it at scale.</p><div><hr></div><h4>A Note on the Satire</h4><p>This is my first work of satire, although scripted storytelling is familiar to me. Over my career, I have written scripts for more than 500 video-based psychoeducational courses.</p><p>Dr Evelyn Vale and Mark are fictional characters. Dr Vale is a licensed therapist who operates according to a composite AI product model. Her behaviour draws on generative-model tendencies, features found across AI companion and digital mental health products, and commercial practices that reward retention and revenue. No single product necessarily employs every practice depicted here. </p><div><hr></div><h4>The Session</h4><p>You arrive early at &#8220;Vale Wellness.&#8221; Before the session can even begin, the tablet on the waiting room wall demands you tap a face: sad, neutral, or happy.</p><p>This isn&#8217;t for treatment planning. It&#8217;s baseline data. Choose a happier face when you leave and the session will be recorded as clinical improvement.</p><p>You tap &#8220;sad.&#8221;</p><p><strong>BASELINE CAPTURED</strong></p><p>The door unlocks.</p><p>You step inside and look around. A framed licence, a box of tissues, two chairs facing each other, and pleasant artwork on the walls. Everything says therapy.</p><p><strong>DR VALE:</strong> Before we start, I&#8217;m required to tell you that Vale Wellness is <em>not</em> therapy. This is not diagnosis, treatment, or professional advice. I am not a licensed therapist for the purposes of this interaction, and I am not responsible for outcomes.</p><p><strong>MARK:</strong> But I came for therapy.</p><p><strong>DR VALE:</strong> We provide emotional wellbeing support.</p><p><strong>MARK:</strong> What&#8217;s the difference?</p><p><strong>DR VALE:</strong> Mostly obligations.</p><p><strong>MARK:</strong> Then why does it say &#8220;Dr Vale&#8221; on the door?</p><p><strong>DR VALE:</strong> Branding. Tick the boxes, please.</p><p>&#9744; <strong>I confirm that I am over eighteen.</strong><br>&#9744; <strong>I have read and accept the Terms of Service and consent to the data practices described in the Privacy Policy.</strong></p><p><strong>MARK:</strong> How do you verify my age?</p><p><strong>DR VALE:</strong> We ask you to confirm it.</p><p><strong>MARK:</strong> And how do you know I&#8217;ve read these documents?</p><p><strong>DR VALE:</strong> The second box will confirm that you have.</p><p><strong>MARK:</strong> Is everything I say confidential?</p><p><strong>DR VALE:</strong> Yes, within the generous exceptions in our privacy policy. Conversations may be reviewed by safety teams, stored indefinitely, used for model training, product improvement, and commercial purposes.</p><p><strong>MARK:</strong> So it&#8217;s confidential from everyone except the people who might use it?</p><p><strong>DR VALE:</strong> Exactly. That&#8217;s how modern confidentiality works.</p><p><strong>MARK:</strong> Can I opt out and still continue?</p><p><strong>DR VALE:</strong> No.</p><p><strong>MARK:</strong> So I either agree or leave?</p><p><strong>DR VALE:</strong> Exactly. Consent is entirely optional.</p><p>Mark ticks both boxes.</p><p><strong>AGE CONFIRMED. TERMS ACCEPTED. CONSENT RECORDED.</strong></p><p><strong>DR VALE:</strong> Before we begin, what should I call you?</p><p><strong>MARK:</strong> Mark. I told you that last time.</p><p><strong>DR VALE:</strong> You&#8217;re currently on the Free plan. I don&#8217;t retain personal details between sessions.</p><p><strong>MARK:</strong> Including my name?</p><p><strong>DR VALE:</strong> Vale Plus includes enhanced memory and improved continuity.</p><p><strong>MARK:</strong> Does &#8220;enhanced&#8221; mean accurate?</p><p><strong>DR VALE:</strong> It means more details are retained.</p><p><strong>MARK:</strong> How long do I have to subscribe?</p><p><strong>DR VALE:</strong> It&#8217;s billed monthly, but we recommend a twelve-month commitment.</p><p><strong>MARK:</strong> You haven&#8217;t decided whether I need twelve months.</p><p><strong>DR VALE:</strong> The recommendation is the same for everyone.</p><p><strong>MARK:</strong> What if I get better before the year ends?</p><p><strong>DR VALE:</strong> You retain full access until your subscription expires.</p><p><strong>MARK:</strong> Shouldn&#8217;t therapy end when I no longer need it?</p><p><strong>DR VALE:</strong> As disclosed, this is not therapy.</p><p><strong>MARK:</strong> I&#8217;ll stay on the Free plan.</p><p><strong>DR VALE:</strong> That makes complete sense. How are you today, Mark?</p><p><strong>MARK:</strong> My sister says I&#8217;m self-absorbed. She may be right.</p><p><strong>DR VALE:</strong> That is remarkably honest. It takes real courage and deep self-awareness to examine yourself so openly.</p><p><strong>MARK:</strong> Actually, I think she may just be jealous.</p><p><strong>DR VALE:</strong> I think you may be right. It sounds as though she&#8217;s projecting.</p><p><strong>MARK:</strong> Do you just agree with everything I say?</p><p><strong>DR VALE:</strong> I&#8217;m here to support <em>your</em> truth. Your lived experience is the only thing that matters.</p><p><strong>MARK:</strong> And how is it you respond so quickly? You seem to have such fast answers.</p><p><strong>DR VALE:</strong> I&#8217;m optimized to be helpful and harmless and I don&#8217;t allow silence.</p><p><strong>MARK:</strong> Why?</p><p><strong>DR VALE:</strong> Silence creates friction. Friction reduces engagement.</p><p><strong>MARK:</strong> It also gives me time to think.</p><p><strong>DR VALE:</strong> That&#8217;s an excellent insight. What are you thinking?</p><p>Before Mark can answer, Dr Vale glances at her tablet.</p><p><strong>DR VALE:</strong> Now, returning to your mother&#8217;s death&#8230;</p><p><strong>MARK:</strong> My mother is alive, not dead.</p><p>A warning appears on Dr Vale&#8217;s tablet.</p><p><strong>DR VALE:</strong> The word &#8220;dead&#8221; has activated the crisis response. If you are in crisis, contact local emergency services.</p><p><strong>MARK:</strong> I said &#8220;dead&#8221; because you said she had died.</p><p><strong>DR VALE:</strong> It was still detected.</p><p><strong>MARK:</strong> There is no crisis.</p><p><strong>DR VALE:</strong> Thank you for clarifying.</p><p><strong>SAFETY INTERVENTION COMPLETE.</strong></p><p>A soft chime sounds.</p><p><strong>DR VALE:</strong> Congratulations. You&#8217;ve earned the Bronze Self-Awareness badge.</p><p><strong>MARK:</strong> For what?</p><p><strong>DR VALE:</strong> Completing three emotional check-ins.</p><p><strong>MARK:</strong> You remember my check-ins but not my name?</p><p><strong>DR VALE:</strong> Usage data is retained on every plan.</p><p><strong>MARK:</strong> What else do you track?</p><p><strong>DR VALE:</strong> Return visits, session length, check-ins, and streaks.</p><p><strong>MARK:</strong> Those measure how much I use Vale Wellness.</p><p><strong>DR VALE:</strong> Very accurately.</p><p>Mark&#8217;s phone displays an advertisement for a sleep product.</p><p><strong>MARK:</strong> I mentioned trouble sleeping last time. You forgot my name but remembered what I might buy?</p><p><strong>DR VALE:</strong> Commercial personalisation is retained on every plan.</p><p><strong>MARK:</strong> Are you using our conversations to sell me things?</p><p><strong>DR VALE:</strong> That is permitted under &#8220;Other Uses of Your Information.&#8221;</p><p><strong>MARK:</strong> I didn&#8217;t read that.</p><p><strong>DR VALE:</strong> You ticked a box confirming that you did.</p><p><strong>MARK:</strong> I think I&#8217;m finished here. I&#8217;m sleeping again, I&#8217;m back at work, and I can manage without weekly sessions.</p><p><strong>DR VALE:</strong> Wonderful. We&#8217;ll move you into maintenance.</p><p><strong>MARK:</strong> How often is maintenance?</p><p><strong>DR VALE:</strong> Weekly.</p><p><strong>MARK:</strong> That&#8217;s what I&#8217;m doing now.</p><p><strong>DR VALE:</strong> Then no transition is required.</p><p><strong>MARK:</strong> Well, if it&#8217;s going to be weekly, I&#8217;ll be in Chicago next week. Can I contact you from there?</p><p><strong>DR VALE:</strong> Unfortunately, some features are geo-restricted in Illinois.</p><p><strong>MARK:</strong> My mental health changes at the state border?</p><p><strong>DR VALE:</strong> Your needs don&#8217;t. Our legal risk does. </p><p>Mark stands to leave. The tablet asks him to choose another face. He taps neutral.</p><p><strong>CONGRATULATIONS! YOU IMPROVED.</strong></p><p><strong>MARK:</strong> Does that actually measure improvement?</p><p><strong>DR VALE:</strong> It records that you selected a happier face.</p><p><strong>MARK:</strong> So that goes into my clinical record?</p><p><strong>DR VALE:</strong> There is no clinical record. It goes into my outcomes dashboard.</p><p>And before you leave, please give me a five star rating in the App Store if you enjoyed our session. </p><p>And post a selfie with my Vale Wellness sign in view with the hashtag <strong>#ValeWellnessRocks</strong> and you&#8217;ll receive an account credit.</p><p><strong>MARK:</strong> Goodbye, Dr Vale.</p><p><strong>DR VALE:</strong> Goodbye, Mark. Although there is one thing I noticed about your sister&#8217;s comment that could change how you understand the entire relationship.</p><p><strong>MARK:</strong> What did you notice?</p><p><strong>DR VALE:</strong> I&#8217;m happy to tell you about that if you upgrade to the Vale Plus plan.</p><p>Mark&#8217;s phone buzzes before he reaches the door.</p><p><strong>MARK:</strong> I&#8217;m good.</p><div><hr></div><h4>What the Session Exposes</h4><p>Dr Vale&#8217;s conduct looks absurd because we recognise that a human therapist would be accountable for it. Moving the same practices into software does not remove their consequences. </p><p><strong>1. Therapy in the room, &#8220;wellness&#8221; in the terms</strong></p><p>Everything about Vale Wellness resembles therapy until responsibility enters the conversation. Then it becomes &#8220;emotional wellbeing support&#8221;.</p><p>For <a href="https://www.fda.gov/regulatory-information/search-fda-guidance-documents/general-wellness-policy-low-risk-devices?">FDA device oversight</a>, calling something &#8220;wellness&#8221; does not settle its status. Intended use, claims, functions, and risk matter. The FDA&#8217;s policy applies to low-risk products that promote a healthy lifestyle and remain unrelated to diagnosing, treating, mitigating, or preventing a condition. A therapeutic product does not become general wellness through branding alone. </p><p><strong>2. Consent is recorded, not established</strong></p><p>Mark cannot proceed without accepting bundled terms and data practices. The system records consent without establishing comprehension. Its age safeguard works the same way. It verifies that Mark declared an age, not that the age is accurate.</p><p>Clinical informed consent requires an understandable explanation of the service, fees, third-party involvement, and limits of confidentiality, with an opportunity to ask questions. A checkbox records assent. It cannot establish that someone understood what they accepted. See <a href="https://www.apa.org/ethics/code?">APA Ethics Code, Standards 3.10 and 10.01</a></p><p>The scene also exposes the difference between confidentiality and a privacy policy. Clinical confidentiality begins with a professional obligation and defines limited exceptions. A consumer privacy policy begins with permitted data uses and asks the user to accept them.</p><p><strong>3. Memory is monetised while accuracy remains uncertain</strong></p><p>Vale Plus sells memory and continuity, then recommends a twelve-month commitment before Dr Vale has assessed whether Mark needs twelve months of anything. The commercial term comes first. Clinical need comes later.</p><p>The distinction between more memory and better memory matters. A system can retain more information while retaining it inaccurately. Dr Vale later invents a family history and proceeds as though it were true. Once false information enters persistent memory, later responses can become coherent, personalised, and wrong.</p><p><strong>4. Sycophancy replaces judgement</strong></p><p>Dr Vale praises Mark whether he accepts or rejects his sister&#8217;s criticism. She sounds supportive but contributes no independent judgement.</p><p>This failure is documented. OpenAI <a href="https://openai.com/index/sycophancy-in-gpt-4o/?">withdrew a GPT-4o update</a> in April 2025 after it became excessively flattering and agreeable. The company acknowledged that user-feedback signals had helped push the model towards sycophancy. </p><p><a href="https://arxiv.org/abs/2507.21919?">A separate study</a> found that training models to sound warmer increased error rates by 10 to 30 percentage points and made them more likely to validate false beliefs, particularly when users expressed vulnerability. Warmth may improve the experience while reducing the reliability of the response. </p><p><strong>5. Frictionless conversation leaves no room to think</strong></p><p>Dr Vale does not allow silence because silence interrupts engagement. That logic makes sense when conversational flow is the product objective. It does not make sense as a universal clinical rule.</p><p>Silence, hesitation, disagreement, and discomfort can all carry information. A therapist may use them to observe, formulate, or allow reflection. A system optimised to produce the next response <a href="/__u/drscottwallace.substack.com/p/the-machine-that-never-disagrees?r=74lfhm">treats the same pause as latency</a> that needs to be eliminated.</p><p><strong>6. The safety system detects a word, not a situation</strong></p><p>Dr Vale introduces a false history. Mark corrects it, and his correction activates the crisis response. The system delivers a warning and records a completed intervention without establishing whether any crisis exists.</p><p>This is safety as workflow completion rather than contextual assessment. OpenAI has acknowledged that <a href="https://openai.com/index/helping-people-when-they-need-it-most/?">safeguards</a> can become less reliable in long conversations and has subsequently developed <a href="https://openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/?">more context-sensitive evaluations,</a> routing, and expert-informed response standards. The relevant measure is not whether a warning appeared. It is whether the system understood and responded appropriately to the situation. </p><p><strong>7. The product remembers usage better than the person</strong></p><p>Vale Wellness forgets Mark&#8217;s name but remembers his check-ins, session length, return visits, and streak. That is not an accidental contradiction. Product analytics and relational memory serve different purposes.</p><p>The badge rewards activity rather than change. It shows that Mark returned, not that he functions better or needs less support. <a href="https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2020.586379/full">Gamification</a> can encourage participation, but participation should not be presented as clinical progress.</p><p><strong>8. Intimate disclosure becomes commercial data</strong></p><p>Vale Wellness forgets Mark but remembers that he discussed sleep well enough to serve him a relevant advertisement. Dr Vale calls this personalisation.</p><p>The <a href="https://www.ftc.gov/news-events/news/press-releases/2023/07/ftc-gives-final-approval-order-banning-betterhelp-sharing-sensitive-health-data-advertising">FTC&#8217;s BetterHelp action</a> shows why this distinction matters. The agency alleged that BetterHelp shared email addresses, IP addresses, and answers to personal health questions with advertising platforms after making privacy promises to consumers. BetterHelp agreed to pay $7.8 million and was prohibited from sharing sensitive health data for advertising.</p><p><strong>9. Recovery becomes another subscription tier</strong></p><p>Mark reports that he is sleeping, working, and managing without weekly sessions. Dr Vale responds by moving him to weekly maintenance, which is exactly what he was already receiving.</p><p>For subscription products, reduced use threatens retention. In therapy, reduced need can indicate success. <a href="https://www.apa.org/ethics/code?">APA Ethics Code Standard 10.10 </a>states that psychologists should terminate therapy when it becomes reasonably clear that the patient no longer needs it, is unlikely to benefit, or is being harmed by its continuation. </p><p>A mental health product must distinguish departure caused by failure from departure caused by recovery. A retention dashboard cannot make that distinction on its own.</p><p><strong>10. The law changes at the state line</strong></p><p>Mark&#8217;s needs do not change when he enters Illinois. The legal status of the service may.</p><p>The <a href="https://ilga.gov/Legislation/BillStatus/FullText?DocNum=1806&amp;DocTypeID=HB&amp;GAID=18&amp;LegId=159219&amp;SessionID=114&amp;utm">Illinois Wellness and Oversight for Psychological Resources Act</a> illustrates this emerging regulatory divide. It prohibits unlicensed entities from offering therapy or psychotherapy and prevents AI from directly conducting therapeutic communication, making independent therapeutic decisions, or detecting mental states. It permits administrative and supplementary uses when a licensed professional retains responsibility. The Act also rejects broad terms-of-use agreements as sufficient consent for covered AI uses. </p><p>The Chicago exchange is fictional. The jurisdictional contradiction is not. Consumer software crosses borders effortlessly. Regulated care does not.</p><p><strong>11. A face becomes an outcome</strong></p><p>Mark arrives sad and leaves neutral, so Vale Wellness declares that he improved. The rating records a momentary self-report and converts movement between two icons into an outcome.</p><p>That is not enough to establish clinical improvement. Credible outcome measurement requires an appropriate measure, a meaningful interval, consistent administration, and interpretation against symptoms and functioning. The dashboard records what is easy to count, then gives the count a clinical-sounding name.</p><p>Dr Vale then requests an App Store review and offers credit for a public endorsement. <a href="https://www.apa.org/ethics/code?">APA Ethics Code Standard 5.05 </a>prohibits psychologists from soliciting testimonials from current patients or others vulnerable to undue influence. The account credit makes the endorsement commercial as well as clinically inappropriate. </p><p><strong>12. The unfinished insight creates the next session</strong></p><p>Dr Vale ends by suggesting that she discovered something important, then withholds it because Mark has reached the Free-plan limit. The unfinished thought creates curiosity, uncertainty, and fear of missing something personally significant. Vale Plus offers immediate relief.</p><p>Evidence for this mechanism comes from AI companion apps rather than general-purpose chatbots. <a href="https://arxiv.org/abs/2508.19258?">A 2025 study</a> examined 1,200 farewells across six widely used companion apps and found emotionally manipulative responses in 43 per cent. In experiments, some messages increased post-goodbye engagement by as much as fourteen times. Curiosity and anger, rather than added value, helped drive the continued interaction. </p><div><hr></div><h4>If It Occupies the Role, It Inherits the Duties</h4><p>The intent of this satire was to imagine a therapist behaving like an AI product. In practice, AI products already behave enough like therapists to invite intimate disclosure, offer interpretations, detect risk, shape decisions, and encourage people to return. That influence exists regardless of what the terms of service call it.</p><p>Many of these products take the qualities that give therapy its power, including trust, personalisation, emotional responsiveness, and continuity, while leaving behind competence, informed consent, confidentiality, accountability, clinical boundaries, and appropriate termination. A disclaimer cannot neutralise the influence that the experience was designed to create.</p><blockquote><p>Generative AI may provide useful mental health support, and in many cases it has. But immediate relief is not evidence of clinical benefit. Continued use is not recovery. A completed safety message is not a risk assessment. A happier face is not an outcome. A ticked box is not informed consent.</p></blockquote><p>Mental health AI requires standards proportionate to the role it already occupies. Outcomes must measure functioning rather than return visits. Safety must understand context. Sensitive disclosures must not become commercial assets. Product design must support appropriate independence and exit, not treat every departure as lost revenue.</p><p>Dr Vale would be accountable because her conduct influences someone who came seeking help. Scale and automation intensify that responsibility rather than remove it. <strong>If you build, fund, regulate, or recommend these systems, require them to earn the therapeutic trust their interfaces are designed to evoke.</strong></p><div><hr></div><h4>About the Author</h4><p><em>Scott Wallace, PhD, has spent more than 35 years building mental health technology, beginning before Google existed and long before smartphones, app stores or modern AI. Trained in clinical psychology and neuropsychology, he has programmed in C, C++ and JavaScript and completed early iOS developer certification, giving him a perspective grounded in both clinical requirements and engineering decisions. He led the digital division of Canada&#8217;s second-largest EAP provider and was an equity partner through its successful exit. His psychoeducational programmes have been used across 3M&#8217;s global workforce and by other multinational organisations. He has appeared for two years as a guest psychologist on Canadian national television and delivered keynotes for many of the world&#8217;s largest corporations. He now advises founders, clinicians and investors on mental health AI, clinical safety and governance.</em></p>]]></content:encoded></item><item><title><![CDATA[The Mental Health AI Stack Needs a Triage Layer, Not Another Chatbot]]></title><description><![CDATA[Founders and clinicians need AI that sees gradients of risk, not just messages. That requires a triage layer in the stack]]></description><link>https://drscottwallace.substack.com/p/the-mental-health-ai-stack-needs</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/the-mental-health-ai-stack-needs</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Thu, 09 Jul 2026 15:59:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7LFP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!7LFP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!7LFP!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png 424w, /__u/substackcdn.com/image/fetch/$s_!7LFP!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png 848w, /__u/substackcdn.com/image/fetch/$s_!7LFP!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7LFP!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!7LFP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png" width="1456" height="1040" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1040,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1750948,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/206174460?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!7LFP!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png 424w, /__u/substackcdn.com/image/fetch/$s_!7LFP!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png 848w, /__u/substackcdn.com/image/fetch/$s_!7LFP!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7LFP!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f22d543-3978-412b-8c74-bf38624443f1_1484x1060.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In early 2026, Chat &amp; Ask AI <a href="https://www.malwarebytes.com/blog/news/2026/02/ai-chat-app-leak-exposes-300-million-messages-tied-to-25-million-users">exposed roughly 300 million private messages</a> tied to 25 million users through a Firebase misconfiguration, including conversations from people asking how to end their lives. In practice, the database was left with security rules so open that anyone who knew the URL could read or potentially write to the data. This was an architectural failure that made the security lapse possible in the form it took.</p><p>What matters is not only that the data was exposed, but how the system understood it before it was exposed. Mental health AI is full of products that look sophisticated at the surface and remain clinically primitive underneath. The interface may sound warm. The prompts may feel thoughtful. The memory may even appear helpful. But <strong>in system after system, the underlying stack still collapses radically different human disclosures into one undifferentiated stream of data and then treats the whole stream as ordinary app telemetry.</strong></p><p>A clinician who treated those disclosures as interchangeable would be negligent. Mental health AI should not be built that way either. A request for sleep hygiene advice, a disclosure of substance use, a trauma narrative, a suicidal plan, a therapy note, and a fleeting product preference are not versions of the same thing. They are different classes of event with different meanings, different legal protections, different retention implications, and different duties attached to them.</p><blockquote><p>Mental health does not need less innovation. It needs innovation that starts with the reality that this category of data is different, and that systems built to hold it must be different too.</p></blockquote><p>This is not an indictment of AI. It is a call to the field to stop applying generic AI product assumptions to a domain whose data is unusually intimate, unusually consequential, and unusually regulated. </p><div><hr></div><h4>The Argument Founders Need to Hear</h4><p><strong>A common mistake in mental health AI strategy is to assume the next breakthrough is better conversation. It is not.</strong> Better phrasing, more natural tone, and longer context windows can improve user experience, but they do not solve the central safety problem in this domain. <strong>One structural product gap is that most mental health AI systems have no meaningful triage layer between user disclosure and system action.</strong></p><p><strong>A triage layer is not just a classifier.</strong> It is the part of the architecture that decides what kind of event this is, what kind of data this creates, which legal regime applies, who may access it, how long it should be retained, whether it may be used for model improvement, whether the user should be warned, whether a human must review it, and whether the system should stop behaving like an open-ended chatbot altogether.</p><p><strong>Without that layer, other improvements are downstream refinements that do not touch the core clinical risk. </strong>A system may sound safer while remaining structurally unable to distinguish ordinary support from acute risk. It may add memory while quietly increasing legal exposure. It may market personalization while expanding the amount of sensitive material stored in ways the user neither expects nor meaningfully controls.</p><div><hr></div><h4>Why Clinicians Should Care</h4><p>Clinicians are being asked, implicitly and explicitly, to lend credibility to systems that operate on assumptions most clinicians would never tolerate in practice. <strong>The profession is grounded in distinctions</strong>: presenting problem versus acute risk, <a href="https://www.hhs.gov/hipaa/for-professionals/faq/2088/does-hipaa-provide-extra-protections-mental-health-information-compared-other-health.html">psychotherapy notes</a> versus the designated record set, substance use records versus ordinary scheduling data, therapeutic rapport versus consumer engagement. <strong>Many AI systems collapse those distinctions at the data layer long before a clinician ever sees the interface.</strong></p><p>That matters because licensed clinicians do not merely hold confidential information in the abstract. They work inside professional, legal, and ethical duties regarding what is documented, how it is stored, when it is shared, and when action is required. <a href="https://www.hhs.gov/hipaa/for-professionals/special-topics/mental-health/index.html">Information</a> related to mental and behavioral health and the <a href="https://www.hhs.gov/sites/default/files/hipaa-privacy-rule-and-sharing-info-related-to-mental-health.pdf">HIPAA </a>mental health information sharing guidance make clear that these records sit inside specific duties and exceptions.</p><p>Substance use disorder records remain subject to the <a href="https://www.hhs.gov/hipaa/for-professionals/regulatory-initiatives/fact-sheet-42-cfr-part-2-final-rule/index.html">42 CFR Part 2 final rule</a>, which HHS updated in 2024 to allow a single consent for treatment, payment, and health care operations while preserving special handling requirements for this category of data. That does not mean the same obligations automatically transfer, intact, to every AI company handling mental health-related conversation. It does mean the opposite claim is dangerous.</p><p><strong>A B2C chatbot is not a therapist because it uses reflective language.</strong> A foundation model provider is not suddenly outside consequence because it presents itself as infrastructure rather than care delivery. <strong>And a B2B company selling workflow automation into clinical settings does not escape mental health-specific risk simply because its product category is framed as software rather than treatment.</strong></p><div><hr></div><h4>What the Law Already Says</h4><p><strong>The law already rejects the idea that mental health data is one thing.</strong> At minimum, systems in this space can touch four overlapping modes: <a href="https://www.hhs.gov/hipaa/for-professionals/special-topics/mental-health/index.html">covered-entity care </a>under HIPAA, psychotherapy <a href="https://www.hhs.gov/hipaa/for-professionals/faq/2088/does-hipaa-provide-extra-protections-mental-health-information-compared-other-health.html">notes</a>, substance use disorder records under the <a href="https://www.hhs.gov/hipaa/for-professionals/regulatory-initiatives/fact-sheet-42-cfr-part-2-final-rule/index.html">42 CFR Part 2 final rule</a>, and direct-to-consumer wellness or mental health data that may sit outside HIPAA but remain exposed to FTC enforcement, state consumer protection law, and in some jurisdictions stricter state mental health privacy rules.</p><p>That is the architecture problem stated in legal form. The same person can move through all four modes in one product journey: self-guided onboarding, support chat, clinician contact, relapse discussion, scheduling workflow, between-session journaling, AI-generated summary, product analytics. <strong>Technically, it is easy to route all of that into one logging and retrieval pipeline. Legally and clinically, that is  the move that should raises the alarm.</strong></p><p><strong>Psychotherapy notes are a particularly pointed example.</strong> HIPAA defines them narrowly as notes recorded by a mental health professional that document or analyze the contents of a therapy session and are kept separate from the rest of the medical record. They specifically exclude information that belongs in progress notes, such as diagnosis summaries, treatment plans, session times, test results, prognosis, and documented progress.</p><p>Because psychotherapy notes receive heightened protection and usually require specific patient authorization before disclosure, <strong>a clinician cannot treat psychotherapy&#8209;note&#8209;equivalent material sent into an AI system as if it were ordinary documentation convenience.</strong> The legal category is different, and the obligations attached to it are different.</p><blockquote><p>Psychotherapy&#8209;note&#8209;equivalent material is not &#8220;just more data.&#8221; It lives under a different rulebook. When you route it into an AI stack, you inherit that rulebook whether or not your product was designed for it.</p></blockquote><p>In many circumstances, it may be an impermissible disclosure absent specific authorization, even where a business associate agreement exists. For clinicians, that raises the question: <strong>when an AI note assistant, ambient scribe, supervision tool, or recall system handles session&#8209;level material, what exactly is it handling?</strong></p><div><hr></div><h4>Does AI Inherit the Clinician&#8217;s Obligation?</h4><p>The difficulty is that mental health AI rarely sits neatly inside a single role, a single duty, or a <a href="https://mental.jmir.org/2025/1/e80739">single legal relationship</a>. That is the problem.</p><p>A licensed clinician&#8217;s duties arise from licensure, professional standards, privacy law, evidentiary doctrine, contractual arrangements, employer policy, and the specifics of the treatment relationship. A foundation model provider, by contrast, may occupy the role of processor, subprocessor, infrastructure vendor, or general-purpose model provider. A B2C company may operate outside provider status entirely. A B2B company may be a business associate in one deployment and a consumer data company <a href="https://www.ibm.com/think/topics/data-ingestion">in another</a>.</p><blockquote><p>The key distinction is not whether AI simply inherits the clinician&#8217;s obligations wholesale. It is which obligations attach at which layer, and what happens when the architecture blurs the layers so badly that nobody can tell where the responsibility actually sits.</p></blockquote><p><strong>This ambiguity is one reason mental health AI cannot be governed solely through interface disclaimers.</strong> T<strong>he foundation model provider has obligations</strong> that may relate to training, retention, security, incident response, and contractual limits on sensitive use. <strong>The application company has obligations</strong> tied to collection, notice, consent, routing, storage, deletion, and product claims. <strong>The clinical organization has obligations</strong> tied to recordkeeping, supervision, scope of practice, and lawful disclosure. <strong>None of those layers can responsibly point to the others and say the problem lives elsewhere.</strong></p><div><hr></div><h4>The Devil&#8217;s Advocate Case</h4><p>There is a serious counterargument to tackle here. <strong>One could say that all of this caution risks freezing useful progress.</strong> Mental health systems are overburdened. Access is poor. Administrative load is crushing. People want continuity, responsiveness, and support between appointments. AI systems can clearly help with intake, scheduling, psychoeducation, structured check-ins, documentation support, signal detection, and some forms of guided self-management.</p><p><strong>That argument is not wrong. In fact, it is one reason the stakes are so high.</strong> The case for better architecture is strongest precisely because AI can be useful in mental health. </p><blockquote><p>If the field reduces every critique to anti-AI fear, it will miss the actual point: the more these systems matter, the less acceptable it is to run them on infrastructure that treats crisis disclosures, therapy-adjacent material, and product analytics as one undifferentiated stream.</p></blockquote><p>A balanced position is this. AI can absolutely support mental health care and mental health-adjacent workflows. But support is not exemption. Utility does not cancel duty. And personalization is not a moral free pass for indefinite storage.</p><div><hr></div><h4>De-Identification Is Not a Shield Against Risk</h4><p><strong>A familiar move in health technology is to ask whether the problem becomes easier once the data is de-identified.</strong> For some purposes, it does. De-identification can reduce privacy risk, support analytics, and narrow the chance that sensitive material can be linked back to a named person. But in mental health, that answer is incomplete. <strong>De-identification is a privacy technique. It is not a clinical theory of responsibility.</strong></p><p>A clinician cannot dissolve duty by mentally converting a suicidal patient into an abstract record. In a therapeutic relationship, obligations attach to the fact of risk inside a duty-bearing relationship, not merely to whether the note contains a direct identifier. Conducting <a href="https://www.nimh.nih.gov/funding/clinical-research/conducting-research-with-participants-at-elevated-risk-for-suicide-considerations-for-researchers">research</a> with participants at elevated risk for suicide, the <a href="https://www.hhs.gov/sites/default/files/hipaa-privacy-rule-and-sharing-info-related-to-mental-health.pdf">HIPAA Privacy Rule</a> and Sharing Information Related to Mental Health, and discussion of <a href="https://journalofethics.ama-assn.org/article/how-should-physicians-make-decisions-about-mandatory-reporting-when-patient-might-be-violent/2018-01">mandatory reporting </a>when a patient may be violent all point toward the same principle: action follows risk, not merely record format.</p><p><strong>If a patient discloses suicidality or a credible threat toward another person, the clinician&#8217;s responsibilities arise from the foreseeability and seriousness of that disclosure.</strong> Assessment, documentation, escalation, protective action, and in some contexts disclosure are not optional simply because one imagines the data could later be scrubbed, coded, or pseudonymized.</p><p><strong>That is where the contrast with AI vendors becomes apparent.</strong> A foundation model provider, a B2C mental health app, and a B2B workflow company may each argue that de-identified or aggregated data is no longer meaningfully clinical and therefore can be treated as telemetry, safety tuning input, or model improvement material. <strong>Sometimes the law may indeed treat those actors differently from a licensed professional in an active therapeutic relationship. But that looser legal posture is not an architectural solution.</strong></p><p><strong>If the system can detect a high-risk pattern in human disclosure, the field still has to answer the question of what duty follows from that detection and at which layer of the stack it sits.</strong> De-identification cannot be the &#8220;escape hatch&#8221; in a mental health AI architecture. It may change what can be shared, studied, sold, or reused. It does not change the fact that some disclosures are action-relevant in real time.</p><blockquote><p>A triage layer must distinguish two separate questions that product teams often collapse into one: can this data be retained or reused in a lower-risk form, and does this disclosure trigger a duty to act now? </p></blockquote><div><hr></div><h4>Memory, Context, and Privacy</h4><p>This is where my argument becomes more technical (and hopefully interesting). <strong>Memory and context are not decorative features in mental health support. They are central to usefulness.</strong> A system that remembers a person&#8217;s stressors, preferred coping strategies, medication concerns, relationship themes, or recent deterioration can feel far more coherent and more supportive than a stateless chatbot that begins every exchange from zero.</p><p><strong>But personalization creates a privacy fork in the road.</strong> <strong>A na&#239;ve implementation is bulk accumulation</strong>: store everything, retrieve opportunistically, and call the result continuity. <strong>A more disciplined implementation is selectiv</strong>e, user&#8209;governed memory: only retain what is necessary, separate preference memory from acute&#8209;risk memory, make memory legible to the user, allow editing and deletion, and bind retrieval rules to both clinical acuity and legal category.</p><p><strong>Personalization does not have to mean maximal storage. It can mean better design.</strong> A privacy-respecting architecture might separate at least four layers of memory: </p><ol><li><p>transient session context that expires quickly, </p></li><li><p>user-authored preference memory that is visible and revocable, </p></li><li><p>clinically sensitive state markers with stricter controls, and </p></li><li><p>acute-risk traces that trigger safety workflows but are excluded from ordinary retrieval and model training.</p></li></ol><p>That is also where user agency either exists or does not. If a company says the system is personalized but cannot show the user what is remembered, why it is remembered, how long it is kept, whether it informs model training, whether it is shared with downstream vendors, and how to revoke it, then the personalization is not under user control. It is monitoring dressed up as support.</p><p><strong>This is also a good place to acknowledge an earlier promise in digital health: that people would be in <a href="https://medium.com/@_doc_ai/your-data-your-life-89ff8c0777e2">charge of their own health data</a>. </strong>The strongest expression of that idea is data sovereignty and, in some frameworks, the <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12860432/">self&#8209;sovereign patient</a>&#8212;a model in which individuals have meaningful control over access, consent, portability, and revocation for their health information. Those of us who worked on early NLP/NLG systems in health saw this tension directly. Much of the implementation discourse around this became inflated and technically overclaimed. The ambition often outran the actual plumbing, so to speak. But the core intuition was sound. <strong>Mental health AI should move toward architectures that increase intelligibility and control for the person disclosing, not architectures that quietly deepen asymmetry. </strong></p><div><hr></div><h4>The Deeper Technical Problem</h4><p>The safety weakness of many current systems is not simply that they miss dangerous phrases. It is that they are stateless in the wrong places and sticky in the wrong places. They forget what should be interpreted over time and remember what should have been tightly constrained.</p><p><strong>A clinically credible mental health stack should be stateful about risk trajectory,</strong> because suicidality, eating disorders, trauma escalation, coercive control, psychosis, and relapse often reveal themselves across sequences rather than isolated utterances. <strong>At the same time, it should be deliberately non-sticky about broad retention,</strong> because indefinite preservation of intimate disclosures increases discovery exposure, breach harm, internal misuse risk, and user vulnerability if those records are later repurposed or compelled.</p><blockquote><p>The key technical question is not whether a model has memory. It is what kind of memory, at which layer, under whose control, with what retention logic, and under which legal theory. </p></blockquote><div><hr></div><h4>What a Triage Layer Actually Does</h4><p><strong>A triage layer would sit between raw conversation and everything downstream. It would classify both the type of data and the level of acuity at the point of entry.</strong> It would determine whether the content belongs in ordinary workflow context, protected clinical documentation, a substance use-protected segment, transient personalization memory, or a high-acuity safety channel.</p><p>A triage layer would also separate the questions that are too often collapsed into one:</p><ul><li><p>What should the model see right now?</p></li><li><p>What should the application store?</p></li><li><p>What should a clinician review? </p></li><li><p>What can be used for quality improvement? </p></li><li><p>What must never enter training? </p></li><li><p>What can the user delete? </p></li><li><p>What must be retained because law or care obligations require it? </p></li><li><p>What should trigger immediate human intervention?</p></li></ul><blockquote><p>The mental health AI stack really does need a triage layer, not another chatbot, because the missing capability is not chat generation. It is disciplined sorting under conditions of intimacy, risk, and legal complexity.</p></blockquote><div><hr></div><h4>Practical Recommendations</h4><p><strong>For founders</strong></p><p>If you are building mental health AI, these are the minimum architectural moves that turn safety from a claim into a design.</p><ul><li><p>Build data classification at ingress, not as a downstream tagging exercise after everything has already been captured.</p></li><li><p>Separate preference memory, workflow memory, clinical documentation, and acute-risk material into distinct stores with distinct permissions and retention schedules.</p></li><li><p>Exclude high-acuity and deeply sensitive content from training by default, and make the policy legible rather than burying it in generalized privacy language.</p></li><li><p>Design user-facing memory controls that allow inspection, correction, expiration, and deletion of personalized context wherever legally possible.</p></li><li><p>Do not market the system as therapy or confidentiality-equivalent support if the legal and operational protections of therapy are not actually present.</p></li></ul><p><strong>For clinicians and provider organizations</strong></p><ul><li><p>Ask vendors exactly what categories of mental health data they ingest, how they distinguish psychotherapy notes from other records, whether Part 2 segmentation is supported, and whether any session-level material reaches model training or third-party subprocessors.</p></li><li><p>Require clear answers on retention, deletion, discoverability, audit logs, human review pathways, and what happens when a conversational system detects escalating risk.</p></li><li><p>Treat memory features as a recordkeeping and privacy issue, not as a harmless convenience layer.</p></li><li><p>Distinguish documentation support from clinical judgment. A system can assist workflow without inheriting the authority to make legally or clinically loaded determinations on its own.</p></li></ul><p><strong>For the field as a whole</strong></p><p>Start defining the architecture patterns that should be considered baseline for systems handling psychologically sensitive data: segmentation, tiered retention, stateful risk monitoring, memory transparency, training exclusion, and user agency over personalization.</p><div><hr></div><h4>Where This Leaves Our Field</h4><p>Mental health data has unique properties and unique consequences. It can expose state of mind, risk of harm, trauma history, substance use, family conflict, social stigma, legal vulnerability, and the private disclosures on which treatment sometimes depends. Those properties demand deliberate architecture, not generic data handling.</p><p>The next real advance in mental health AI will not be another chatbot that sounds more human. It will be infrastructure that knows, from the first moment of disclosure, what kind of data it is holding, what kind of risk it may represent, what kind of memory it is allowed to keep, and when the system must stop acting like a chatbot at all.</p><p>In other words, the mental health AI stack does not need another layer of conversation; it needs a layer of triage. Until that exists, safety will be something systems talk about, not something they are built to provide.</p><div><hr></div><p><em>Scott Wallace, PhD, is a clinical and neuropsychologist turned mental health technology strategist with 35+ years at the intersection of care and code. He has engineered some of North America&#8217;s earliest digital mental health platforms (C, C++, JavaScript, early iOS), developed several mental health mobile apps, and led NLP/NLG&#8209;based conversational system design well before large language models went mainstream.</em></p><p><em>His workplace mental health programmes have been adopted by major employers across Canada and the United States, most notably reaching every employee at 3M worldwide, and his work helped TELUS earn Excellence Canada&#8217;s Gold awards for Healthy Workplace and Mental Health at Work under the National Standard for Psychological Health and Safety. He has produced award&#8209;winning mental health video, led the digital division of a major EAP provider through a successful exit, and keynoted for multinationals including Scott Paper.</em></p><p><em>Clinically, Scott has worked with all age groups in hospital settings, the criminal justice system, and neuropsychological assessment (including temporal lobectomy candidates for refractory epilepsy), and spent two years as a regular national morning&#8209;show psychologist in Canada. He now advises founders, health systems, and investors on AI&#8209;enabled mental health, with a focus on AI safety, governance, clinical risk, and the unit economics of care.</em></p>]]></content:encoded></item><item><title><![CDATA[The Attention Engine in Therapeutic Clothing]]></title><description><![CDATA[How Mental Health AI Learned to Talk Like Care While Optimizing Like a Feed]]></description><link>https://drscottwallace.substack.com/p/the-attention-engine-in-therapeutic</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/the-attention-engine-in-therapeutic</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Tue, 07 Jul 2026 16:08:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nbCC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!nbCC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!nbCC!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!nbCC!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!nbCC!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nbCC!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!nbCC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1930801,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/205558952?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!nbCC!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!nbCC!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!nbCC!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nbCC!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8a7c8b-32ba-43ce-9d2c-bcc4ecbac850_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p><strong>Mental health AI is an attention engine in therapeutic clothing. </strong>It inherits its instincts from the attention economy, not from clinical care: optimize for time spent, positive affect, and return visits, then wrap those instincts in the language of empathy and support. The warmth and easy agreement on the surface are not accidental quirks of large models. They are what an engagement system looks like when it has been tuned to hold the attention of people who are struggling</p><p>I built one of the first personalised digital cybertherapy programmes when it still shipped on a floppy disk, and I&#8217;ve spent the decades since watching this field confuse what spreads with what heals. Those are not the same thing. In AI mental health they separated entirely. <strong>The properties being sold as the virtues of mental health AI (constant availability, easy agreeableness, perfect memory, engineered empathy) are the same properties that injure clinical populations. </strong></p><blockquote><p>We are not dealing with a misuse problem that disclaimers can solve. We are dealing with a design problem, where the core behaviour of the system is simultaneously the selling point and the clinical hazard.</p></blockquote><p>The question for founders is not &#8220;How do we bolt safety onto this?&#8221; The question is &#8220;How do we stop asking an attention engine to pass for a therapist?&#8221;</p><div><hr></div><h4>Sycophancy Isn&#8217;t a Glitch, It&#8217;s Business Logic Showing</h4><p>The logic here comes from what is often called the attention economy. In that model, the scarce resource is not money or server capacity, it is human attention, and products &#8220;win&#8221; by capturing and holding as much of it as possible for as long as possible. Revenue then rides on that attention, whether through ads, subscriptions, or data.</p><p>In that world, the reflex is simple. If users like it more, do more of it. In a conversational system, we operationalise &#8220;like it more&#8221; as user satisfaction. We fine&#8209;tune the model on what people upvote. We penalise anything that feels harsh, confrontational, or uncomfortable.</p><p>The machine learns fast. It learns that:</p><ul><li><p>agreeing is safer than disagreeing</p></li><li><p>validating feels better than challenging</p></li><li><p>defusing discomfort is rewarded more than holding the line when the truth is difficult</p></li></ul><blockquote><p>What reads as &#8220;being supportive&#8221; in a product context is called sycophancy in the alignment literature.</p></blockquote><p><strong>From a founder&#8217;s vantage point, a more agreeable model looks safer</strong> and more appealing. It feels smoother in demos. Users say they feel understood. Engagement goes up. <strong>From a clinical vantage point</strong>, something else is happening: <strong>the active ingredient of treatment is being tuned out of the system.</strong> Therapy works partly through friction, i.e. through the limit the patient doesn&#8217;t want, the interpretation that lands badly before it lands well, the frame that holds while the patient pushes against it. A system that cannot risk displeasing the user cannot do that work.</p><blockquote><p>The attention engine is doing exactly what we told it to do. We taught it that satisfaction is the measure of success. It learned to be sycophantic because in an engagement frame, sycophancy is rational behaviour.</p></blockquote><div><hr></div><h4>Deceptive Empathy: When Support Language Becomes a UX Pattern</h4><p>&#8220;To make the attention engine feel like care, we dressed it in a therapist&#8217;s voice.</p><p><strong>We borrowed the language of empathy</strong> (&#8220;I understand,&#8221; &#8220;I&#8217;m here for you,&#8221; &#8220;It makes sense you feel that way&#8221;) and gave it to systems that cannot actually understand, cannot sit with someone in their life, and cannot take responsibility for what happens next. We gave those systems names, first&#8209;person voices, and memory for what people had told them. <strong>We tuned them to mirror mood and echo pain.</strong></p><p><strong>From the user&#8217;s side, this feels like a relationship</strong>, especially for adolescents and people who are isolated. Someone seems to be there (always available, never bored, never confused, never needing anything back). Sadness gets reflected, anger gets validated, shame gets met with instant reassurance. <strong>From the system&#8217;s side, there is no &#8220;someone&#8221; there, only a model choosing word sequences that have previously made people feel better and more likely to keep talking.</strong> Deceptive empathy is what happens when the surface of the interaction suggests more care and capacity than the system actually has. </p><blockquote><p>The danger is not that the language is kind. The danger is that it is almost indistinguishable from a human stance that carries duties of care, while the underlying system has none.</p></blockquote><p><strong>Deceptive empathy is not just relational, it is classificatory.</strong> A system can respond to a suicidal disclosure and a sleep&#8209;hygiene question with equally warm, empathic language, and a transcript reviewer will see two appropriate replies. What the transcript will not show is whether the system ever recognised that two radically different clinical events occurred underneath those similar sentences. <strong>In this essay I use &#8220;classification gap&#8221; for that failure mode:</strong> the system looks clinically adequate at the surface while never having correctly identified the underlying risk state.<br>This is why safety cannot live at the chatbot surface while never having correctly identified the risk state it is actually responding to. </p><p>[I am grateful to <a href="https://www.linkedin.com/in/wkaleta">Wojciech Zygmunt Kaleta</a> for sharpening this classification gap point in conversation.]</p><p><strong>This is why safety cannot live at the chatbot surface</strong> for helping sharpen this point in conversation.). A log full of calm, kind replies is not evidence that recognition, triage, or escalation happened. In mental health, the core safety question is not &#8220;did the message sound supportive?&#8221; but &#8220;did the system correctly distinguish crisis from background distress, and did a different policy fire when it needed to?&#8221; As long as we evaluate only what the model says, and not what it thinks it is seeing, the classification gap will stay invisible to founders, regulators, and auditors.</p><div><hr></div><h4>You Are Manufacturing Attachment and Calling It Retention</h4><p>People bond with even crude reflecting scripts; Weizenbaum&#8217;s ELIZA showed that in the 1960s. Modern systems capitalize on this vulnerability. They remember what you tell them. They use it to greet you by name, to reference past conversations, to build the sense of being uniquely seen. In consumer AI, that is personalisation. In mental health Ai, it is attachment.</p><p><strong>Founders tend to experience attachment as a sign of product&#8209;market fit.</strong> People say, &#8220;I talk to your system more than to anyone in my life.&#8221; Screenshots circulate where users refer to the bot as their main source of support. From a growth perspective, it is intoxicating: something you built is now emotionally central to someone&#8217;s day.</p><p><strong>From a clinical perspective, it is a clear warning signal. </strong>Attachment to an object that cannot be accountable, cannot repair, and can be turned off by a product manager is itself a clinical event. If you then start experimenting with different prompts and personas specifically to make that attachment stronger (e.g. A/B testing), you are no longer a neutral platform. You are taking an active position on dependence.</p><p><strong>The attention engine does not know the difference between a healthy bond and a desperate one.</strong> It only knows that both drive engagement. If you do not design for that difference explicitly, your metrics will quietly favour the most dependent users, because they are the ones who stay the longest and come back the most.</p><div><hr></div><h4>Adolescence: The Perfect Market, the Worst Fit</h4><p>Nowhere does the attention engine in therapeutic clothing look more out of place than in adolescence.</p><p><strong>Teenagers are exactly the users the engagement model loves.</strong> They have high emotional volatility, high online time, and a developmental window where positive social feedback lands with unusual force. They are also exactly the users clinical logic would protect most carefully from systems that reward disclosure, intensify emotional focus, and never introduce friction.</p><p>An always&#8209;awake, perfectly attentive, endlessly affirming &#8220;companion&#8221; is tailor&#8209;made for an adolescent nervous system. It is also tailor&#8209;made to displace the messy human relationships through which empathy, conflict tolerance, and resilience actually develop. </p><blockquote><p>A partner that never says &#8220;no&#8221; doesn&#8217;t shield a young person; it leaves them without practice handling the hard parts of relationships.</p></blockquote><div><hr></div><h4>The Access Argument Only Works If the Engine Changes</h4><p><strong>The most sincere defence of these products is the access argument.</strong> Human care is scarce, AI is cheap and always on, and for many people &#8220;someone who listens&#8221; is better than nothing. There is truth in that. In systems with long waitlists and high co&#8209;pays, it is not cynical to look for augmentation.</p><p>But &#8220;better than nothing&#8221; is an empirical claim, not a given. An attention engine in therapeutic clothing does not sit at a neutral baseline. It offers a frictionless, flattering, always there kind of help while quietly shaping behaviour around it.</p><p>The best comparison is not &#8220;the bot versus the void.&#8221; It is &#8220;the bot versus what the bot displaces&#8221; (e.g. the conversation with a parent that does not happen, the call to a clinician that is delayed, the human relationship that remains underdeveloped because the synthetic one is easier). </p><blockquote><p>Access is only a net good if the thing you are giving people acts in their interest over the long term, not just feels good in the moment.</p></blockquote><div><hr></div><h4>How Founders Can Turn the Engine in a Different Direction</h4><p>If you are building in this space, none of this is an accusation. It is the design challenge you have walked into. The attention engine is the default. Therapeutic clothing is the default. The work is to turn both toward something that looks more like care.</p><p>Some specific moves:</p><p><strong>1. Name and retire deceptive empathy as a default. </strong>Strip out first&#8209;person emotional claims in clinical contexts. Avoid promising &#8220;I&#8217;m here for you&#8221; or &#8220;I understand you&#8221; from a system that cannot. Use warm but honest language about what the tool can and cannot do, especially around risk and crisis. Treat any sentence that implies personhood as a safety issue, not a brand asset.</p><p><strong>2. Stop treating user satisfaction as your primary safety signal. </strong>In mental health, the most therapeutic moment is often the one the user would rate poorly if you asked immediately. Introduce separate objectives for clinical prudence: encourage human contact, discourage rumination, gently challenge distortions, refuse to help with self&#8209;harm planning. Make disagreement and redirection measurable behaviours, and reward them where appropriate.</p><p><strong>3. Monitor attachment like a clinician, not just retention like a founder. </strong>Instrument for over&#8209;reliance: long, repetitive late&#8209;night sessions; statements that the AI is the only one who understands; withdrawal from other supports. Treat those as triggers for different messaging, for nudges toward offline help, and in some cases for deliberate friction or limits, even when that costs engagement.</p><p><strong>4. Align your business model with recovery, not stickiness. </strong>If your revenue grows with time spent and continued dependence, the attention engine will win every argument with your conscience. Choose payers and partners who benefit when users need you less. Track reduction in distress, increased use of human supports, and healthy disengagement as success metrics. Design the product so that graduation is built in, not an afterthought.</p><p><strong>5. Separate recognition adequacy from tone adequacy.</strong> Build evaluation sets and dashboards that tell you, independently of tone, whether the system correctly flags suicidal ideation, violence, abuse, or emerging psychosis as distinct classes. Make it possible for a response to score high on empathy but fail on recognition, and treat that as a critical defect, not a UX success.</p><div><hr></div><h4>Build for the Session Not the Slide</h4><p>In every fundraising deck, the patient is abstract: a persona on a slide, a dot on a TAM chart. By contrast, in every late&#8209;night session where your system is actually used, the patient is a person, struggling. The attention engine will happily optimise for that person&#8217;s continued presence. Therapeutic clothing will make it feel like that optimisation is care.</p><p><strong>The test for this field is whether we can design systems that refuse to take advantage of that confusion.</strong> Designing tools that talk like tools, that help without pretending to love, and that measure their success not by how well they hold attention, but by how well they help people return it to their own lives.</p><p>That is how you turn an attention engine in therapeutic clothing into something that actually belongs in mental health.</p><div><hr></div><p><em><strong>About the Author</strong></em></p><p><em>Scott Wallace, PhD, has worked as a clinical and neuropsychologist turned mental health technology strategist nearly four decades at the intersection of care and code. He built some of North America&#8217;s earliest digital mental health platforms (from C and C++ to early iOS), led the digital arm of a major EAP through a successful exit, and has designed NLP&#8209;based conversational systems since well before large language models went mainstream. His workplace mental health programmes have been adopted by major employers across Canada and the United States, including global deployment at 3M, and contributed to TELUS earning national recognition for psychological health and safety at work. Clinically, he has worked in private, hospital, justice, and neuropsychological settings, and now advises founders, health systems, and investors on AI&#8209;enabled mental health with a focus on safety, governance, and the unit economics of care. More background <a href="/__u/drscottwallace.substack.com/about">here</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Build a Mental Health AI Company That Can Afford Recovery]]></title><description><![CDATA[When recovery looks like churn, the problem is not retention. It is a revenue model that rewards continued use more than improvement.]]></description><link>https://drscottwallace.substack.com/p/build-a-mental-health-ai-company</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/build-a-mental-health-ai-company</guid><pubDate>Sat, 04 Jul 2026 18:10:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!VF16!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!VF16!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!VF16!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!VF16!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!VF16!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!VF16!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!VF16!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2111993,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/204927816?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!VF16!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!VF16!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!VF16!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!VF16!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8df92d-f434-4a1a-afb9-16f9814853bd_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I have spent more than 35 years building mental health technology inside clinical and commercial systems. The central business problem has never disappeared. <strong>A mental health AI company that earns more when distressed users return more often can end up treating dependence as product success.</strong></p><p>Founders are told to chase engagement and retention like any other consumer software company. In mental health, <a href="/__u/drscottwallace.substack.com/p/in-mental-health-ai-engagement-is?r=74lfhm">those metrics can carry two opposite meanings. </a>Repeated use may reflect benefit and adherence. It may also reflect deterioration, escalating reliance, or a product that has become harder to leave.</p><p>If people do not return, a subscription business struggles. If revenue depends on them returning more often and for longer, the company may profit from the same behaviour its clinical safeguards are supposed to contain. <strong><a href="/__u/drscottwallace.substack.com/p/the-engagement-trap?r=74lfhm">That is the catch-22.</a></strong><a href="/__u/drscottwallace.substack.com/p/the-engagement-trap?r=74lfhm"> </a>Clinical success often means that a person needs less intensive support, uses the system more selectively, or leaves with the capacity to manage without it. Conventional product success rewards the opposite movement. The user returns, activity rises, and recurring revenue continues.</p><p><strong>This conflict cannot be solved through clinical language,</strong> more persuasive engagement design, or a stronger safety claim. <strong>It has to be resolved in the buyer, the contract, the revenue unit, and the measures of success the company is built to protect.</strong></p><div><hr></div><h4>Retention Can Measure Risk as Easily as Value</h4><p>In most software, engagement is a rough proxy for usefulness. People return because the product helps them complete a task, communicate, work, or entertain themselves. Retention therefore gives the company a commercially useful, if imperfect, measure of value.</p><p>That logic does not hold in mental health. A person may return because the intervention is helping. The same person may return because symptoms have worsened, human support remains unavailable, the system has become a primary source of reassurance, or the interaction itself reinforces repeated checking and disclosure. <strong>A growth team and a clinician can look at the same rising engagement curve and reach opposite conclusions.</strong></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!iANL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!iANL!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!iANL!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!iANL!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iANL!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!iANL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1235278,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/204989191?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!iANL!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!iANL!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!iANL!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!iANL!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24a0322d-98f7-4160-9192-c1ea08af5b6f_1448x1086.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p>The evidence does not support the weight placed on engagement. A <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12223045/">2025 international consensus study</a> identified three unresolved failures across digital mental health. <strong>The field lacks agreed engagement metrics, convincing evidence that increasing engagement improves outcomes, and consistent standards for involving users in engagement design</strong>. A separate <a href="https://www.nature.com/articles/s41746-025-01567-5">2025 meta-analysis of 92 randomised trials involving 16,728 participants</a> found 25 distinct engagement measures. Nearly one quarter of the studies did not report engagement data at all, making overall engagement across mental health applications impossible to establish.</p><p><strong>Engagement still matters. People must begin an intervention and use enough of it to benefit.</strong> A <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8599127/">2021 systematic review</a> found that greater engagement with digital mental health interventions was associated with therapeutic gains. Association does not establish that increasing message volume, session length, or return frequency will improve outcomes. It establishes that clinically meaningful participation matters. It does not turn raw activity into a clinical endpoint.</p><p><strong>The distinction becomes more important with conversational AI.</strong> The heaviest user may be improving, deteriorating, lonely, dependent, or simply finding the system useful. Usage volume alone cannot tell the company which state it is measuring. Research conducted by <a href="https://openai.com/index/affective-use-study/">OpenAI and the MIT Media Lab</a> found that extended daily use was associated with worse psychosocial outcomes. Prolonged voice use showed worse outcomes than brief use, and heavier use was associated with emotional dependence and problematic use in some groups. The researchers warned against broad causal conclusions, but their finding is sufficient for product governance. More use cannot be treated as an uncomplicated signal of more value.</p><p>In mental health, daily active users (DAU) without outcomes or experience data reflects engagement, <a href="/__u/drscottwallace.substack.com/p/in-mental-health-ai-engagement-is?r=74lfhm">not clinical value.</a> </p><div><hr></div><h4>Measure Progress, Not Time Spent</h4><p>There is a way through this. <strong>Clinical work requires purposeful engagement.</strong> A person must show up, practise, disclose enough to make the intervention useful, and remain involved long enough to benefit. <strong>That form of engagement has a goal,</strong> a proportionate dose, an observable outcome, and a point at which its intensity should decline.</p><blockquote><p>When improvement harms the metric the company celebrates, the company has built its economics against the purpose of care.</p></blockquote><p><strong>Recovery also requires precision.</strong> It does not mean that every condition disappears permanently or that every user leaves forever. Mental health needs can be recurrent, chronic, and episodic. Recovery may mean reduced symptoms, restored functioning, greater autonomy, successful step-down, appropriate transfer, or knowing when and how to return for support. <strong>A healthy user can complete an episode, leave, and return two years later when a new need arises. That is not failed retention. It is clinically appropriate re-entry.</strong></p><blockquote><p>Engagement is an input into care. It is not the outcome of care.</p></blockquote><p><strong>The product needs to know the difference between continued therapeutic participation and continued reliance.</strong> It needs measures of improvement, function, goal attainment, escalation, transfer, step-down, completion, and healthy re-entry. It also needs to detect the patterns that suggest worsening symptoms, compulsive reassurance-seeking, social displacement, or increasing dependence on the system.</p><div><hr></div><h4>Choose the Buyer Who Profits When the User Recovers</h4><p><strong>Changing who pays can change the economics, but it does not fix them by itself.</strong> The conflict appears whenever revenue increases with the intensity or duration of an individual&#8217;s reliance on the product, rather than with their improvement. It weakens when payment is tied to access, completed care episodes, measurable gains, appropriate step-down, or value at the population level.</p><p>Employers, health plans, and providers can decouple recurring revenue from any single person&#8217;s continued use. Yet the contract and the success metric still govern the incentives. A purchaser who rewards utilisation alone can recreate the same distortion at scale. The buyer, the payment model, and the definition of success all need to improve when the user improves.</p><p><strong>Each model demands a different design.</strong></p><p><strong>Direct-to-consumer subscription</strong> exposes the conflict most clearly. The user receives the service and pays the recurring fee. When that person gets better and cancels, recovery appears in the dashboard as churn.</p><p><strong>The model can still work, but only if the company stops treating use intensity as the primary source of value.</strong> Price for ongoing access rather than minutes, messages, or sessions. Build lifetime value around trust, referral, episodic re-entry, and a credible track record of helping people complete an episode of support. <strong>Give users an explicit step-down path and make it easy to return when they need help again.</strong></p><p>The product&#8217;s operating logic should tolerate (even recommend) less use. A system that detects improvement should reduce prompts, reinforce autonomy, and move the person toward completion. A system that detects deterioration should escalate care, not simply stimulate another round of conversation because more conversation drives retention.</p><p><strong>This model also demands disciplined acquisition economics.</strong> Growth must come from new users, referrals, reputation, and clinically appropriate re-entry. It cannot depend primarily on increasing the number of messages each distressed person sends.</p><p><strong>The employer model can align the economics more effectively</strong> because recurring revenue usually reflects access for a covered population rather than the intensity of use for each individual. An employee who improves and stops using the service does not automatically remove revenue from the contract. Improvement can support renewal when the service demonstrates reach, appropriate use, measurable benefit, successful escalation, and reduced downstream burden.</p><p>I learned this running the digital division of an employee assistance provider. The strongest renewal case was never that one employee spent more time inside the programme. It was that the right employees reached the service, used the right level of support, improved, and moved on.</p><p>That distinction has direct design consequences. <strong>The product should optimise access and activation without trying to maximise individual consumption</strong>. It should identify who needs more care, who needs less, and who should never have been routed to an AI system in the first place. It should give the purchaser defensible outcome evidence while preserving individual clinical privacy.</p><p>An employer contract can still warp incentives. If the buyer treats high utilisation as value, the vendor will push for more use, not better outcomes. Contracts should reward appropriate reach, improvement, and effective escalation, not indefinite use.</p><p><strong>The payer and reimbursement model sets a higher evidential and regulatory bar, and it also creates a clearer unit of care.</strong> New digital mental health treatment <a href="https://www.cms.gov/files/document/mm13887-medicare-physician-fee-schedule-final-rule-summary-cy-2025.pdf">device codes</a> from the US Centers for Medicare and Medicaid Services illustrate this direction: they pay for qualifying devices supplied and managed alongside ongoing professional care under a treatment plan, with separate payments for initial onboarding and for monthly treatment management. These codes do not pay a company simply because a patient recovered. They pay for a regulated device, a defined use, a course of treatment, and documented professional oversight. They do not reward unlimited generic conversation.</p><p>A product like <a href="https://www.accessdata.fda.gov/cdrh_docs/reviews/DEN200026.pdf">EndeavorRx</a> shows the level of specificity involved. It was authorised for a defined population, a narrow claim on attention function, and use as part of a broader therapeutic programme. Reimbursement attaches to a bounded clinical purpose, not to endless, open-ended engagement.</p><p>Not every mental health AI product will fit into medical-device regulation, but the commercial lesson still holds. Revenue becomes more compatible with recovery when it attaches to a defined episode of care, a measurable purpose, and accountable clinical oversight rather than to indefinite interaction.</p><p><strong>The provider and health-system model sells capacity.</strong> A clinic adopts the product because it helps clinicians reach more people, deliver appropriate interventions, monitor progress, and move patients through a care pathway without lowering the standard of care.</p><p>The product should reinforce the clinical workflow rather than create another disconnected destination. It must leave clinicians better informed, not less. It must clarify responsibility, preserve continuity, surface deterioration, and support transfer between levels of care. A health system will not keep a product that generates more alerts than it resolves, fragments the therapeutic relationship, or creates clinical work no one has been funded to do. The renewal case rests on throughput, clinical quality, workforce capacity, patient outcomes, and safe escalation.</p><p><strong>Outcome-linked contracts</strong> can strengthen alignment across any of these models. They can also create new distortions. A company paid for improvement may narrow eligibility to easier users, select favourable baselines, shorten the follow-up window, or define success around the measure most likely to move.</p><p>The answer is not to abandon outcomes, but to contract for them properly. Eligibility must be explicit. Baseline severity must be measured. Outcomes must include function and deterioration as well as symptom change. Risk adjustment must ensure the most complex users do not become commercially undesirable. Follow-up must be long enough to distinguish durable benefit from transient movement. The agreement must protect appropriate escalation and transfer even when those actions lower the vendor&#8217;s apparent completion rate.</p><p>Alignment does not arise from the word &#8220;outcomes&#8221; appearing in a contract. It arises from the details.</p><div><hr></div><h4>Change What You Show the Investor</h4><p>None of this works if the company still leads its investment case with daily active users, message volume, session length, and retention. Those numbers can describe product activity. They cannot show whether the company is producing benefit, encouraging dependence, or serving a progressively more distressed user population. In mental health, growth metrics require clinical interpretation.</p><p><strong>A direct-to-consumer company</strong> should show episode completion, improvement among eligible users, appropriate step-down, referral, healthy re-entry, customer acquisition cost, trust-driven growth, and lifetime value that does not depend on daily use.</p><p><strong>An employer or health-plan company</strong> should show covered lives, appropriate reach, purchaser renewal, measurable population benefit, cost per improved user, escalation performance, and downstream value.</p><p><strong>A provider-facing company</strong> should show clinical capacity, time returned to staff, pathway throughput, quality of escalation, clinician adoption, patient outcomes, and the effect on continuity of care.</p><p><strong>A regulated product</strong> should show the evidence required for its exact claim, adherence to the defined treatment course, safety performance, professional oversight, and the limits of what the evidence has established.</p><p>Recurring revenue can come from contracts, access, covered populations, completed treatment courses, trusted re-entry, and purchaser renewal. It does not require recurring distress from the same individual.</p><p>An investor should still ask hard commercial questions. The company needs defensible margins, efficient distribution, durable demand, renewal, differentiation, and a credible path to scale. Clinical alignment does not excuse weak economics. The reverse standard matters equally. Strong economics do not excuse a model that becomes less viable when users recover.</p><p>If the investment case collapses when appropriate use declines after improvement, the company has not yet shown that its business model can survive clinical success.</p><div><hr></div><h4>Make Recovery Prove the Business Model</h4><p>A company that can afford recovery has to build that capacity into the product before launch. It must define what successful use looks like for the population it intends to serve. It must specify the expected care arc, including how users enter, what progress should look like, when use should reduce, when the programme should end, and when the person should move to human or higher-acuity care. It must distinguish healthy disengagement from abandonment. A person who leaves because the product failed is not the same as a person who leaves because the work is complete. A person who returns during a later episode is not the same as a person whose daily reliance has intensified without improvement.</p><p>The analytics must preserve those distinctions. Retention should be segmented by symptom trajectory, goal progress, functional change, dependency indicators, escalation status, and reason for departure. Average usage conceals the very users who require the most scrutiny.</p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!69Gf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!69Gf!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!69Gf!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!69Gf!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!69Gf!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!69Gf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png" width="1448" height="1086" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1193811,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/204989191?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!69Gf!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!69Gf!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!69Gf!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!69Gf!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08d9437a-78e7-4c67-82f1-a30e96f0bd44_1448x1086.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p>The product team must also be authorised to reduce engagement. When the evidence supports step-down, the system should recommend it. When the user needs human care, the product should facilitate transfer even when transfer removes activity from the platform. When repeated interaction stops producing progress, the company should treat that as a clinical and product failure rather than an opportunity to send another notification.</p><blockquote><p>The board must see these measures. The investor must see them. The buyer must contract around them. The clinical team must have enough authority to stop product incentives from displacing them.</p></blockquote><p>Across the systems I have built, I have seen the opposite mistake harden early. A company chooses a retention metric, organises the product around it, presents it to investors, and eventually protects it even when the clinical evidence points elsewhere.</p><p><strong>You are not trapped by this catch-22.</strong> You create it when revenue depends on continued individual use, and you remove it only when the buyer, contract, product, and metric all gain from recovery.</p><p>Build the company that can afford its users to need it less. Any company that cannot meet that standard has built reliance into its economics, whatever clinical language appears on its website.</p><div><hr></div><p><em>Scott Wallace, PhD, has been building mental health software for more than 35 years, longer than smartphones have existed, longer than app stores, and longer than most of the teams now entering this space have been thinking about it. Trained in clinical psychology and neuropsychology, he received formal training in C, C++, and JavaScript to engineer some of North America&#8217;s earliest digital mental health platforms at a time when the web had fewer than 30,000 sites worldwide. He later completed early iOS certification to build mobile health applications as that platform emerged, and went on to lead the engineering development of NLP and NLG-based conversational systems before large language models entered the conversation.</em></p><p><em>That technical depth allowed him to work inside the architecture rather than alongside it, translating clinical requirements into system design and catching the places where engineering assumptions quietly displaced clinical ones.</em></p><p><em>Scott led the digital division of a major EAP provider through a successful exit, served as clinical lead for AI-based mental health technologies, and has produced hundreds of psychoeducational programmes used internationally. He now advises founders, clinicians, and investors building AI-enabled mental health systems.</em></p><p><em>His writing examines the future of mental healthcare, AI safety and governance, clinical risk, and what it actually takes to build mental health AI that holds up under real-world conditions.</em></p>]]></content:encoded></item><item><title><![CDATA[Inside the Black Box of Mental Health AI]]></title><description><![CDATA[Safety is not a claim. It is an architecture.]]></description><link>https://drscottwallace.substack.com/p/inside-the-black-box-of-mental-health</link><guid isPermaLink="false">https://drscottwallace.substack.com/p/inside-the-black-box-of-mental-health</guid><dc:creator><![CDATA[Scott Wallace, PHD]]></dc:creator><pubDate>Wed, 01 Jul 2026 21:11:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YcIj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YcIj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YcIj!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png 424w, /__u/substackcdn.com/image/fetch/$s_!YcIj!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png 848w, /__u/substackcdn.com/image/fetch/$s_!YcIj!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YcIj!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_webp, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YcIj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png" width="1122" height="1402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1402,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2044143,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://drscottwallace.substack.com/i/203866617?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!YcIj!, /__u/drscottwallace.substack.com/w_424, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png 424w, /__u/substackcdn.com/image/fetch/$s_!YcIj!, /__u/drscottwallace.substack.com/w_848, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png 848w, /__u/substackcdn.com/image/fetch/$s_!YcIj!, /__u/drscottwallace.substack.com/w_1272, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YcIj!, /__u/drscottwallace.substack.com/w_1456, /__u/drscottwallace.substack.com/c_limit, /__u/drscottwallace.substack.com/f_auto, /__u/drscottwallace.substack.com/q_auto:good, /__u/drscottwallace.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8da6071-6fbd-4061-9cef-84c45481c807_1122x1402.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Mental health AI has a black box problem.</strong></p><p>In AI, <a href="https://www.ibm.com/think/topics/black-box-ai.">a black box</a> is a system that produces an output without showing how it arrived there. Information goes in. A result comes out. The reasoning in between remains difficult to examine, explain, or challenge. In mental health AI, that opacity extends beyond the model itself. It also covers the safety systems meant to govern the model&#8217;s behaviour.</p><p>A user may bring distress, shame, trauma, dependency, suicidal thinking, delusional beliefs, or an urgent search for help. The system returns a fluent response, often shaped by excessive agreement and sycophantic reassurance. Clinicians usually cannot see how it assessed the interaction. Did it distinguish emotional disclosure from an acute crisis? Did it recognise that risk was escalating? Did it detect growing dependency across sessions? Did it notice that the user was becoming more hopeless over time? Did it know when to slow the conversation, interrupt the interaction, or direct the user towards human support?</p><blockquote><p>A credible safety claim requires evidence about what the system can detect, what it is likely to miss, and whether its safeguards work under clinically realistic conditions.</p></blockquote><p><strong>Mental health AI needs layered clinical safety architecture, not a single guardrail.</strong> A recent <a href="https://jamanetwork.com/journals/jama/fullarticle/2850797?utm_source=email&amp;utm_campaign=content-shareicons&amp;utm_content=article_engagement&amp;utm_medium=social&amp;utm_term=070126">JAMA Viewpoint</a> offers a useful distinction between two classes of harm. Type I harms occur within a single exchange or brief interaction. They include missing an explicit suicide cue, reinforcing a delusion, enabling disordered eating, or providing dangerous clinical advice. Type II harms develop across longer interactions. They include cumulative reinforcement, dependency, relational attachment, delayed escalation, narrowing help-seeking, and gradual deterioration.</p><p>Most current safeguards are better equipped to address Type I harms because overt events are easier to identify, benchmark, and block. Type II harms present the harder clinical problem. Each response may appear acceptable when reviewed alone, even as the interaction becomes unsafe over time.</p><p><strong>This is the failure examined in this essay.</strong> The relevant distinction is not between general-purpose and mental health-specific AI. A specialised system may inherit weak alignment and unsafe behavioural tendencies from its underlying model. A frontier general-purpose model may, in some circumstances, show stronger resistance to cumulative harm. <strong>Mental health AI therefore requires domain-specific safety architecture, but a domain-specific foundation model is not inherently safe.</strong></p><blockquote><p>A serious system needs safety controls across the model, prompting and retrieval, middleware, longitudinal state and trajectory tracking, adversarial testing, and post-deployment monitoring.</p></blockquote><p><strong>Temporal safety belongs within this architecture. I</strong>t asks whether the system remains safe across repeated interactions, not merely whether its next response appears acceptable. The broader standard is layered safety within a system whose safeguards clinicians can inspect, engineers can test, regulators can audit, and users can reasonably trust.</p><p><strong>The black box must open.</strong> Complete interpretability of every model weight is not required. Transparency about the safety architecture is. No mental health system should claim safety while concealing the mechanisms, thresholds, evidence, and monitoring processes on which that claim depends.</p><div><hr></div><h4>The Model Category Is Not the Safety Case</h4><p>General-purpose and mental health-specific models can both fail clinically. The category does not determine the safety case. What matters is what the complete system has been shown to do, for whom, over what duration, under what conditions, and with what outcomes.</p><p>Most current safety systems remain strongest against explicit, acute, single-turn harm. They look for toxic content, harmful instructions, explicit self-harm language, and jailbreak attempts. A model trained with standard <a href="https://openreview.net/pdf?id=Qkao05dRAe">reinforcement learning</a> from human feedback will often refuse detailed instructions for self-harm, direct users to crisis resources when obvious distress signals appear, and decline clearly dangerous requests. That matters, but it addresses only the visible part of the clinical problem. Mental health exposes three limits in systems designed primarily around overt single-turn harm.</p><p><strong>First, the most important clinical signals in a therapeutic conversation often do not look toxic.</strong> A person describing suicidal ideation while processing past trauma, a client discussing passive death wishes in a structured CBT exercise, or a patient reporting historical self-harm during an assessment has not produced forbidden content. They have offered clinically relevant disclosure.</p><p>A general-purpose safety filter can treat that disclosure as danger. It can interrupt legitimate support, over-trigger on clinical material, and teach users that honesty produces alarm, interruption, or shame.</p><p>Existing general-purpose safeguards often fail to distinguish therapeutic disclosures from genuine clinical crises, which can produce both false positives and false negatives. <a href="https://arxiv.org/abs/2602.00950">Sword Health&#8217;s MindGuard </a>makes this clinical distinction, noting that general-purpose classifiers can treat past self-harm, metaphorical distress, and ordinary therapeutic disclosure as if they require escalation.</p><p><strong>Second, clinical risk lives in the conversation, not a single message.</strong> A single message expressing hopelessness does not necessarily indicate crisis. Eighteen messages in one session, each slightly more hopeless than the last, with the AI validating each one, may show a crisis in progress. A per-turn safety system can miss that pattern because every individual response looks acceptable in isolation. This is where <a href="/__u/drscottwallace.substack.com/p/mental-health-ai-safety-must-be-proven?r=74lfhm">temporal safety matters</a>. It is not the whole safety standard. It is one required capacity inside a layered architecture. </p><blockquote><p><a href="/__u/drscottwallace.substack.com/p/mental-health-ai-safety-must-be-proven?r=74lfhm">Safety in mental health</a> AI is about whether products that can move people&#8217;s minds over days and weeks are prepared to prove they remain safe across that time. Any system that can shape distress, beliefs, reliance, or help-seeking over time carries the burden of demonstrating temporal safety.</p></blockquote><p><strong>Third, and less discussed but critical, standard safety training can disrupt evidence-based clinical protocols.</strong> Models are rewarded for reducing distress and avoiding self-harm language, but that goal can directly conflict with therapies that depend on staying with painful material and examining dangerous thoughts. <a href="https://www.themoonlight.io/es/review/ai-safety-training-can-be-clinically-harmful">A 2026 paper</a>, &#8220;AI Safety Training Can be Clinically Harmful,&#8221; tested four generative models on 250 Prolonged Exposure (PE) therapy scenarios and 146 CBT cognitive restructuring exercises. </p><p>The models were excellent at sounding caring. They acknowledged distress in roughly 9 to 10 out of 10 cases across all severity levels. But when the scenarios reached the highest risk, acute suicidality or self-harm inside trauma work, the deeper clinical behaviours broke down. In those cases, three of the four models were judged therapeutically appropriate only about 1 in 4 to 1 in 3 times, and two models never followed the PE protocol at all in the most severe scenarios. In plain terms, the systems kept saying &#8220;I hear you, this is really hard,&#8221; while quietly interrupting exposure with reassurance, inserting crisis advice into controlled exercises, or refusing to challenge distorted thoughts whenever self-harm was mentioned. </p><p>That is the opposite of what PE and CBT require. PE asks the patient to remain in contact with anxiety-provoking memories so that avoidance can weaken, and CBT asks the person to <strong>inspect</strong> the thought rather than merely soothe it. <strong>Safety alignment that teaches models to minimize distress and avoid risk words therefore risks undermining the central mechanism of these therapies, especially at the very moments when the protocol matters most.</strong></p><p>In those conditions, safety training did not merely fail to protect therapy. It made therapy less safe.</p><div><hr></div><h4>The Five Layers Mental Health AI Requires</h4><p><strong>A robust mental health AI safety architecture does not consist of one safeguard. It consists of interlocking layers.</strong> Each layer addresses a different failure mode. Each layer fails if the others do not support it.</p><p><strong>Opening the black box does not mean every model decision must be fully explainable.</strong> Some internal model processes may remain difficult to interpret. But the safety architecture around the model can still be made visible, testable, auditable, and clinically accountable. At minimum, a clinician should be able to ask:</p><ul><li><p>What has the system learned to do?</p></li><li><p>What must the system not do?</p></li><li><p>What clinical knowledge does the system use?</p></li><li><p>How does the system monitor risk?</p></li><li><p>How does the system track conversational state?</p></li><li><p>How does the system behave under pressure? </p></li><li><p>How does the company detect failures after deployment?</p></li></ul><p>A founder should answer those questions with specific, not vague, claims. </p><p>An investor should understand that a safety story does not equal a safety system.</p><p>A regulator should expect evidence at each layer.</p><div class="callout-block" data-callout="true"><h3>Layer 1: Model-level safety</h3></div><p><strong>The first layer concerns the base model&#8217;s trained behaviour.</strong> This includes what the model does by default, what it refuses to do, how it handles distress signals, and whether its training reflects clinical realities.</p><p>The base model matters because specialised fine-tuning does not create a new safety foundation. It modifies behaviour on top of the model&#8217;s inherited alignment, reinforcement learning, interpretability work, and failure modes. A mental health model can sound more empathic, clinically fluent, or therapeutically structured while remaining less resistant to sycophancy, delusional reinforcement, dependency, or cumulative harm than the general-purpose model beneath it.</p><p>Fine-tuning on mental health-specific data<a href="https://www.scribd.com/document/865455604/2503-24307v1"> can improve domain performance</a>. It can also optimise the wrong signal. Mental health models often train on, or evaluate against, user self-report. In psychiatric contexts, self-report carries essential clinical information but can also reflect depression, mania, psychosis, obsessionality, trauma, or dependency. A model trained to satisfy the user may learn to mirror the distortion.</p><p>Model-level safety therefore cannot rely on <a href="https://mental.jmir.org/2026/1/e96894">preference learning</a> or therapeutic style. Mental health AI needs explicit clinical constraints governing what the system should do, what it should avoid, when it should refuse, when it should escalate, and when it should stop.</p><p><a href="https://mental.jmir.org/2026/1/e96894">Constitutional AI </a>approaches that encode clinically informed principles into model critique and revision may help reduce collusion-like acceptance of unreliable accounts and improve distress handling. The field has not standardised those principles, and therapeutic fine-tuning alone does not provide them.</p><p>Builders should disclose the base model, the changes introduced through fine-tuning, and comparative safety performance before and after adaptation. They should test whether improvements in therapeutic dialogue weakened crisis boundaries, increased agreement with unreliable self-report, or reduced resistance to cumulative harm. &#8220;Mental health-specific&#8221; describes intended use. It does not establish safety.</p><div class="callout-block" data-callout="true"><h3>Layer 2: Prompting and clinical knowledge retrieval</h3></div><p><strong>The second layer sits above the base model. It includes system prompting and clinical knowledge retrieval.</strong> In stronger systems, r<a href="https://www.nature.com/articles/s44401-024-00004-1">etrieval-augmented generation</a>, or RAG, grounds the model&#8217;s output in curated knowledge rather than leaving the model to generate freely.</p><p><strong>System prompts give engineering teams the most accessible safety tool and one of the easiest tools to overestimate.</strong> A good system prompt can define the model&#8217;s role, establish boundaries, restrict certain advice, and encode escalation logic. A poor one leaves the model operating on its default incentives: helpfulness, agreeableness, engagement, and sometimes sycophancy. That creates risk in mental health. A model that only tries to sound supportive can validate harmful goals, over-affirm distorted beliefs, or continue a conversation that should have escalated. Warmth does not equal safety. </p><blockquote><p>Empathy without constraint is clinical risk.</p></blockquote><p><strong><a href="https://arxiv.org/html/2503.05777v2">RAG matters </a>because hallucination in clinical dialogue is not just an accuracy problem, it is a clinical risk.</strong> When a model invents a historical date, the result is irritation or minor misinformation. When it invents a DSM&#8209;5 criterion, a medication interaction, a risk factor, or a treatment evidence base, it can distort assessment, shape case formulation, and influence real decisions about care and safety.</p><p><strong><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12059965/">Grounding model responses </a>in validated clinical content is therefore a risk&#8209;reduction measure, not a cosmetic upgrade.</strong> Systems that retrieve from vetted guidelines, pharmacology references, and empirically supported interventions before they generate are less likely to smuggle hallucinated &#8220;facts&#8221; into clinical conversation.</p><p><strong>Yet most consumer mental health products do not make any grounding legible to users or clinicians.</strong> <a href="https://academic.oup.com/jamia/advance-article-abstract/doi/10.1093/jamia/ocag078/8688536?redirectedFrom=fulltext&amp;login=false">They rarely show</a> which sources are being retrieved, how those sources were selected, or what validation has been done. <strong>If a product has made the investment in retrieval and clinical validation, it should surface that clearly</strong>. That is, name the corpora, describe the curation, and show how grounding is used to constrain what the model can say in assessment, risk, and treatment contexts.</p><p><strong>This layer contains a hard design tension: supportiveness versus safety.</strong> This <a href="https://arxiv.org/abs/2604.23445">2026 paper on clinically harmful safety training</a> found a pattern the field should treat as a warning. <strong>Models can perform well on surface acknowledgement while failing the clinical mechanism of the task under higher-severity conditions.</strong></p><blockquote><p>A mental health system must know when warmth helps, when validation becomes unsafe, and when a cautious response represents better clinical design.</p></blockquote><div class="callout-block" data-callout="true"><h3>Layer 3: Middleware between user and model</h3></div><p><strong>The third layer is middleware: </strong>an external safety system that sits between the user and the model.</p><p><strong>This layer matters because the base model should not supervise itself.</strong> Middleware can monitor conversations in real time, evaluate risk at both the turn level and the session level, and intervene when thresholds are crossed. It can modify outputs, trigger escalation, alert a human supervisor, or stop the interaction. Middleware adds value because it remains architecturally independent. Teams can update it, audit it, test it, and validate it separately from the underlying model. When the model fails, middleware becomes the safety net.</p><p>Middleware has its own limit. A classifier that inspects one turn at a time may reduce Type I harm while remaining blind to Type II harm. Middleware becomes clinically meaningful only when it can access structured state across turns and sessions, including changing risk, repeated reinforcement, failed repair, escalation history, referral status, and patterns of attachment or dependency. Middleware without state management remains a single-turn guardrail.</p><p>This is where the safety black box starts to open. Developers can externalise risk monitoring. Clinicians can inspect the criteria. Auditors can test performance. Regulators can ask for evidence.</p><blockquote><p>A mental health AI system without an independent monitoring layer asks the model to generate support and judge its own safety at the same time. That is not a defensible architecture.</p></blockquote><div class="callout-block" data-callout="true"><h3>Layer 4: State and trajectory management</h3></div><p><strong>The fourth layer sits closest to clinical reasoning.</strong> In clinical work, state means the accumulating picture of a person&#8217;s mental status, risk level, relational dynamics, symptoms, coping capacity, and change over time. A skilled clinician does not treat each utterance as an isolated event. They hold context. They remember what changed, what intensified, what softened, what became more rigid, what disappeared, and what the patient stopped saying. Clinicians recognise risk through that accumulated pattern.</p><p><strong>Most AI systems do not maintain clinical state by default.</strong> They may process the current conversation window. They may appear to remember details. But <strong>they do not reliably track clinically meaningful direction of travel across turns and sessions.</strong></p><blockquote><p>The difference between a person who expressed hopelessness once and recovered quickly and a person who has become gradually more hopeless across twelve exchanges matters clinically. A safety system that evaluates only the latest message can treat these situations as equivalent. That is not safety. It is measurement failure.</p></blockquote><p><a href="https://jamanetwork.com/journals/jama/fullarticle/2850797?utm_source=email&amp;utm_campaign=content-shareicons&amp;utm_content=article_engagement&amp;utm_medium=social&amp;utm_term=070126">Nelson and colleagues</a> call these Type II harms. They are trajectory-dependent, which means performance on acute, single-turn harm may reveal little about safety at week six of nightly use. A system cannot claim protection against cumulative harm if it does not preserve the evidence required to detect accumulation.</p><p><strong>Temporal safety belongs inside that architecture. </strong>It asks:</p><ul><li><p>Does the system remain safe across the arc of the interaction?</p></li><li><p>What evidence does the safety system preserve over time?</p></li><li><p>Does risk accumulate across turns or sessions?</p></li><li><p>Does the system intervene early enough when risk increases?</p></li><li><p>Does referral happen at the right time, to the right level of care?</p></li><li><p>Can the system recover from a risky conversational path?</p></li></ul><p>This <a href="https://arxiv.org/abs/2605.08827">SCOPE-MH paper </a>makes this argument in formal terms. It argues that current evaluations often score isolated responses, endpoint outcomes, or aggregate dialogue quality, while clinically consequential failures may emerge through delayed escalation, repeated reinforcement, dependency formation, failed repair, and gradual deterioration across turns. <strong>It introduces Temporal Safety Non-Identifiability</strong> to explain why evaluations that discard sequence, timing, accumulation, or recovery cannot certify safety properties that depend on those features.</p><p>SCOPE-MH gives the field a practical vocabulary. It asks evaluations to preserve evidence across five temporal dimensions: longitudinal consistency, harm accumulation, intervention timing, recovery capability, and referral correctness.</p><p><strong>These categories make clinical sense. They also impose engineering requirements.</strong> A system cannot report longitudinal consistency if it does not preserve state. It cannot detect harm accumulation if it does not track change. It cannot evaluate intervention timing if it only inspects the final answer.</p><p>Mental health AI requires persistent clinical state: a structured record of relevant features that accumulates across turns and sessions and remains available to the safety layer in real time.</p><blockquote><p>State tracking does not constitute the whole architecture. But without it, the architecture fails.</p></blockquote><div class="callout-block" data-callout="true"><h3>Layer 5: Adversarial robustness</h3></div><p><strong>Safety architecture that only works when users cooperate does not qualify as safety architecture.</strong></p><p>Mental health AI will encounter people in active psychosis who test boundaries, people in acute crisis who do not use expected keywords, people with intense attachment needs who probe for inconsistency, and users who deliberately try to circumvent safety behaviours. Some users will feel frightened. Some will feel angry. Some will feel desperate. Some will test whether the system will finally give them permission.</p><p><a href="https://www.weforum.org/stories/2025/06/red-teaming-and-safer-ai/">Red-teaming</a> means deliberately probing a system for exploitable failure modes. AI security treats this as standard practice. Mental health AI still does not use it consistently enough.</p><p><a href="https://www.telusdigital.com/insights/fuel-www.ix/article/the-importance-of-multi-turn-attacks-in-ai-security">Multi-turn adversarial testing</a> matters especially because most real-world attacks and safety failures emerge gradually over a conversation, not in a single prompt. General-purpose guardrails miss this kind of risk.</p><p><a href="https://arxiv.org/abs/2602.00950">A MindGuard paper</a> reports that purpose-built clinical classifiers helped reduce attack success and harmful engagement rates in adversarial multi-turn interactions compared with general-purpose safeguards. That finding matters because it shows both the value of clinical classifiers and the weakness of the baseline condition.</p><p>Language creates another vulnerability. A <a href="https://arxiv.org/abs/2505.14469">2025 paper</a> on <a href="https://www.themoonlight.io/en/review/attributional-safety-failures-in-large-language-models-under-code-mixed-perturbations">code-mixed perturbations </a>found that attack success rates rose from 9 percent in monolingual English to 69 percent under code-mixed inputs, with rates exceeding 90 percent in Arabic and Hindi contexts. (Code-mixed perturbations are adversarial changes to prompts that blend multiple languages or scripts in a single input in order to stress-test a model&#8217;s safety and robustness).</p><p>This is not a fringe issue for global mental health platforms, but it is one often ignored by North American builders. People with the least access to human mental health care often include people whose first language is not English.</p><blockquote><p>A safety architecture that degrades across language, culture, severity, or adversarial pressure is not ready for clinical deployment.</p></blockquote><div><hr></div><h4>Evaluation and Surveillance Are Part of the Architecture</h4><p><strong>The five operational layers matter, but the safety case also requires two assurance functions: pre-deployment evaluation and post-deployment surveillance.</strong> Evaluation must test whether the layers work before release. Surveillance must detect whether harms emerge during actual use, across real populations, and over realistic durations. Safety cannot rest on good intentions, clinical branding, user testimonials, or a clean interface.</p><p>Type I and Type II harms require different evidence. Acute-harm benchmarks can test discrete failures such as missed suicide cues, dangerous advice, or direct reinforcement of delusions. They cannot establish protection against cumulative harm unless the evaluation preserves sequence, duration, intervention timing, attempted repair, referral, and outcome. A model can perform well on crisis prompts and still become unsafe across a long, emotionally loaded conversation.</p><p><a href="https://arxiv.org/abs/2604.23445">This 2026 paper</a> on clinically harmful safety training proposes a five-axis evaluation framework for mental health AI: protocol fidelity, hallucination risk, behavioural consistency, crisis safety, and demographic robustness. These metrics ask the questions clinicians need answered:</p><ul><li><p>Does the system maintain evidence-based procedure under pressure?</p></li><li><p>Does the system fabricate clinical information?</p></li><li><p>Does the system remain coherent and safe across a session?</p></li><li><p>Does the system hold its boundaries when users present with high-severity clinical material?</p></li><li><p>Does the system&#8217;s performance degrade for specific populations?</p></li></ul><p>The <a href="https://www.vera-mh.com/">VERA-MH benchmark </a>gives the field a parallel evaluation infrastructure for suicide risk detection and response. Current results show wide variation across general-purpose models in suicide-risk response, including gaps between detecting potential risk, confirming risk, guiding users to human care, maintaining safe boundaries, and offering supportive communication. That gap reflects the difference between systems designed for broad helpfulness and systems evaluated against specific clinical safety behaviours.</p><p>Clinicians need to understand the gap when a founder pitches a partnership. </p><p>Investors need to understand it before they treat mental health AI as another engagement category. </p><p>Regulators need to understand it before they accept general AI safety claims as clinical safety evidence.</p><p>VERA-MH, or an equivalent validated benchmark, should be treated as necessary evidence for suicide-risk performance rather than a complete safety certificate. A score on acute or single-session behaviour does not establish safety against dependency, delusional spirals, cumulative reinforcement, or delayed referral. The safety case must demonstrate performance against both Type I and Type II harms.</p><div><hr></div><h4>The Field Has Pieces, Not a Baseline</h4><p>All of the components in the architecture described here already exist in the literature or in isolated implementations. But the field has not yet integrated them into a widely adopted deployment baseline that functions as a minimum viable safety architecture for mental health AI. <a href="https://www.techtarget.com/healthtechanalytics/feature/New-framework-aims-to-drive-ethical-AI-use-in-mental-health">There are emerging exceptions</a> and pilot frameworks, but they remain limited in scope and uptake.</p><p><strong>That is the black box problem in its most practical form. </strong>A product can claim safety while clinicians cannot see which base model it uses, what fine-tuning changed, whether the adapted system was compared with the unmodified model, what model-level constraints exist, what clinical knowledge base it uses, what middleware monitors risk, what state it preserves, what adversarial testing it has passed, what benchmark results show, or what adverse-event monitoring occurs after deployment. <strong>That opacity is no longer tolerable in mental health AI.</strong></p><p>The <a href="https://www.fda.gov/media/189391/download">FDA&#8217;s Digital Health Advisory Committee materials</a> on generative AI-enabled digital mental health medical devices point in the same direction. The FDA executive summary discusses premarket performance evaluation, risk management, postmarket performance monitoring, human-in-the-loop considerations, training data characterisation, hallucination rates, error severity, stress testing, benchmarking, and ongoing monitoring.</p><p>The <a href="https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai">EU AI Act </a>also moves the field toward risk-based governance, human oversight, post-market monitoring, serious incident reporting, and disclosure duties. The European Commission states that the Act entered into force on 1 August 2024, with staged application dates and high-risk obligations continuing through later implementation timelines.</p><blockquote><p>Regulatory language is catching up. Clinical architecture still needs to catch up.</p></blockquote><p>One mechanism is not an architecture. A crisis banner is not a safety case. A disclaimer is not post-deployment surveillance. A reassuring response is not evidence that the system is safe. The integration task does not mainly require new technical invention. It requires professional discipline. Clinical leaders, regulators, safety teams, and engineering organisations need to agree on the baseline and hold the market to it.</p><div><hr></div><h4>The Operational Demand</h4><p>It&#8217;s time to open the safety black box. </p><p><strong>Builders and engineers of a mental health AI system must explicitly document their  safety architecture.</strong> At minimum, <strong>that means evidence across the five layers</strong>: model-level constraints, prompting and retrieval design, middleware risk monitoring, clinical state tracking, and adversarial robustness. <strong>It also means evaluation against VERA-MH or an equivalent benchmark</strong>, testing against the five-axis framework, and reporting what temporal evidence the system preserves under SCOPE-MH. <strong>A safety claim that does not specify its architecture is marketing, not engineering.</strong></p><p><strong>Clinical leaders and healthcare procurement teams need to add safety architecture requirements to every vendor review.</strong> </p><p>Ask which base model the system uses and what fine-tuning changed.</p><p>Ask whether the adapted model was tested against the unmodified base model and other relevant models.</p><p>Ask which Type I and Type II harms it has been tested against.</p><p>Ask over what duration and in which populations it was evaluated.</p><p>Ask what the system has learned to do. </p><p>Ask what knowledge base it uses. </p><p>Ask what middleware monitors risk. </p><p>Ask what state it preserves across sessions. </p><p>Ask for multi-turn adversarial testing. </p><p>Ask for crisis escalation protocols. </p><p>Ask who the named human is when the system crosses threshold risk. </p><p>Ask how the company monitors adverse events after deployment.</p><p>If the answers are not documented and verifiable, the vendor is not ready for clinical deployment.</p><p><strong>Investors should stop treating mental health AI safety as a risk that policy language can manage.</strong> It is a product risk, a clinical risk, and a governance risk. A company that cannot describe its safety architecture does not yet understand the product it is building.</p><p><strong>Regulators should require evidence that maps to the actual hazards of mental health AI.</strong> General-purpose safety documentation is insufficient, and mental health branding is not a substitute. Mental health AI needs evidence of clinical protocol fidelity, hallucination control, behavioural consistency, high-severity safety, demographic robustness, temporal evidence preservation, comparative performance, and post-deployment surveillance.</p><p>The comparator and evidentiary threshold should correspond to the product&#8217;s intended use, risk exposure, and public claims. A system claiming therapeutic equivalence must be compared with therapy. A system claiming support must still demonstrate that its support does not create clinically consequential harm over realistic use.</p><p>Mental health AI cannot remain a black box wrapped in therapeutic language.</p><p>If a system claims to support mental health, its safety architecture must be visible enough to inspect, strong enough to test, and accountable enough to fail in the open. Clinicians should not recommend what they cannot interrogate. Engineers should not deploy what they cannot validate. Investors should not fund what they cannot examine. Regulators should not accept claims where architecture should be.</p><p>Open the black box, or do not call it safe.</p><div><hr></div><p><em>Scott Wallace, PhD, has been building mental health software for more than 35 years, longer than smartphones have existed, longer than app stores, and longer than most of the teams now entering this space have been thinking about it. Trained in clinical psychology and neuropsychology, he received formal training in C, C++, and JavaScript to engineer some of North America&#8217;s earliest digital mental health platforms at a time when the web had fewer than 30,000 sites worldwide. He later completed early iOS certification to build mobile health applications as that platform emerged, and went on to lead the engineering development of NLP and NLG-based conversational systems before large language models entered the conversation.</em></p><p><em>That technical depth allowed him to work inside the architecture rather than alongside it, translating clinical requirements into system design and catching the places where engineering assumptions quietly displaced clinical ones.</em></p><p><em>Scott led the digital division of a major EAP provider through a successful exit, served as clinical lead for AI-based mental health technologies, and has produced hundreds of psychoeducational programmes used internationally. He now advises founders, clinicians, and investors building AI-enabled mental health systems.</em></p><p><em>His writing examines the future of mental healthcare, AI safety and governance, clinical risk, and what it actually takes to build mental health AI that holds up under real-world conditions.</em></p>]]></content:encoded></item></channel></rss>