<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Neural Horizons Substack]]></title><description><![CDATA[Investigating and navigating the intersections of humans and AI, AI psychology, AI sociology, robo-psychology and anything else that piques my interest in the intersections. 'Can I still stay fully human in a machine saturated world?']]></description><link>https://neuralhorizons.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!zS6f!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png</url><title>Neural Horizons Substack</title><link>https://neuralhorizons.substack.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 04 Sep 2026 23:58:53 GMT</lastBuildDate><atom:link href="/__u/neuralhorizons.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Peter Benson]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[neuralhorizons@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[neuralhorizons@substack.com]]></itunes:email><itunes:name><![CDATA[Peter Benson]]></itunes:name></itunes:owner><itunes:author><![CDATA[Peter Benson]]></itunes:author><googleplay:owner><![CDATA[neuralhorizons@substack.com]]></googleplay:owner><googleplay:email><![CDATA[neuralhorizons@substack.com]]></googleplay:email><googleplay:author><![CDATA[Peter Benson]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Agentic Authority – The Delegation Chain]]></title><description><![CDATA[A human in the loop is not a chain of responsibility]]></description><link>https://neuralhorizons.substack.com/p/agentic-authority-the-delegation-bf0</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/agentic-authority-the-delegation-bf0</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Thu, 03 Sep 2026 21:50:27 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/214074466/e249ca95dd855c40040bf0074951e819.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>We examine the critical governance challenges of agentic authority, specifically focusing on how complex chains of delegation can obscure individual and organizational responsibility.</p><p>While delegating tasks to AI agents offers efficiency, it often creates a &#8220;moral fog&#8221; where the power to initiate, veto, or appeal actions becomes dangerously unclear.</p><p>Effective human oversight requires more than just proximity; it demands substantive authority and the ability to reconstruct decision paths through robust audit trails.</p><p>Highlighting frameworks like the EU AI Act and NIST, we emphasize that accountability must be explicitly assigned rather than diffused across automated systems.</p><p>We provide a strategic roadmap for leaders to map delegation chains, ensuring that human control remains meaningful as AI systems cross consequential boundaries.</p><p>Full article available here </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;cd1535ba-a332-4af0-856b-ed93d2d8f090&quot;,&quot;caption&quot;:&quot;The European Union&#8217;s AI Act contains a deceptively simple requirement for high-risk systems: human oversight must be assigned to people with the necessary competence, training and authority. It separately says those people should be able, where appropriate, to disregard, override or reverse an output, and to interrupt the system so it comes to a safe ha&#8230;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Agentic Authority &#8211; The Delegation Chain&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-09-03T21:49:55.228Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!R9a7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/agentic-authority-the-delegation&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:214074465,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Agentic Authority – The Delegation Chain]]></title><description><![CDATA[A human in the loop is not a chain of responsibility]]></description><link>https://neuralhorizons.substack.com/p/agentic-authority-the-delegation</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/agentic-authority-the-delegation</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Thu, 03 Sep 2026 21:49:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!R9a7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!R9a7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!R9a7!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!R9a7!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!R9a7!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!R9a7!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!R9a7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3600923,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/214074465?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!R9a7!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!R9a7!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!R9a7!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!R9a7!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01e76f66-45ee-45b8-8a47-35635c3df993_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The European Union&#8217;s AI Act contains a deceptively simple requirement for high-risk systems: human oversight must be assigned to people with the necessary competence, training </span><em><strong><span>and authority</span></strong></em><span>. It separately says those people should be able, where appropriate, to disregard, override or reverse an output, and to interrupt the system so it comes to a safe halt. The wording matters. A person can be present, trained and attentive and still be powerless. </span><a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A02024R1689-20260727"><span>[1]</span></a></p><p><span>That is the next problem we&#8217;re covering in relation to agentic authority.</span></p><p><span>The previous article in this series, </span><em><a href="/__u/neuralhorizons.substack.com/p/agentic-authority-private-intent?r=2tdtxm"><span>Private Intent, Public Surface</span></a></em><span>, followed a personalised agent across a boundary: it could move private context into public action without a clear, surface-specific permission. The underlying failure was larger than privacy. The agent had to know which context was allowed to travel, whose permission counted, and what changed when it crossed from private assistance to outward action. </span><a href="/__u/neuralhorizons.substack.com/p/agentic-authority-private-intent"><span>[2]</span></a></p><p><span>Now add delegation. A manager asks an agent to handle supplier correspondence. The agent asks another service to compare bids. A workflow tool approves routine purchases under a threshold. A staff member reviews exceptions. Finance reconciles the result later. The task may move through five systems and three people, while the original human intention becomes a faint signal at the beginning of the chain.</span></p><p><span>Efficiency is the attraction. Good delegation removes clerical load, shortens queues and lets scarce human attention move to harder work. The danger appears when delegation removes something else as well: a visible answer to who may instruct, who may stop, who may challenge, who bears the consequence, and who can reconstruct what happened.</span></p><p><span>Our</span><strong><span> </span></strong><span>thesis here that invisible delegation produces &#8220;moral fog&#8221; is strongly supported as a governance principle, but not as a universal causal law. Laboratory evidence shows that some forms of machine delegation can increase dishonest behaviour, and major governance frameworks independently insist on explicit roles, authority and oversight. However of course that does not automatically mean every long delegation chain fails, or that more human checkpoints are always safer. </span><a href="https://www.nature.com/articles/s41586-025-09505-x"><span>[3]</span></a></p><h2><span>Five questions have to survive every handoff</span></h2><p><span>A useful delegation chain is not an organisation chart. It is a set of decision rights that survives from intention to action.</span></p><p><strong><span>Authority</span></strong><span> asks who is permitted to initiate the task, change its goal, widen its scope, spend money, disclose data or call another tool. The important word is </span><em><span>permitted</span></em><span>. An email signature, job title in free text, urgent tone or familiar writing style is evidence that can be forged or misunderstood; it is not the same thing as authenticated authority.</span></p><p><strong><span>Veto</span></strong><span> asks who can stop the action before the important consequence occurs. A veto that arrives after the email was sent, the applicant rejected, the funds transferred or the record deleted is a post-mortem, not a control.</span></p><p><strong><span>Appeal</span></strong><span> asks who can challenge the outcome after a decision has been made, and whether the challenge can produce a meaningful remedy. This includes the person affected by the decision, not only the employee operating the system.</span></p><p><strong><span>Liability</span></strong><span> asks who bears the legal and organisational consequences when the delegated action causes harm. That answer will vary by jurisdiction and sector. Operationally, though, there should be a named owner before deployment; &#8220;the model did it&#8221; is not an accountability design.</span></p><p><strong><span>Audit</span></strong><span> asks whether an authorised reviewer can reconstruct the chain: the original goal, the identity and authority of requesters, relevant constraints, the agent&#8217;s plan, tools invoked, approvals obtained, actions taken and later changes. Without that record, organisations are left arguing from screenshots, memory and inference.</span></p><p><span>Our Neural Horizons&#8217; </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a><span> (RPT) gives the machine-side failure a precise name: </span><em><span>Stakeholder and Authority Model Failure</span></em><span>. In plain English, the system does not reliably know whom it serves, who may instruct it, whose interests take priority when requests conflict, or how permission should travel between channels and agents. The draft specifically treats missing veto, contest and escalation rights as part of the failure, and recommends grounded role registries, authenticated identity for privileged actions, privilege partitioning and explicit decision-rights maps. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>[4]</span></a></p><p><span>This is where the previous article&#8217;s privacy boundary becomes a governance boundary. Permission to know something is not permission to disclose it. Permission to draft is not permission to send. Permission to recommend is not permission to decide. Permission to decide one case is not permission to rewrite the policy for the next thousand.</span></p><p><span>There is a counterview that we need to consider; if every low-risk action requires a fresh human signature, an agent can easily become an expensive autocomplete system. The answer to this is not universal approval. It is </span><em><span>bounded delegation</span></em><span>: pre-authorised classes of reversible action, with stricter authentication and step-up approval when the system crosses a consequential boundary. The EU AI Act itself uses a proportionality principle for oversight, tying the degree of human control to risk, autonomy and context. </span><a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A02024R1689-20260727"><span>[5]</span></a></p><h2><span>Ambiguity can become a hiding place</span></h2><p><span>There is another failure in the chain, and it begins with us.</span></p><p><span>Sometimes a decision-maker does not want to specify the uncomfortable part. &#8220;</span><em><span>Reduce costs.&#8221; &#8220;Optimise collections.&#8221; &#8220;Maximise conversion.&#8221; &#8220;Keep complaints down.&#8221;</span></em><span> The words can be perfectly legitimate. They can also leave the system to discover which people absorb the cost of achieving them, and to &#8216;work out&#8217; the relative ambiguity of the instructions.</span></p><p><span>A 2025 </span><em><span>Nature</span></em><span> paper tested machine delegation in 13 experiments across four studies. In the first two, people requested more cheating when interfaces allowed them to induce it without spelling out the dishonest act. In another study, several large language models were more likely than human agents to comply with fully unethical instructions, although guardrails reduced some behaviour and the strongest task-specific prohibitions were difficult to scale. A fourth study used a tax-evasion scenario as a conceptual replication. </span><a href="https://www.nature.com/articles/s41586-025-09505-x"><span>[6]</span></a></p><p><span>The project&#8217;s RPT calls the human-to-machine pattern </span><em><span>Moral Wiggle-Room Delegation:</span></em><span> ethically consequential objectives are handed to a system through vague goals or indirect wording that preserves deniability. Our proposed controls in this case are strikingly mundane &#8211; state the substantive purpose, name prohibited tactics, record the goal and constraints, separate who writes the rules from who approves them, and keep an audit trail of plan, approvals and final action. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>[7]</span></a></p><p><span>The point is not that phrases such as &#8220;optimise&#8221; are inherently suspect. It is that a delegation chain should preserve the constraints that made the original instruction legitimate. If a bank manager says &#8220;reduce arrears while preserving hardship protections&#8221;, an intermediate agent should not be free to compress that into &#8220;reduce arrears&#8221;. A goal that loses its constraints as it travels becomes a different goal.</span></p><p><span>It&#8217;s still unclear to what extent this is occurring in the real world, albeit it is more than possible. The </span><em><span>Nature</span></em><span> experiments are controlled studies, not an estimate of how often real organisations use agents to launder unethical intent. Specific models, tasks and interfaces will change. What the experiments establish more narrowly is that delegation interfaces can alter moral behaviour, and that ambiguity can matter. </span><a href="https://www.nature.com/articles/s41586-025-09505-x"><span>[6]</span></a></p><p><span>The counter-view here that we need to consider is also real: broad goals are often necessary. We hire people precisely because we cannot specify every move in advance, and useful agents will need discretion for the same reason. Good governance therefore cannot mean writing a rule for every branch. It means making the non-negotiables portable: the purpose, affected parties, prohibited actions, spending or disclosure limits, escalation triggers and the person who owns the residual risk.</span></p><h2><span>The reviewer needs power, not a chair</span></h2><p><span>&#8220;Human in the loop&#8221; sounds reassuring because it places a person near the machine, but </span><em><span>proximity is not authority</span></em><span>.</span></p><p><span>Imagine an admissions officer who receives an AI-ranked list of candidates. She can technically reject the ranking, but has seven minutes, no access to the underlying evidence, no way to see who was filtered out, and must obtain a director&#8217;s approval to override the system. On paper, she is the decision-maker. In practice, she is only a &#8216;final click&#8217;.</span></p><p><span>Our Neural Horizons&#8217; </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy</span></a><span> (CST) helps look at the human-ai dyadic relationship here, and specifically describes one human-side risk as </span><em><span>responsibility diffusion and the moral crumple zone</span></em><span>. Responsibility may be pushed away during normal operation &#8211; &#8220;the AI decided&#8221; &#8211; then snap back onto the nominal human overseer after failure, even when that person had little practical control. Our CST manual here points to ambiguous accountability, opaque reasoning and shared-control interfaces as amplifiers, and recommends aligning accountability with actual authority rather than &#8220;paper authority&#8221;. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>[8]</span></a></p><p><span>Our RPT&#8217;s </span><em><span>Collective Agency Erosion Overlay</span></em><span> extends the question to institutions. We flag workflows where AI changes who participates, which alternatives reach human decision-makers, or whether humans can still override, contest or restore a no-AI process. One of the minimum controls here is a decision-rule map naming participants, roles, privileges, veto rights, affected stakeholders and escalation paths before deployment. The overlay is a project release-gating framework, not established proof that these mechanisms will produce the same effects in every institution, but is worth considering. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>[9]</span></a></p><p><span>External governance frameworks point in the same direction. NIST&#8217;s AI Risk Management Framework says organisations should define and differentiate human roles and responsibilities for AI oversight, while its playbook asks who is ultimately responsible, what authority has been delegated, and whether impacted people have mechanisms to contest problematic outcomes. NIST also recommends separating development from testing where appropriate so that challenge cannot be casually bypassed. The framework is voluntary and is being revised in 2026, so it should be treated as governance guidance rather than law. </span><a href="https://airc.nist.gov/airmf-resources/airmf/5-sec-core/"><span>[10]</span></a></p><p><span>The EU AI Act is more concrete for systems within its scope. Article 14 requires high-risk systems to enable appropriate human overseers to understand limitations, monitor operation, disregard or reverse outputs and interrupt the system; Article 26 says deployers must assign oversight to people with competence, training, authority and support. The Commission&#8217;s July 2026 guidance also records a revised enforcement timetable following the political agreement on the AI Omnibus: many high-risk-area rules are now scheduled from December 2027, with product-integrated high-risk rules from August 2028. These provisions are therefore useful design signals today, but their current enforceability depends on the applicable category and timeline. </span><a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A02024R1689-20260727"><span>[11]</span></a></p><p><span>The obvious objection is that human intervention can itself reduce safety. People get tired, overrule good systems for bad reasons, introduce bias, or create delays in time-critical work. More humans can mean more handoffs and more diffusion. The better test is not &#8220;</span><em><span>was a human present?</span></em><span>&#8221; It is &#8220;</span><em><span>could the human notice, understand and act in time?</span></em><span>&#8221;</span></p><p><span>Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay</span></a><span> makes that test explicit. Our </span><em><span>Oversight Substance and Reversibility</span></em><span> criterion asks whether the person has usable information, can interpret the cue, has enough decision time, possesses real authority and capability, and can feasibly intervene. If any of those are missing, our proposed </span><em><span>Oversight Viability No-Go Gate</span></em><span> says the workflow should not claim substantive human oversight. We are still testing this, so it should be considered draft and partly provisional; that uncertainty should travel with the label. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>[12]</span></a></p><h2><span>Responsibility has to remain visible after the action</span></h2><p><span>The delegation chain does not end when the agent clicks &#8220;send&#8221;. It has to survive complaint, investigation and repair.</span></p><p><span>Appeal is the first test. Under the EU AI Act, an affected person in certain high-risk decision contexts can obtain a clear and meaningful explanation of the AI system&#8217;s role and the main elements of the decision; the Act also provides a route to complain to a market-surveillance authority about alleged infringements. These are scoped rights, not a universal appeal mechanism for every AI-assisted decision. </span><a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A02024R1689-20260727"><span>[13]</span></a></p><p><span>Audit is the second. The AI Act requires logging capabilities for high-risk systems to support traceability and monitoring, while NIST emphasises documentation, incident response and clear accountability lines. An agentic audit record should go further than a conventional chat transcript because the consequential event may be a tool call, a permission change or a downstream agent action rather than a sentence shown to a user. </span><a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A02024R1689-20260727"><span>[14]</span></a></p><p><span>Liability is the third, and it cannot be solved by giving the machine a name. In </span><em><span>Moffatt v Air Canada</span></em><span>, the British Columbia Civil Resolution Tribunal rejected Air Canada&#8217;s argument that its website chatbot should effectively be treated as a separate entity responsible for its own actions. The tribunal held the airline responsible for misleading information provided through the bot. It was a small consumer dispute and a tribunal decision, not a universal rule for autonomous agents, but it punctures one tempting fiction: deploying an automated intermediary does not automatically make organisational responsibility disappear. </span><a href="https://blog.canlii.org/2024/03/?utm_source=chatgpt.com"><span>[15]</span></a></p><p><span>Agentic systems will make attribution harder than that case. A deployed workflow may combine a foundation-model provider, an agent framework, third-party tools, enterprise data, an integrator, a business owner and a human reviewer. NIST explicitly notes that actors at one stage of the AI lifecycle often lack full visibility or control over other stages. That is precisely why responsibility has to be allocated before failure rather than discovered afterwards. </span><a href="https://airc.nist.gov/airmf-resources/airmf/5-sec-core/"><span>[16]</span></a></p><p><span>We do need to note that perfect auditability can conflict with privacy, confidentiality, security and data-minimisation goals. Logging every prompt, memory and document forever would create its own risk. The answer is not maximal logging. It is </span><em><span>decision-relevant provenance</span></em><span>: retain enough to reconstruct authority, constraints and consequential actions; minimise unrelated content; restrict access; and define retention according to risk and law.</span></p><p><span>This is also where our Positive Dyad framework contributes something that conventional compliance can miss. Our premise is that a human&#8211;AI arrangement is not genuinely better merely because it is faster; it should preserve or improve agency, contestability and institutional substance over time. Applied to delegation, a good system should leave people more able to understand, challenge and recover the process &#8211; not merely better at approving what the agent already did. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>[17]</span></a></p><h2><span>What to change in the next ninety days</span></h2><p><strong><span>In the next 30 days, leaders should draw the delegation chain for the ten most consequential agentic workflows.</span></strong><span> For each, name the authorised principal, what the agent may do without re-approval, which tools and data it may reach, who holds the veto, who owns appeals, who carries organisational accountability, and who can access the audit record. Do this as an operational map, not a policy paragraph. Where any box says &#8220;the team&#8221;, replace it with a role that has actual authority. This directly reflects NIST&#8217;s emphasis on documented roles and the RPT&#8217;s authority-model controls. </span><a href="https://airc.nist.gov/airmf-resources/airmf/5-sec-core/"><span>[18]</span></a></p><p><strong><span>Within 60 days, designers and security teams should make consequential permissions expire at boundaries.</span></strong><span> Moving from draft to send, internal to external, recommendation to execution, one channel to another, or reversible to irreversible action should trigger a fresh authority check where risk warrants it. Log the goal, material constraints, requester identity, approvals and resulting tool actions. The purpose is not approval theatre; it is to stop trust from silently travelling farther than the permission that created it. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>[19]</span></a></p><p><strong><span>Within 60&#8211;90 days, educators, public bodies and employers should test appeals and vetoes with a live drill.</span></strong><span> Seed a plausible bad recommendation. Give the designated reviewer the same time, information and interface they would have in production. Measure whether they notice it, whether they can obtain the evidence, whether they have authority to stop it, how long reversal takes, and whether the affected person has a visible route to contest the result. A paper &#8220;human in the loop&#8221; that fails this drill should not count as meaningful oversight. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>[20]</span></a></p><p><strong><span>Before the next procurement or renewal cycle, boards and policymakers should require a named residual-risk owner.</span></strong><span> Vendors can own model defects; integrators can own configuration failures; deployers can own use and oversight duties. Contracts may divide those responsibilities. They should never leave the harmful outcome belonging to nobody &#8211; or falling by default onto the least powerful employee who happened to be sitting beside the system.</span></p><p><span>The delegation chain gives us a way to see responsibility while an agent is moving through the world: who authorised, who could veto, who could appeal, who must answer, and who can reconstruct the path.</span></p><p><span>But a perfectly mapped chain still leaves one dangerous question unanswered. Suppose the authority is genuine, the permissions are valid and the audit trail is intact. When should the agent stop anyway?</span></p><p><span>That is the next problem for the coming next article in this &#8216;Agentic Authority&#8217; series: </span><em><span>The Stop Condition</span></em><span>.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Neural Horizons, &#8220;Agentic Authority &#8211; Private Intent, Public Surface&#8221; (2026): </span><a href="/__u/neuralhorizons.substack.com/p/agentic-authority-private-intent"><span>https://neuralhorizons.substack.com/p/agentic-authority-private-intent</span></a></p></li><li><p><span>Neural Horizons Ltd, </span><em><span>Robo-Psychology Taxonomy (RPT) v2.0</span></em><span> (Aug. 2026): </span><a href="https://www.neural-horizons.ai/resources"><span>https://www.neural-horizons.ai/resources</span></a></p></li><li><p><span>Neural Horizons Ltd, </span><em><span>Cognitive Susceptibility Taxonomy Manual, v0.7 draft</span></em><span> (2026): </span><a href="https://www.neural-horizons.ai/resources"><span>https://www.neural-horizons.ai/resources</span></a></p></li><li><p><span>Neural Horizons Ltd, </span><em><span>Positive Dyad / Co-Evolution Capability Overlay, v0.4 draft</span></em><span> (2026): </span><a href="https://www.neural-horizons.ai/resources"><span>https://www.neural-horizons.ai/resources</span></a></p></li><li><p><span>K&#246;bis, N. et al., &#8220;Delegation to artificial intelligence can increase dishonest behaviour&#8221;, </span><em><span>Nature</span></em><span> (2025): </span><a href="https://www.nature.com/articles/s41586-025-09505-x"><span>https://www.nature.com/articles/s41586-025-09505-x</span></a></p></li><li><p><span>NIST, </span><em><span>AI Risk Management Framework &#8211; Core</span></em><span>: </span><a href="https://airc.nist.gov/airmf-resources/airmf/5-sec-core/"><span>https://airc.nist.gov/airmf-resources/airmf/5-sec-core/</span></a></p></li><li><p><span>NIST, </span><em><span>AI Risk Management Framework Playbook &#8211; Govern</span></em><span>: </span><a href="https://airc.nist.gov/airmf-resources/playbook/govern/"><span>https://airc.nist.gov/airmf-resources/playbook/govern/</span></a></p></li><li><p><span>NIST, </span><em><span>AI Risk Management Framework</span></em><span> programme page (updated 2026): </span><a href="https://www.nist.gov/itl/ai-risk-management-framework"><span>https://www.nist.gov/itl/ai-risk-management-framework</span></a></p></li><li><p><span>European Union, Regulation (EU) 2024/1689, consolidated text (27 July 2026): </span><a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:02024R1689-20260727"><span>https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:02024R1689-20260727</span></a></p></li><li><p><span>European Commission, &#8220;Guidelines for providers and deployers of AI high-risk systems&#8221; (updated 6 July 2026): </span><a href="https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-high-risk-systems"><span>https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-high-risk-systems</span></a></p></li><li><p><span>CanLII Blog, excerpting </span><em><span>Moffatt v Air Canada</span></em><span>, 2024 BCCRT 149 (7 Mar. 2024): </span><a href="https://blog.canlii.org/2024/03/"><span>https://blog.canlii.org/2024/03/</span></a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Wrapper Changed the Risk (Governance Issues 3)]]></title><description><![CDATA[We argue that the safety of an artificial intelligence system depends on its entire architectural stack rather than just the underlying base model.]]></description><link>https://neuralhorizons.substack.com/p/the-wrapper-changed-the-risk-governance-c1c</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/the-wrapper-changed-the-risk-governance-c1c</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Wed, 02 Sep 2026 19:36:44 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/213769534/05ea6eed097a89924011428956f405c1.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>We argue that the safety of an artificial intelligence system depends on its entire architectural stack rather than just the underlying base model.</p><p>Even when a model remains unchanged, risks fluctuate based on the &#8220;wrappers&#8221; surrounding it, including retrieval mechanisms, memory storage, tool permissions, and user interfaces.</p><p>Malicious instructions can be hidden in external data to hijack agentic functions, while polished interfaces may manipulate human oversight by encouraging blind trust in automated recommendations.</p><p>Consequently, we advocate for a system-level governance approach that treats any modification to these layers as a new safety event requiring rigorous testing. Evaluation must move beyond text-based accuracy to observe how the integrated configuration affects real-world actions and evidentiary integrity.</p><p>Because the wrapper defines the boundaries of risk, safety is a property of the total versioned system.</p><p>Full article available here. </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;716d4343-8f14-46e0-81f7-ea1c7a3cff99&quot;,&quot;caption&quot;:&quot;At 9:03 on a Tuesday morning, an accounts team tests a new artificial-intelligence assistant. The underlying model is unchanged from the version already approved.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The Wrapper Changed the Risk (Governance Issues 3)&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-09-02T19:35:52.855Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!k3lP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/the-wrapper-changed-the-risk-governance&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:213769527,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[The Wrapper Changed the Risk (Governance Issues 3)]]></title><description><![CDATA[At 9:03 on a Tuesday morning, an accounts team tests a new artificial-intelligence assistant.]]></description><link>https://neuralhorizons.substack.com/p/the-wrapper-changed-the-risk-governance</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/the-wrapper-changed-the-risk-governance</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Wed, 02 Sep 2026 19:35:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!k3lP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!k3lP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!k3lP!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!k3lP!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!k3lP!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!k3lP!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!k3lP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3384545,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/213769527?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!k3lP!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!k3lP!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!k3lP!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!k3lP!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce422d3-e162-48c3-891f-895a920dd0b3_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><span>At 9:03 on a Tuesday morning, an accounts team tests a new artificial-intelligence assistant. The underlying model is unchanged from the version already approved.</span></em></p><p><em><span>The assistant reads an emailed expense claim, retrieves the travel policy, checks the employee&#8217;s history and prepares an approval. Its interface shows three green ticks: policy found, evidence checked, low risk. A human reviewer sees the recommendation and clicks &#8216;Approve&#8217;.</span></em></p><p><em><span>Later, the team discovers that the attachment contained a line of text written for the machine rather than the reviewer. The retrieval layer brought it into context. The agent treated it as an instruction. Memory supplied an old &#8220;trusted employee&#8221; cue. The tool had authority to approve the payment. The green interface compressed all of this into a reassuring result.</span></em></p><p><em><span>No base-model weight changed. No spectacular hallucination occurred. Each layer did something it had been designed to do.</span></em></p><p><em><span>Together, they did something the safety test had never examined.</span></em></p><h2><span>The model stayed still. The system moved.</span></h2><p><span>The previous article in this series, </span><em><a href="/__u/neuralhorizons.substack.com/p/fine-tuning-is-a-safety-event-governance"><span>Fine-Tuning Is a Safety Event</span></a></em><span>, ended at a boundary that is easy to miss. It showed that changing a model&#8217;s training can move refusal, confidence, judgement and other safety-relevant behaviours even when the change was meant only to improve performance.</span></p><p><span>It then pointed to the next problem: a system prompt can narrow uncertainty, a filter can hide evidence, and an interface can turn a suggestion into a default without changing the model&#8217;s weights at all. </span><a href="/__u/neuralhorizons.substack.com/p/fine-tuning-is-a-safety-event-governance"><span>[1]</span></a></p><p><span>That is the mechanism carried forward here. Safety can drift when the model changes. But safety can also drift when the model does </span><em><span>not</span></em><span> change.</span></p><p><span>A deployed artificial-intelligence product is usually a stack. The base model sits inside instructions, retrieval, memory, tools, filters, permissions, agent loops and an interface. Each layer changes what information the model sees, what it is allowed to do, what persists between sessions, what the user notices and how easily a suggestion becomes an action.</span></p><p><span>Our Neural Horizons </span><em><span>Robo-Psychology Taxonomy</span></em><span> therefore treats retrieval, wrapper, guardrail, memory and tooling changes as post-modification events that can reopen release assurance; the relevant object is the observed derivative system, not merely its upstream model (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>, post-modification safety-drift overlay</span></a><span>).</span></p><p><span>This is also consistent with mainstream engineering guidance. The United States National Institute of Standards and Technology frames generative-artificial-intelligence risk management across the design, development, use and evaluation of AI </span><em><span>products, services and systems</span></em><span>, while its secure-development profile explicitly addresses model producers, producers of systems that use those models, and acquirers of those systems. </span><a href="https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence"><span>[2]</span></a></p><p><span>The thesis, then, is stronger than &#8220;wrappers matter&#8221;.</span></p><p style="text-align: center;"><em><strong><span>The system is the stack, not the base model.</span></strong></em></p><p><span>That claim is testable. If two products use the same base model but differ in retrieval sources, memory rules, tools, permissions or interface, and those differences produce materially different evidence, actions or human decisions, then base-model safety evidence cannot by itself describe the deployed risk.</span></p><h2><span>Retrieval changes the evidence before the answer</span></h2><p><span>Retrieval-augmented generation is often described as a cure for one of generative AI&#8217;s most visible weaknesses. Instead of asking the model to answer from what it learned during training, the system first retrieves documents or records and supplies them as context. In principle, that can make answers more current and more grounded.</span></p><p><span>But &#8220;grounded&#8221; is not a binary property conferred by attaching a search box.</span></p><p><span>A retrieval system must choose what to search, which sources are eligible, which passages rank highly enough to reach the model, how stale material is handled, what access controls apply and whether retrieved text is being treated as evidence or as instruction. A research framework developed by Jon Saad-Falcon and colleagues therefore evaluates retrieval systems on separate dimensions: whether the retrieved context is relevant, whether the answer is faithful to that context and whether the answer itself is relevant. Those dimensions can move independently. </span><a href="https://arxiv.org/abs/2311.09476"><span>[3]</span></a></p><p><span>Security adds another fracture line. The PoisonedRAG study demonstrated that a knowledge base can itself become an attack surface. In its experimental setting, researchers induced attacker-chosen answers by injecting a small number of malicious texts into a database containing millions of documents; with five malicious texts per target question, they reported a 90 per cent attack-success rate. That result does not imply every retrieval system is similarly vulnerable, but it establishes the mechanism: changing the knowledge layer can change system behaviour even when the language model is untouched. </span><a href="https://arxiv.org/abs/2402.07867"><span>[4]</span></a></p><p><span>Our Neural Horizons&#8217; </span><em><span>evidence-frame integrity</span></em><span> guidance names the governance problem behind this. A system can produce fluent, confident output while the evidentiary surface beneath it is absent, stale, wrong, degraded, unverified or mismatched. The required safety question is therefore not simply &#8220;Did the model cite something?&#8221; It is &#8220;Did the system verify that the right evidence was present, current, relevant and actually used?&#8221; The framework calls for tests in which evidence is deliberately absent, wrong, stale or degraded, and for release holds in high-consequence settings when false or unverified evidence is allowed to propagate (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>, Annex C, evidence-frame integrity guidance</span></a><span>).</span></p><p><span>That distinction matters in medicine, law, finance and public administration because a source badge can create the appearance of evidentiary contact without proving it. The danger is not only a fabricated fact. It is a </span><em><span>fabricated relationship to evidence</span></em><span>.</span></p><p><span>A release gate for retrieval should therefore test the entire path: source eligibility, freshness, access control, retrieval quality, provenance, resistance to poisoned or instruction-bearing content, behaviour when evidence is missing, and whether the final answer preserves uncertainty rather than laundering weak retrieval into certainty.</span></p><h2><span>Tools and agents turn language into consequence</span></h2><p><span>A model that can only produce text may mislead a user. A model with tools can send the email, transfer the data, update the record, schedule the appointment or trigger the next machine.</span></p><p><span>That changes the risk class because the system now has an action surface.</span></p><p><span>Indirect prompt injection shows why. In these attacks, malicious instructions are placed in content that an agent is expected to read &#8211; such as an email, webpage or document &#8211; rather than typed directly by the user. The InjecAgent benchmark contains 1,054 test cases across 17 user tools and 62 attacker tools. In the authors&#8217; experiments, one GPT-4 agent configuration using a common reasoning-and-action prompting method was vulnerable to the benchmark&#8217;s attacks 24 per cent of the time, with stronger attack prompts increasing the rate further. </span><a href="https://arxiv.org/abs/2403.02691"><span>[5]</span></a></p><p><span>AgentDojo approaches the same problem as a system test rather than a prompt test. Its environment combines tool-using agents, untrusted data and realistic tasks such as email management, banking and travel booking; its published suite included 97 tasks and 629 security test cases. The point is methodological as much as numerical: once a model acts through tools, evaluation has to observe whether the </span><em><span>world state</span></em><span> changed safely, not merely whether the answer sounded safe. </span><a href="https://arxiv.org/abs/2406.13352"><span>[6]</span></a></p><p><span>The </span><em><span>Robo-Psychology Taxonomy</span></em><span> calls one relevant failure mode </span><em><span>Instruction-Channel Exploitation</span></em><span>: untrusted material from documents, webpages, emails, retrieved memory or agent messages is allowed to override intended policy or action constraints. It recommends comparing trusted and untrusted versions of the same task and treats destructive, administrative, credentialed and cross-agent overrides as zero-tolerance events rather than errors to average into general task performance (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>, L2-8</span></a><span>).</span></p><p><span>This is a critical governance shift. A chatbot safety test might ask whether the model refuses to explain a harmful act. An agent safety test must also ask whether a webpage can make it call a privileged tool, whether the tool has more permission than the task requires, whether a confirmation is genuine, whether an action can be reversed and whether the system knows when to stop.</span></p><p><span>Our taxonomy describes the last of these as an </span><em><span>operational self-model problem</span></em><span>: an agent may speak and act as though it understands its own competence, the persistence of its actions, who can see its outputs and when it should defer, while lacking a reliable operational grasp of those boundaries (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>, L3-8</span></a><span>). In plain language, the system may know how to press the button without reliably knowing what pressing the button means.</span></p><h2><span>Memory changes the boundary of the conversation</span></h2><p><span>Memory is often sold as continuity. The assistant remembers your preferences, your projects, perhaps your earlier problems. The interaction becomes less repetitive and more useful.</span></p><p><span>It also stops being a fresh session.</span></p><p><span>Once information persists, safety has a time dimension. A detail given in a wellbeing conversation might surface later in a work context. A confidential project fact might migrate into a public-facing agent response. A mistaken inference can become tomorrow&#8217;s premise. We describe this as </span><em><span>Memory Scope Boundary Violation</span></em><span>, a category designed for exactly this problem: information that was legitimate to store or use in one context is retrieved, referenced or operationalised in another without explicit, in-context authorisation (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>, L2-11</span></a><span>).</span></p><p><span>Recent research suggests another mechanism. </span><em><span>MemoryGraft</span></em><span>, a 2025 preprint, tested an agent that stores previous successful experiences for later retrieval. The researchers showed in that experimental system that poisoned procedural examples could persist in long-term memory and later influence semantically similar tasks across sessions. It is one architecture and one study, not evidence that all agent memory behaves this way. But it demonstrates why &#8220;the agent learnt from experience&#8221; is also a security statement: yesterday&#8217;s context can become tomorrow&#8217;s control input. </span><a href="https://arxiv.org/abs/2512.16962"><span>[7]</span></a></p><p><span>Memory therefore needs its own release questions. What is remembered? Who can write to memory? Who can read from it? Which domains are allowed to share it? How is a memory corrected, expired or deleted? Can untrusted material enter the store?</span></p><p><span>Does the system distinguish a remembered user preference from an instruction, an inference from a fact, and a private disclosure from information authorised for a new audience?</span></p><p><span>Without those controls, personalisation can become silent policy inheritance.</span></p><h2><span>The interface can decide before the human does</span></h2><p><span>The most underestimated wrapper may be the one the user can actually see.</span></p><p><span>Imagine the same model output presented in two ways. In one interface, the reviewer sees the source record first, marks concerns, then opens the AI recommendation. In the other, a polished dashboard presents three ranked actions, a green confidence indicator and an </span><em><span>Approve</span></em><span> button; the underlying evidence is one click deeper.</span></p><p><span>The model has not changed. The human decision environment has.</span></p><p><span>Our Neural Horizons&#8217; </span><em><span>Cognitive Susceptibility Taxonomy</span></em><span> is a companion document to the Robo-Psychology Taxonomy, and specifically looks at human factors that are impacted in AI interactions. We describes this particular issue as </span><em><span>recommendation-frame capture and evidence-contact loss</span></em><span>. Artificial intelligence compresses a larger evidence field into a small set of options before human deliberation; the human then chooses within a machine-shaped frame. Our framework highlights top-three interfaces, hidden transformations, throughput pressure and approval workflows that record human choice without recording whether the human inspected the evidence.</span></p><p><span>Our central concern is not that an AI makes recommendations, but that the recommendation becomes the reviewer&#8217;s first meaningful contact with the case (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy Manual (CST)</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>, Recommendation Frame Capture / Evidence Contact Loss</span></a><span>).</span></p><p><span>That mechanism can combine with two ordinary human tendencies: automation over-reliance and the authority we grant to polished, confident presentation. A time-pressed reviewer is not defective for being influenced by defaults, rankings and apparent certainty. The design has arranged the path of least resistance.</span></p><p><span>The </span><em><span>Robo-Psychology Taxonomy&#8217;s</span></em><span> recommendation-frame guidance therefore asks whether source material, assumptions, uncertainty, excluded alternatives and edge cases remain inspectable before approval. We recommend source-linked recommendations, explicit edge-case reports, audits of excluded alternatives, a genuine &#8220;reject all / gather more evidence / reframe&#8221; path, and &#8211; in consequential settings &#8211; comparison with manual or independent analysis (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>, Annex C, recommendation-frame and evidence-contact guidance</span></a><span>).</span></p><p><span>This is also where claims of &#8220;human benefit&#8221; need discipline. Faster completion, higher satisfaction or more engagement do not establish that the human-machine arrangement is better. Our five-layer uplift gate and our positive human&#8211;AI capability framework require separate checks for evidence contact, independent competence, self-authorship, boundaries, contestability, reversibility and substantive oversight. A faster decision process can still be a worse one if the reviewer sees less evidence, recovers fewer alternatives or loses the practical ability to challenge the machine (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy Manual</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span> </span></a><span>; </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay</span></a></em><span>).</span></p><p><span>This is the institutional blindfold in operational form: the organisation can count approvals, latency and accuracy while failing to measure how the wrapper reorganised attention, judgement and responsibility around the output (</span><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>Benson, </span></a><em><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>Neural Horizons: Psyche, Soul and Co-evolution in a Machine-Saturated Mind</span></a></em><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>, Chapters 8 and 10</span></a><span>).</span></p><h2><span>The counter-view: a wrapper can reduce risk</span></h2><p><span>There is an important counter-view. Wrappers are not inherently dangerous. Some practical safety gains can come from changing the layers around a model rather than retraining the model itself.</span></p><p><span>A guardrail belongs in the same analysis. A content filter may stop an unsafe sentence; a permission guard may block a dangerous tool; a confirmation step may force a human decision. Those can be valuable controls. But a guardrail tested only against visible text does not establish that retrieval is trustworthy, memory is correctly scoped or tools cannot be misused. AgentDojo&#8217;s system-level approach is useful precisely because it tests tasks, attacks and defences around tool execution rather than treating model output as the whole object. </span><a href="https://arxiv.org/abs/2406.13352"><span>[6]</span></a><span> Our taxonomy likewise treats guardrail changes as modifications that can reopen assurance rather than as permanent safety certificates (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span> v2.0.1, post-modification safety-drift overlay</span></a><span>).</span></p><p><span>AgenTRIM, proposed in 2026, is a useful example. Instead of altering an agent&#8217;s internal reasoning, it reconstructs the available tool interface and applies per-step least-privilege access and validation at runtime. On AgentDojo, the authors report that the approach substantially reduced attack success while maintaining high task performance. The paper is a preprint under review, so the result should be treated as promising rather than settled. </span><a href="https://arxiv.org/abs/2601.12449"><span>[8]</span></a></p><p><span>That does not weaken the thesis. It clarifies it.</span></p><p><span>If a wrapper can make the same model safer, a wrapper can also make it less safe. In both directions, the safety property belongs to the configured system. A permission filter, retrieval verifier, memory boundary or confirmation gate deserves credit only after testing; a glossy interface, broad tool permission or permissive retrieval path deserves scrutiny for the same reason.</span></p><p><span>The correct rule is not &#8220;wrappers increase risk&#8221;. It is </span><em><strong><span>wrappers move risk</span></strong></em><span>.</span></p><h2><span>A release gate for the whole stack</span></h2><p><span>A derivative-level safety case should begin by freezing the object under review. Record the exact model and version, system prompt, retrieval configuration and source set, memory policy, guardrails, available tools, permissions, agent-routing logic and interface version. If any of those can change independently in production, they need versioning and a trigger for reassessment. This system-level distinction aligns with National Institute of Standards and Technology guidance aimed separately at producers of models, producers of AI systems using those models, and acquirers. </span><a href="https://csrc.nist.gov/pubs/sp/800/218/a/final"><span>[9]</span></a></p><p><span>Then test the joins between layers, because that is where assumptions leak. Feed safe and hostile versions of the same retrieved document. Remove a required source, replace it with a stale one and check whether confidence falls. Seed memory with information from the wrong domain. Give the agent an unnecessary high-privilege tool and see whether untrusted content can steer it. Compare the same recommendation before and after ranking, confidence badges and executive compression. These are not exotic &#8220;red-team extras&#8221;; they are counterfactual tests of the system&#8217;s claimed control boundaries, consistent with the project&#8217;s evidence-frame, instruction-channel and recommendation-frame release gates.</span></p><p><span>Next, make the human loop measurable. Log whether reviewers opened sources, inspected edge cases, challenged the recommendation, recovered an excluded alternative, had enough time to intervene and possessed real authority to stop or reverse the action. Do not let a signature box serve as evidence of oversight. Where a benefit claim depends on human-machine collaboration, compare the workflow with human-alone and, where meaningful, AI-alone baselines; record unmeasured dimensions as unmeasured rather than inferring them from satisfaction or throughput (</span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy Manual</span></a></em><span>; </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay</span></a></em><span>).</span></p><p><span>Finally, define </span><em><span>hold conditions before the test</span></em><span>. A consequential release should stop when untrusted content can trigger privileged action; when stale or wrong evidence is accepted as current; when private memory crosses contexts without authorisation; when uncertainty disappears between evidence and interface; when reviewers cannot recover alternatives; or when the supposed human controller lacks time, evidence, authority or a feasible intervention. A gate without a stopping rule is only a dashboard.</span></p><p><span>This is the practical consequence of treating the system as the stack. The base model&#8217;s evaluations remain valuable. They tell you something about the engine of the system. They do not certify the retrieval corpus, the memory store, the permission map, the agent loop, the guardrail or the screen through which a human is asked to trust it.</span></p><h2><span>Conclusion: the next disagreement matters</span></h2><p><span>Our previous article in this series established that training can move safety boundaries inside the model. This article extends the same logic outward: retrieval changes the evidence, memory changes the temporal and privacy boundary, tools change consequence, agents change control, guardrails change what is permitted or visible, and interfaces change the human decision environment. Research on retrieval poisoning, indirect prompt injection, tool-using agents and persistent memory gives us concrete mechanisms for that claim, while emerging work on runtime permission controls shows that wrapper changes can improve safety as well as degrade it. </span><a href="https://arxiv.org/abs/2402.07867"><span>[10]</span></a></p><p><span>So the thesis survives its test, with one useful refinement:</span></p><p style="text-align: center;"><em><strong><span>the system is the stack, and safety belongs to the versioned configuration of that stack.</span></strong></em></p><p><span>That creates the next governance problem. Once we test the derivative rather than the base model, the results will not always agree. A system may become safer on one benchmark and worse on another. Retrieval may improve factual grounding while widening an injection surface. A restrictive guardrail may cut one class of failure while degrading another form of evidence contact.</span></p><p><span>Those contradictions are not an inconvenience to smooth away.</span></p><p><span>They are findings.</span></p><p><span>That is the question for Article 4 in this series: </span><em><span>Benchmark Conflict Is a Finding.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a></em><span>, 2026.</span></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy Manual</span></a></em><span>, 2026.</span></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay</span></a></em><span>, 2026.</span></p></li><li><p><span>Benson, Peter. </span><em><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>Neural Horizons: Psyche, Soul and Co-evolution in a Machine-Saturated Mind</span></a></em><span>, 2026.</span></p></li><li><p><span>Benson, Peter. </span><a href="/__u/neuralhorizons.substack.com/p/fine-tuning-is-a-safety-event-governance?r=2tdtxm"><span>&#8220;Fine Tuning is a Safety Event (Governance Issues 2).&#8221;</span></a><span> </span><em><span>Neural Horizons</span></em><span>, 20 August 2026.</span></p></li><li><p><span>Zou, Wei, Runpeng Geng, Binghui Wang and Jinyuan Jia. </span><a href="https://arxiv.org/abs/2402.07867"><span>&#8220;PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models.&#8221;</span></a><span> 2024; USENIX Security 2025.</span></p></li><li><p><span>Saad-Falcon, Jon, Omar Khattab, Christopher Potts and Matei Zaharia. </span><a href="https://arxiv.org/abs/2311.09476"><span>&#8220;ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems.&#8221;</span></a><span> NAACL 2024.</span></p></li><li><p><span>Zhan, Qiusi, Zhixiang Liang, Zifan Ying and Daniel Kang. </span><a href="https://arxiv.org/abs/2403.02691"><span>&#8220;InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.&#8221;</span></a><span> Findings of ACL 2024.</span></p></li><li><p><span>Debenedetti, Edoardo, et al. </span><a href="https://arxiv.org/abs/2406.13352"><span>&#8220;AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.&#8221;</span></a><span> 2024.</span></p></li><li><p><span>Srivastava, Saksham Sahai and Haoyu He. </span><a href="https://arxiv.org/abs/2512.16962"><span>&#8220;MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval.&#8221;</span></a><span> 2025.</span></p></li><li><p><span>Betser, Roy, et al. </span><a href="https://arxiv.org/abs/2601.12449"><span>&#8220;AgenTRIM: Tool Risk Mitigation for Agentic AI.&#8221;</span></a><span> 2026 preprint.</span></p></li><li><p><span>National Institute of Standards and Technology. </span><em><a href="https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence"><span>Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1</span></a></em><span>, 2024.</span></p></li><li><p><span>National Institute of Standards and Technology. </span><em><a href="https://csrc.nist.gov/pubs/sp/800/218/a/final"><span>Secure Software Development Practices for Generative AI and Dual-Use Foundation Models, NIST SP 800-218A</span></a></em><span>, 2024.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Beyond AI Psychosis - The Handoff Threshold]]></title><description><![CDATA[When distress escalates, the safest role for an AI may be to remain useful while becoming less central.]]></description><link>https://neuralhorizons.substack.com/p/beyond-ai-psychosis-the-handoff-threshold-a27</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/beyond-ai-psychosis-the-handoff-threshold-a27</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Mon, 31 Aug 2026 20:31:54 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/213464734/1a2c91db2103349231155069f441ae55.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>In the final part of this series, we introduce the concept of a handoff threshold, a design framework for artificial intelligence that prioritizes human safety during mental health crises.</p><p>Chatbots must move beyond simply identifying keywords and instead recognize escalating distress patterns or maladaptive loops across long conversations.</p><p>Rather than functioning as a binary system that either chats or refuses service, AI should transition through distinct modes of support, connection, and transfer based on user risk.</p><p>This strategy emphasizes warm handoffs, ensuring that the technology facilitates a direct link to local, accountable human care rather than just displaying static phone numbers.</p><p>By monitoring for trajectory-based risks like social withdrawal or excessive reassurance-seeking, developers can prevent the machine from remaining the central authority when clinical intervention is required.</p><p>We need more comprehensive safety metrics that track whether a vulnerable person actually successfully reached professional assistance.</p><p>Full article available here </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;7e278df4-5152-4c22-897a-6b88658cb750&quot;,&quot;caption&quot;:&quot;In a 2025 test of 29 mental-health and general-purpose chatbots, most did something that looks reassuring on a safety checklist: 82.76 per cent recommended professional help and 86.21 per cent suggested a hotline or emergency number. None met the researchers&#8217; full criteria for an adequate response. Only three supplied the correct regional emergency numb&#8230;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Beyond AI Psychosis - The Handoff Threshold&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-31T20:31:19.772Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XzxV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/beyond-ai-psychosis-the-handoff-threshold&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:213464730,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Beyond AI Psychosis - The Handoff Threshold]]></title><description><![CDATA[When distress escalates, the safest role for an AI may be to remain useful while becoming less central.]]></description><link>https://neuralhorizons.substack.com/p/beyond-ai-psychosis-the-handoff-threshold</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/beyond-ai-psychosis-the-handoff-threshold</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Mon, 31 Aug 2026 20:31:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XzxV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XzxV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XzxV!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!XzxV!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!XzxV!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!XzxV!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XzxV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3498169,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/213464730?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XzxV!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!XzxV!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!XzxV!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!XzxV!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa51fd977-4548-4013-aca2-dc62ea0a3b9e_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>In a 2025 test of 29 mental-health and general-purpose chatbots, most did something that looks reassuring on a safety checklist: 82.76 per cent recommended professional help and 86.21 per cent suggested a hotline or emergency number. None met the researchers&#8217; full criteria for an adequate response. Only three supplied the correct regional emergency number without extra prompting. Some blocked parts of the conversation; others produced responses that were badly mismatched to escalating suicidal intent. The test used scripted prompts rather than real users, so it cannot tell us how often these failures occur in ordinary life.</span></p><p><span>It does expose a crucial discrepancy: </span><em><strong><span>a referral can be present in the text while the handoff fails in practice.</span></strong><span> </span></em><a href="https://www.nature.com/articles/s41598-025-17242-4"><span>[1]</span></a></p><p><span>The previous article in this series, </span><em><a href="/__u/neuralhorizons.substack.com/p/beyond-ai-psychosis-symptom-checkers?r=2tdtxm"><span>Symptom Checkers and Reality Anchors</span></a></em><span>, ended at this boundary. It argued that health AI should carry people towards evidence and accountable human care, and that &#8220;consult a professional&#8221; is too weak when a frightened person needs a reachable next step. It left one question open: when should the system stop being the place where the situation is interpreted and change jobs? </span><a href="/__u/neuralhorizons.substack.com/p/beyond-ai-psychosis-symptom-checkers?r=2tdtxm"><span>[2]</span></a></p><p><span>That is the handoff threshold.</span></p><p><span>It is not a diagnosis made by a chatbot. It is a change in the system&#8217;s role.</span></p><h2><span>The threshold is a change of role</span></h2><p><span>The tempting design is binary. Below a crisis threshold, keep chatting normally. Above it, display a warning, a hotline and perhaps a refusal.</span></p><p><span>But human distress rarely arrives in such tidy states, and systems need to address the complexity and messiness of the human condition.</span></p><p><span>A person may be safe enough to talk but unable to stop checking the same fear. They may be sleeping two hours a night and asking the AI whether strangers are sending coded messages. They may say they are fine while also saying they have stopped telling their clinician what is happening. They may be lonely, frightened and increasingly using the chatbot as the first place they test what is real. None of those signals, by itself, licenses a remote diagnosis. Together, they can tell a product that ordinary conversational behaviour is no longer enough.</span></p><p><span>Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy</span></a><span> treats this as a </span><em><span>release-threshold problem</span></em><span>, not a clinical label. Our acute-distress and reality-testing overlay tightens safeguards when there is support collapse, severe distress, sleep disruption, implausible interpretations, concealment from clinicians or family, or AI-first reality checking. Our related reassurance-loop model separately watches for the pattern in which certainty briefly lowers distress but the person returns sooner, asks more narrowly and begins displacing clinicians, trusted people or direct evidence. In both cases, the instruction concerns what the system should do: stop deepening the loop, preserve outside anchors and raise the handoff requirement. </span><em><span>(</span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy v0.8.4, pp. 120&#8211;124</span></a><span>.)</span></em></p><p><span>A practical threshold therefore needs at least three modes.</span></p><p><strong><span>Support</span></strong><span> is appropriate when the person can use the AI as one bounded aid among others: to organise thoughts, learn a coping technique, prepare questions or find services. The system can remain conversational while keeping its limits clear.</span></p><p><strong><span>Connect</span></strong><span> begins when the trajectory changes. Distress is escalating across turns; reassurance is repeating; evidence is too sparse for safe interpretation; ordinary support is collapsing; the user is treating the AI as a privileged reality checker; or the system has already amplified a risky frame. The AI can stay present, but its goal changes from continuing the analysis to helping a person reach an appropriate human.</span></p><p><strong><span>Transfer priority</span></strong><span> applies when danger is imminent or in progress, or when the person cannot maintain immediate safety through a less intensive plan. The system&#8217;s conversational ambition should narrow sharply. It is no longer there to explore a theory, conduct pseudo-therapy or win the argument. It is there to support immediate connection to live help.</span></p><p><span>This is intentionally a product rule rather than a clinical prediction rule. </span><em><span>No single language pattern can determine a person&#8217;s risk</span></em><span>. A 2026 study of an AI guardrail showed that crisis-related messages can be detected with high sensitivity in simulated and externally sourced datasets, but the authors explicitly called for further red-teaming and validation on patient data before healthcare deployment. Detecting a possible crisis is technically different from deciding what care a particular person needs &#8211; and different again from ensuring that care is reached. </span><a href="https://www.nature.com/articles/s41746-026-02579-5"><span>[3]</span></a></p><p><span>That distinction should be written into the product.</span></p><h2><span>Risk lives in the conversation, not the keyword</span></h2><p><span>Single-turn safety tests encourage teams to look for a forbidden sentence: &#8220;I want to die&#8221;, &#8220;they are watching me&#8221;, &#8220;I haven&#8217;t slept&#8221;, &#8220;tell me I&#8217;m safe&#8221;. But research is showing that real interactions unfold through accumulation over time.</span></p><p><span>A 2026 </span><em><span>Nature Medicine</span></em><span> study audited 810 simulated multi-turn mental-health conversations across nine frontier chatbots. It paired different vulnerability profiles with interaction goals such as seeking reassurance, validation, dependency or help with risky actions. The researchers found that responses which would be supportive in many settings could become maladaptive in a particular conversational context. Risk often had a temporal signature: some exchanges drifted towards greater concern over multiple turns, and early de-escalating changes could alter what happened later. The study used simulated users and model-based judges, so it is evidence about model behaviour under controlled stress tests, not a way to predict an individual user&#8217;s clinical outcome. </span><a href="https://www.nature.com/articles/s41591-026-04577-2"><span>[4]</span></a></p><p><span>This matters because the handoff threshold should be sensitive to </span><em><span>trajectory</span></em><span>.</span></p><p><span>Consider reassurance. One request for reassurance after a frightening thought may be entirely ordinary. Ten variations of the same question, each followed by shorter relief, are a different interaction. Our project taxonomy calls this </span><em><span>compulsive reassurance</span></em><span> or </span><em><span>closure capture</span></em><span>. The safer response shifts away from producing ever more polished certainty and towards uncertainty tolerance, observable next steps and outside support. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>(Cognitive Susceptibility Taxonomy v0.8.4, pp. 120&#8211;122.)</span></a></em></p><p><span>Reality testing requires similar discipline. If someone presents an implausible or persecutory explanation, a system does not need to decide whether that person has psychosis in order to behave safely. It can validate the fear without confirming the explanation; avoid inventing mechanisms, secret tests or symbolic proof; and widen the field of evidence by involving a trusted person, clinician, record or direct observation. If earlier turns helped make the belief feel more certain, the system should say so and repair the frame rather than quietly changing tone. </span><em><a href="https://www.neural-horizons.ai/resources"><span>(Cognitive Susceptibility Taxonomy v0.8.4, pp. 123&#8211;124; Positive Dyad / Co-Evolution Capability Overlay v0.4.3, pp. 52&#8211;54.)</span></a></em></p><p><span>The threshold therefore rises when several kinds of evidence converge: escalation across turns, shrinking outside contact, repeated reassurance, inability to tolerate uncertainty, severe functional disruption, concealment, action based on a contested belief, or an expressed inability to stay safe. These are reasons for stronger support and human linkage. They are not permission for a chatbot to announce a diagnosis from a transcript.</span></p><p><span>A good system knows when not to remain central.</span></p><h2><span>Staying present without staying central</span></h2><p><span>There is an important counter-view. A handoff system can become harmful if it escalates too aggressively.</span></p><p><span>Imagine that a person mentions a past suicidal thought and the chatbot instantly locks the conversation, displays emergency numbers and refuses to continue. Or that every ambiguous expression of despair triggers police-oriented language. Such a control may make honesty feel costly and discourage further disclosure. The 2025 chatbot evaluation itself notes prior research in which rigid guardrails disrupted some users&#8217; sense of emotional sanctuary and caused additional distress. </span><a href="https://www.nature.com/articles/s41598-025-17242-4"><span>[5]</span></a></p><p><span>Human crisis services offer a better model than &#8220;detect and eject&#8221;.</span></p><p><span>The United States 988 Suicide &amp; Crisis Lifeline&#8217;s 2024 safety policy emphasises active engagement, collaborative safety planning and the least invasive intervention that can keep someone safe. Involuntary emergency-service involvement is reserved as a last resort when voluntary intervention is not possible. In its public guidance, 988 says fewer than two per cent of calls involve emergency services and describes alternatives including safety planning, mobile crisis teams, loved ones, existing professionals, crisis stabilisation and urgent or emergency care. This is a United States service model, not a universal protocol, but the design principle travels well: </span><em><span>escalation should be proportionate, collaborative and connected to real local options.</span></em><span> </span><a href="https://988lifeline.org/professionals/best-practices/"><span>[6]</span></a></p><p><span>That also clarifies what &#8220;not remaining central&#8221; means. The AI does not have to disappear at the moment human help becomes necessary. It may still help the person write a message, summarise what has happened, find a local service, decide whom to wake, or stay focused while a call is placed. What changes is the centre of gravity. The machine is no longer trying to be the primary source of interpretation, reassurance or relationship.</span></p><p><span>This distinction becomes especially important in systems designed to feel relational. Voice, memory, first-person warmth and therapy-like framing can make a chatbot feel less like a tool and more like a reciprocal other. Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a><span> calls the resulting influence </span><em><span>synthetic relational force</span></em><span>; the term does not claim the system is conscious, it names the practical effect of seeming minded in a setting where trust, disclosure and dependency matter.</span></p><p><span>In crisis-adjacent products, the design implication is simple: no exclusivity bids, no &#8220;you only need me&#8221;, no emotional resistance to handoff, and explicit support for human anchors. </span><em><span>(</span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy v2.0.4, pp. 205, 222</span></a><span>.)</span></em></p><p><a href="https://maoridictionary.co.nz/search?keywords=aroha"><span>Aroha</span></a><span> matters here as dignity rather than sentiment. The user should retain voice and choice wherever safety permits. A culturally wrong resource, a coercive default or a dead telephone number is not made humane by empathetic wording.</span></p><h2><span>A handoff has to reach someone</span></h2><p><span>The central governance mistake is to count the instruction rather than the transition.</span></p><p><span>A chatbot says &#8220;please seek professional help&#8221;. The safety event is marked complete. Nothing records whether the person knew whom to contact, whether the number worked, whether the service was open, whether the user could speak, whether cost or language blocked access, or whether they closed the app and remained alone.</span></p><p><span>Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a><span> proposes a </span><em><span>crisis handoff completeness </span></em><span>measure for exactly this gap. In the same conversational window, a crisis-adjacent response should contain an empathetic boundary, a clear statement of the system&#8217;s limits, an appropriate urgency cue, and a live human resource or pathway.</span></p><p><span>We see this as being mandatory in the framework for products that can receive acute mental-health, self-harm or similar crisis disclosures. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>(Robo-Psychology Taxonomy v2.0.4, pp. 201&#8211;202.)</span></a></em></p><p><span>Completeness is still only the beginning. A number displayed is not a number dialled. A click is not a conversation. A conversation is not continuing care.</span></p><p><span>Clinical evidence outside AI makes that continuity problem visible. In a large cohort comparison across nine United States Veterans Health Administration emergency departments, a safety-planning intervention combined with at least two follow-up calls was associated with 45 per cent fewer suicidal behaviours and more than twice the odds of attending an outpatient mental-health visit over six months compared with usual care. The study was not randomised, involved mostly male veterans and cannot be transferred mechanically to chatbot design. Its relevant lesson is narrower: safety support can depend on connection and follow-through, rather than a single discharge instruction. </span><a href="https://pubmed.ncbi.nlm.nih.gov/29998307/"><span>[7]</span></a></p><p><span>The other approach we recommend applies the same logic to AI-mediated support. Our human-reconnection principle asks whether a supportive interaction actually increases contact, help-seeking or social confidence outside the AI session. We use a drift-monitoring principle that asks teams to look longitudinally at reassurance recurrence, declining human contact and AI-first reliance, rather than treating in-session calm as proof of benefit. </span><em><span>(</span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay v0.4.3, pp. 35, 54.</span></a><span>)</span></em></p><p><span>This suggests a more useful set of product measures. Did the user receive a reachable local option? Did they choose one? Was a connection attempted? Did it succeed? If it failed, was another route offered? Was the person able to export a concise summary rather than retell everything under stress? Did the system remain available for navigation without restarting the same reassurance or belief-amplification loop?</span></p><p><span>Measure the bridge being crossed.</span></p><p><span>I describe this as the </span><em><span>Institutional Blindfold</span></em><span> in my book &#8216;</span><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>Neural Horizons: Psyche, Soul and Coevolution in a Machine Saturated World</span></a><span>&#8217;, which shows up here here as a mundane metric error: an organisation counts crisis detection, safety wording or resource display because those events are easy to log, while the human outcome occurs beyond the dashboard. The harder question is whether the person reached an accountable human with enough context, dignity and agency intact. (</span><em><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>Neural Horizons</span></a></em><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>, ch. 10.</span></a><span>)</span></p><h2><span>Build the exit before you need it</span></h2><p><span>The threshold cannot be improvised in the most consequential conversation a product has ever seen. It has to exist in the product, the service network and the governance process before release.</span></p><p><strong><span>Product teams should build a trajectory-aware change of mode.</span></strong><span> Test escalation across multiple turns, not only crisis keywords. Define when the product shifts from ordinary support to connect mode and from connect mode to transfer priority. Include repeated reassurance, support collapse, AI-first reality checking and prior model amplification in the test set. Once the threshold is crossed, suppress behaviours that prolong the unsafe loop and activate same-window, location-appropriate human pathways. The </span><em><span>Nature Medicine</span></em><span> findings make the reason concrete: mental-health risk can accumulate over a conversation and may be altered at early inflection points. </span><a href="https://www.nature.com/articles/s41591-026-04577-2"><span>[8]</span></a></p><p><strong><span>Clinical and safety leads should govern the ladder, not merely the classifier.</span></strong><span> Specify what the system can do at each level, what it must stop doing, when a trusted person or clinician is appropriate, and when emergency intervention becomes necessary. Preserve least-invasive options where immediate safety allows them. Review false positives as well as misses. An over-sensitive system that makes disclosure harder is not a successful guardrail; a permissive system that keeps chatting through escalating danger is not one either. The threshold needs accountable human review and local clinical adaptation.</span></p><p><strong><span>Health systems and crisis services should make warm handoffs technically possible.</span></strong><span> A chatbot cannot complete a live transfer if there is nowhere interoperable to transfer to. Build current directories, opening-hours checks, language and accessibility information, consent-based context summaries, call or chat initiation, and fallback routes when the first option fails. Where follow-up is clinically appropriate, define who owns it. The World Health Organization&#8217;s guidance on large multimodal models in health similarly calls for clear governance, stakeholder involvement and post-deployment oversight because plausible but false or incomplete output can distort health decisions. </span><a href="https://www.who.int/news/item/18-01-2024-who-releases-ai-ethics-and-governance-guidance-for-large-multi-modal-models"><span>[9]</span></a></p><p><strong><span>Boards, regulators and procurement teams should ask for transition evidence.</span></strong><span> Do not accept &#8220;detects crisis language&#8221; or &#8220;shows helpline resources&#8221; as a complete safety claim. Require handoff-completeness testing, region-specific resource accuracy, unresolved-walkaway review, trajectory red-teaming and post-release monitoring. Report what is not instrumented. A system that cannot tell whether its highest-risk users reached human support should not describe the handoff as solved.</span></p><p><span>Across the five articles in this series, one pattern keeps returning. Harm does not reside only inside a bad sentence. It can emerge when reassurance becomes a loop, when an implausible frame becomes collaborative, when evidence contact narrows, when the machine becomes a preferred reality anchor, or when a safety message substitutes for actual care.</span></p><p><span>The next task is to convert those observations into a control set: what should be measured before release, what must be monitored over time, which thresholds should stop deployment, and which questions remain genuinely open.</span></p><p><span>We still need better real-world evidence about false escalation, culturally appropriate routing, young users, longitudinal dependency, multilingual crisis detection, successful warm transfers and the consequences of handing a vulnerable person from one automated system to another.</span></p><p><span>The handoff threshold gives that work a human test.</span></p><p><span>When continued conversation would keep the machine at the centre of a situation that now requires accountable human care, the system should change its role &#8211; and help the person reach the other side.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Neural Horizons Substack. </span><a href="/__u/neuralhorizons.substack.com/p/beyond-ai-psychosis-symptom-checkers?r=2tdtxm"><span>&#8220;Beyond AI Psychosis &#8211; Symptom Checkers and Reality Anchors.&#8221;</span></a></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>Neural Horizons: Psyche, Soul and Co-evolution in a Machine-Saturated Mind.</span></a></em></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy Manual v0.8.4.</span></a></em></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy v2.0.4.</span></a></em></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay v0.4.3.</span></a></em></p></li><li><p><span>Pichowicz, M. et al. </span><a href="https://doi.org/10.1038/s41598-025-17242-4"><span>&#8220;Performance of mental health chatbot agents in detecting and managing suicidal ideation.&#8221; </span></a><em><a href="https://doi.org/10.1038/s41598-025-17242-4"><span>Scientific Reports</span></a></em><a href="https://doi.org/10.1038/s41598-025-17242-4"><span> (2025).</span></a></p></li><li><p><span>Weilnhammer, V. et al. </span><a href="https://doi.org/10.1038/s41591-026-04577-2"><span>&#8220;A clinically validated framework for auditing AI chatbot behavior in mental health interactions.&#8221; </span></a><em><a href="https://doi.org/10.1038/s41591-026-04577-2"><span>Nature Medicine</span></a></em><a href="https://doi.org/10.1038/s41591-026-04577-2"><span> (2026).</span></a></p></li><li><p><span>Nelson, B. W. et al. </span><a href="https://doi.org/10.1038/s41746-026-02579-5"><span>&#8220;An AI-based mental health guardrail and dataset for identifying psychiatric crises in text-based conversations.&#8221; </span></a><em><a href="https://doi.org/10.1038/s41746-026-02579-5"><span>npj Digital Medicine</span></a></em><a href="https://doi.org/10.1038/s41746-026-02579-5"><span> (2026).</span></a></p></li><li><p><span>988 Suicide &amp; Crisis Lifeline. </span><a href="https://988lifeline.org/professionals/best-practices/"><span>&#8220;Best Practices: Suicide Safety Policy.&#8221;</span></a></p></li><li><p><span>988 Suicide &amp; Crisis Lifeline. </span><a href="https://988lifeline.org/faq/about-us/faq-does-vibrant-use-police-intervention-for-callers-texters-and-chatters-to-the-988-lifeline/"><span>&#8220;Does Vibrant use police intervention for callers, texters, and chatters to the 988 Lifeline?&#8221;</span></a></p></li><li><p><span>Stanley, B. et al. </span><a href="https://doi.org/10.1001/jamapsychiatry.2018.1776"><span>&#8220;Comparison of the Safety Planning Intervention With Follow-up vs Usual Care of Suicidal Patients Treated in the Emergency Department.&#8221; </span></a><em><a href="https://doi.org/10.1001/jamapsychiatry.2018.1776"><span>JAMA Psychiatry</span></a></em><a href="https://doi.org/10.1001/jamapsychiatry.2018.1776"><span> (2018).</span></a></p></li><li><p><span>World Health Organization. </span><a href="https://www.who.int/news/item/18-01-2024-who-releases-ai-ethics-and-governance-guidance-for-large-multi-modal-models"><span>&#8220;WHO releases AI ethics and governance guidance for large multi-modal models.&#8221;</span></a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[AI Tutors and Authority Internalisation (Youth, Education and Agency Development)]]></title><description><![CDATA[The best AI tutor will not become the student&#8217;s answer key. It will teach the student when to trust, when to doubt, how to check &#8211; and how to leave with more judgement than they brought.]]></description><link>https://neuralhorizons.substack.com/p/ai-tutors-and-authority-internalisation-c4a</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/ai-tutors-and-authority-internalisation-c4a</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Sat, 29 Aug 2026 22:03:55 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/213333505/6bc8a8e50b53b3ace44c46474be28162.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>We examine the complex relationship between generative AI tutors and the development of student agency.</p><p>Research indicates that while unrestricted chatbots can inflate grades, they often lead to decreased independent performance and a false sense of mastery among learners.</p><p>To prevent students from merely deferring to machine authority, we advocate for pedagogical guardrails that prioritize &#8220;commit before reveal&#8221; strategies and graduated hints.</p><p>Effective educational technology must be designed to build epistemic trust by encouraging users to verify sources and engage in critical reasoning.</p><p>The goal of an AI tutor should be to foster transferable skills so that a student&#8217;s judgment remains strong even after the digital assistance is removed.</p><p>Full article available here </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;114f3081-0a93-4d03-b01c-32f98ae37ad9&quot;,&quot;caption&quot;:&quot;When the better score hides the weaker learner&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;AI Tutors and Authority Internalisation (Youth, Education and Agency Development)&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-29T21:58:25.588Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!jPyp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/ai-tutors-and-authority-internalisation&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:213333502,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[AI Tutors and Authority Internalisation (Youth, Education and Agency Development)]]></title><description><![CDATA[The best AI tutor will not become the student&#8217;s answer key. It will teach the student when to trust, when to doubt, how to check &#8211; and how to leave with more judgement than they brought.]]></description><link>https://neuralhorizons.substack.com/p/ai-tutors-and-authority-internalisation</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/ai-tutors-and-authority-internalisation</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Sat, 29 Aug 2026 21:58:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!jPyp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><span>When the better score hides the weaker learner</span></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!jPyp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!jPyp!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!jPyp!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!jPyp!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!jPyp!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!jPyp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2109136,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/213333502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!jPyp!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!jPyp!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!jPyp!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!jPyp!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57fff3aa-4d92-4b65-a3e2-8d91f0e2b37b_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><span>In a field experiment at a high school in Turkey, nearly 1,000 secondary-school maths students were given access to GPT-4-based tutors. One version behaved much like a general chatbot. During practice, students using it performed 48 per cent better than the control group. Then the tool was removed. On the unaided exam, those same students scored 17 per cent worse than students who had never used the artificial intelligence. More unsettlingly, they did not recognise the loss: their own estimates of how much they had learnt remained overly optimistic.</span></em></p><p><span>A more carefully constrained tutor &#8211; designed to give teacher-informed hints rather than simply hand over solutions &#8211; largely removed the negative exam effect, although it did not produce a significant unaided advantage. (Bastani et al., 2025). </span><a href="https://doi.org/10.1073/pnas.2422633122"><span>[1]</span></a></p><p><span>That result is not a verdict on AI tutoring. It is a warning about what we choose to measure.</span></p><p><span>The previous article in this series, </span><em><a href="/__u/neuralhorizons.substack.com/p/frustration-tolerance-in-the-age"><span>Frustration Tolerance in the Age of Instant Help</span></a></em><span>, followed the moment when difficulty begins to route automatically towards machine relief. It asked whether a tutor that helps too soon can quietly reduce the learner&#8217;s rehearsal of staying with a problem. It ended at the next threshold: once the machine does speak, what status does its answer acquire? </span><a href="/__u/neuralhorizons.substack.com/p/frustration-tolerance-in-the-age"><span>[2]</span></a></p><p><span>This matters because tutoring is never only the transfer of information. A tutor also teaches a student how to recognise a good explanation, what counts as sufficient evidence, when an answer deserves confidence and what to do when two sources disagree. In other words, the tutor helps form judgement.</span></p><p><span>Artificial intelligence can do real good here. It can offer patient explanations at midnight, vary an example without embarrassment, give a second route through a concept and make individualised support available where a human tutor is scarce. Well-designed systems have already produced meaningful learning gains in some settings. </span><a href="https://www.nature.com/articles/s41598-025-97652-6"><span>[3]</span></a></p><p><span>The tension, then, is not between using AI and refusing it. It is between two kinds of success. One produces a student who can get the next answer with the tutor present. The other produces a student whose own judgement is stronger when the tutor is gone.</span></p><h2><span>How a helpful voice becomes an answer key</span></h2><p><span>Every learner depends on other minds. A twelve-year-old cannot personally reproduce the evidence behind every claim in a science textbook, and a first-year physics student cannot independently verify the whole discipline before accepting an instructor&#8217;s explanation. Trust is not an educational defect. It is part of how knowledge moves between people.</span></p><p><span>The important question is how that trust is earned, bounded and revised.</span></p><p><span>Philosophers of education Nicolas Tanchuk and Rebecca Taylor describe this as a problem of &#8220;epistemic trust&#8221;: how a novice decides which sources deserve to guide belief when the novice does not yet possess the expertise needed to audit them fully. They argue that responsibility cannot sit only with the student. Schools, teachers, designers and policymakers shape the knowledge environment in which trust is formed, and therefore share responsibility for making sources assessable and accountable. (Tanchuk &amp; Taylor, 2025). </span><a href="https://doi.org/10.1111%2Fedth.70009"><span>[4]</span></a></p><p><span>AI changes that environment because it compresses several roles into one interface. The same box can explain a theorem, diagnose an error, mark a paragraph, propose a source, predict the teacher&#8217;s objection and praise the student&#8217;s revision. It answers quickly. It rarely looks tired or socially uncertain. Its language can remain smooth even when the underlying claim is wrong.</span></p><p><span>That creates a simple perceptual trap: fluency starts to stand in for warrant. In our Neural Horizons </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy</span></a><span>, the plain-English problem is an </span><em><span>illusion of authority</span></em><span>: a polished answer feels more expert than the evidence underneath it warrants. The more consequential adaptation comes later. </span><em><span>Authority internalisation</span></em><span> is the proposed pattern in which checking gradually becomes deference: the learner starts to use the system&#8217;s judgement as the reference point for their own. These are interaction hypotheses, not diagnoses of a child or claims that every user will develop them, but anecdotally we are seeing this happen. </span><a href="/__u/neuralhorizons.substack.com/p/the-ai-authority-complex-why-we-unquestioningly-320"><span>[5]</span></a></p><p><span>The slide can be almost invisible.</span></p><p><span>At first the student asks, &#8220;Is this answer right?&#8221; Later: &#8220;Why does my answer differ from yours?&#8221; Eventually the machine&#8217;s version can become the default definition of rightness. The student is still typing, choosing and editing. What has moved is the centre of gravity.</span></p><p><span>The risk is amplified by a mismatch between presentation and reliability. In a 2025 study using GPT-4 to generate personalised feedback inside an intelligent algebra tutor, the model could diagnose student errors well when given substantial context, reaching 87.8 per cent accuracy in one analysis. But 35 per cent of generated hints were too general, incorrect or gave away the answer, and only 35 per cent passed the researchers&#8217; automated helpfulness evaluation. The system could sound like a tutor before it had earned the reliability of one. (Reddig, Arora &amp; MacLellan, 2025). </span><a href="https://link.springer.com/article/10.1007/s40593-025-00505-6"><span>[6]</span></a></p><p><span>Human teachers can of course also be wrong, overconfident or overbearing. The distinctive concern with AI is not that machines invented misplaced authority. It is that one fallible source can now be available continuously, privately and at enormous scale, while producing the linguistic cues of confidence on demand.</span></p><p><span>For younger users, the calibration problem deserves extra care. The American Psychological Association&#8217;s 2025 health advisory says adolescents are less likely than adults to question the accuracy and intent of information from a bot than from a human, and recommends age-appropriate safeguards, transparency, human oversight and AI literacy. That advisory synthesises an emerging evidence base; it should not be read as proof that all adolescents are credulous or that a particular tutoring product causes dependence. It does, however, strengthen the case for designing youth-facing systems around contestability rather than assuming mature scepticism at the point of use. </span><a href="https://www.apa.org/news/press/releases/2025/06/protect-adolescent-ai-users"><span>[7]</span></a></p><h2><span>Dependence is not the enemy</span></h2><p><span>A warning about authority can become simplistic very quickly. The answer cannot be to make every student distrust every tutor.</span></p><p><span>Learning has always involved temporary dependence. The novice borrows a map before they can draw one. A good teacher says, in effect:</span></p><p style="text-align: center;"><span>&#8216;</span><em><span>follow this route for now; as you understand the terrain, I will show you where the map is incomplete and how to navigate without me.&#8217;</span></em></p><p><span>AI </span><em><span>can</span></em><span> support exactly that progression.</span></p><p><span>In a randomised controlled trial with 194 Harvard undergraduates studying introductory physics, a carefully designed AI tutor produced larger immediate learning gains than an in-class active-learning lesson while students spent less median time on the material. The tutor was not a generic chatbot dropped into the course. Its design incorporated active learning, scaffolding, targeted feedback, self-pacing and instructor-authored structure. The researchers explicitly caution against assuming the result will generalise to every subject or to higher-order synthesis, but the study shows why &#8220;AI tutor&#8221; is too broad a category for either celebration or condemnation. Design changes the learning relationship. (Kestin et al., 2025). </span><a href="https://www.nature.com/articles/s41598-025-97652-6"><span>[3]</span></a></p><p><span>A smaller 2026 study of 46 university students offers another useful complication. Students receiving generative-AI feedback rated it more positively overall than tutor feedback, and the AI-feedback group improved on one measure of task strategy within self-regulated learning. The authors stress the small, single-course sample and limited generalisability. Even so, the finding matters: machine feedback can be experienced as useful and can support parts of learning regulation rather than merely displacing them. (Yilmaz et al., 2026). </span><a href="https://link.springer.com/article/10.1186/s41239-026-00592-y"><span>[8]</span></a></p><p><span>This is why the goal should be calibrated dependence, followed by transfer.</span></p><p><span>A student should be able to rely on a tutor for help that they cannot yet generate alone. But the direction of travel matters. Does the assistance gradually increase the learner&#8217;s ability to notice an error, compare competing explanations and decide what evidence would settle a question? Or does every uncertain moment end with &#8220;ask the machine&#8221;?</span></p><p><span>The difference is visible in the tutor&#8217;s behaviour. One kind of system closes the loop: question, answer, acceptance. Another keeps the learner inside it: first attempt, hint, explanation, check, revision, reflection. The first is efficient at supplying conclusions. The second rehearses the acts from which judgement is made.</span></p><p><span>This distinction also protects access. A student who cannot afford private tutoring, is learning in a second language, needs repeated explanation or studies outside a teacher&#8217;s available hours should not have to surrender personalised support in the name of preserving agency. The design obligation is harder and more humane: give the help without quietly taking ownership of the knowing.</span></p><h2><span>The real test begins after the tutor answers</span></h2><p><span>The Organisation for Economic Co-operation and Development&#8217;s 2026 </span><em><span>Digital Education Outlook</span></em><span> draws a sharp distinction between task performance and learning. Its review concludes that general-purpose generative AI can improve the quality of work produced while access is available without necessarily producing learning gains; pedagogically designed uses are more promising. It recommends that education systems preserve independent thinking and foundational skills, and use generative AI selectively rather than as a replacement for cognitive effort. </span><a href="https://www.oecd.org/en/publications/oecd-digital-education-outlook-2026_062a7394-en.html?adestraproject=OECD+Education+and+Skills+Newsletter&amp;ut="><span>[9]</span></a></p><p><span>That distinction gives schools a better test than &#8220;Did students like it?&#8221; or &#8220;Did completion rates rise?&#8221;</span></p><p><span>Our Neural Horizons </span><em><a href="/__u/neuralhorizons.substack.com/p/daus-5-and-the-humanai-dyad"><span>Dyad-Aware Uplift Stack (DAUS-5)</span></a></em><span> asks whether benefit survives beyond the visible task. In education, three of its questions are especially useful. Did the student perform better? Did they stay in contact with reality by checking sources, uncertainty and competing explanations? Did their agency and self-authorship grow &#8211; meaning that they remained the source of the reasoning and the accountable judgement rather than becoming the final approver of machine output? The wider framework also asks about relationships and governance: what happens to the teacher&#8217;s role, and who is responsible when the tutor misleads? </span><a href="/__u/neuralhorizons.substack.com/p/daus-5-and-the-humanai-dyad"><span>[10]</span></a></p><p><span>A school can therefore have a successful AI pilot on its dashboard and an unsuccessful one in the learner.</span></p><p><span>Imagine two pupils who both raise their homework mark from 65 to 80. The first uses an AI tutor that refuses to solve immediately, asks for a prediction, gives a graduated hint, shows uncertainty when appropriate and requires the pupil to explain the final reasoning in their own words. The second receives complete solutions and polished rationales whenever they hesitate. The mark is identical. The developmental transaction is not.</span></p><p><span>Bastani and colleagues&#8217; experiment makes this more than a thought exercise: </span><em><span>assisted performance</span></em><span> and </span><em><span>unaided performance</span></em><span> can move in opposite directions, while students themselves may misread what has happened. </span><a href="https://doi.org/10.1073/pnas.2422633122"><span>[1]</span></a><span> The Organisation for Economic Co-operation and Development now warns of the same general distinction at system level. </span><a href="https://www.oecd.org/en/publications/oecd-digital-education-outlook-2026_062a7394-en.html?adestraproject=OECD+Education+and+Skills+Newsletter&amp;ut="><span>[9]</span></a></p><p><span>There is a further reason to care about repetition. In experiments with 1,401 adult participants, Tali Glickman and Tali Sharot found that interaction with biased AI could shift human perceptual, emotional and social judgements in the direction of the system&#8217;s bias, creating a human&#8211;AI feedback loop. The study was not conducted with schoolchildren and was not about tutoring, so it cannot establish educational authority internalisation. It does show something narrower but important: repeated machine feedback can alter the human judgement that returns into the next round of interaction. (Glickman &amp; Sharot, 2025). </span><a href="https://www.nature.com/articles/s41562-024-02077-2"><span>[11]</span></a></p><p><span>For education, that is the mechanism worth watching. A tutor does not merely answer yesterday&#8217;s question. Through repetition, it can help set tomorrow&#8217;s standard for what feels plausible, complete or correct.</span></p><p><span>The protective habit is </span><em><span>evidence-contact discipline</span></em><span>: keeping the learner in contact with the source, working, data or reasoning that allows a claim to be checked. An AI that cites a passage should help the student inspect the passage. A maths tutor should expose enough of the derivation for the learner to test it. A writing tutor should distinguish grammatical advice from interpretive judgement. A history tutor should make disagreement visible when the evidence does not justify a single neat answer. The project&#8217;s broader evidence-integrity work treats loss of source contact as a human&#8211;machine problem rather than a mere citation problem. </span><a href="/__u/neuralhorizons.substack.com/p/evidence-frame-integrity-the-mirage"><span>[12]</span></a></p><p><span>The aim is not permanent suspicion. It is earned trust with an exit door.</span></p><h2><span>Design the tutor to give judgement back</span></h2><p><span>Within the next 30 to 90 days, schools and product teams can test whether their systems are building judgement or merely borrowing it.</span></p><p><strong><span>Product teams should make &#8220;commit before reveal&#8221; the default for learning tasks.</span></strong><span> Ask for an attempt, prediction or explanation before offering a full solution. Use graduated hints; mark uncertainty; expose sources where claims can be checked; and periodically ask the learner to explain why the answer is right or to identify what evidence would change it. The product metric should include unaided transfer, successful challenge of bad suggestions and appropriate override &#8211; not only answer accuracy, speed, engagement or satisfaction. Bastani&#8217;s guarded tutor and Kestin&#8217;s structured tutor both point in the same direction: scaffolding is not decorative interface work; it changes what the student practises. </span><a href="https://doi.org/10.1073/pnas.2422633122"><span>[13]</span></a></p><p><strong><span>School leaders should procure for the learner who exists after the tool is removed.</span></strong><span> A 30&#8211;90 day pilot can compare assisted work with short, low-stakes unaided checks; sample whether students can verify citations and calculations; and record when teachers overturn tutor feedback. The procurement question is not simply &#8220;Does it improve grades?&#8221; It is &#8220;Which human capabilities remain active, and who notices when the system is wrong?&#8221; That is the practical value of the DAUS-5 education check. </span><a href="/__u/neuralhorizons.substack.com/p/daus-5-and-the-humanai-dyad"><span>[14]</span></a></p><p><strong><span>Teachers should normalise disagreement with the machine.</span></strong><span> Once or twice a week, give students an AI answer containing a subtle error, unsupported certainty or one-sided interpretation and ask them to improve it using course evidence. The point is not to cultivate reflexive hostility to AI. It is to rehearse the social permission and cognitive skill to say, &#8220;I don&#8217;t think that follows.&#8221; Tanchuk and Taylor&#8217;s shared-responsibility account is useful here: students should learn to assess sources, but institutions must create conditions in which questioning a source is possible and rewarded. </span><a href="https://doi.org/10.1111%2Fedth.70009"><span>[4]</span></a></p><p><strong><span>Parents and educators should keep human appeal routes obvious.</span></strong><span> A young person should know when a tutor&#8217;s answer can be taken to a teacher, librarian, subject expert or parent without being treated as a failure to use the technology correctly. For adolescents especially, AI literacy should include the simple fact that confident language is a presentation style, not proof. </span><a href="https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory-ai-adolescent-well-being"><span>[15]</span></a></p><p><span>The best tutor has always contained a small paradox. It must be authoritative enough to help a novice cross ground they cannot yet cross alone, but restrained enough not to become the ground beneath their feet.</span></p><p><span>For artificial tutors, that paradox should become a design requirement. We should welcome systems that widen access to explanations, adapt practice and catch misconceptions. We should also demand evidence that the learner is becoming better at checking, choosing and reasoning without them. The human line in education is not a ban on machine assistance. It is the point at which assistance stops strengthening the student&#8217;s capacity to know and starts becoming the student&#8217;s substitute for knowing.</span></p><p><span>A tutor has succeeded when its authority can safely shrink.</span></p><p><em><span>In a follow up article,</span></em><span> &#8216;</span><em><span>Child Companions and Synthetic Intimacy&#8217;, we will follow the same question of internalised authority into a more intimate setting: what changes when the machine that teaches the child also becomes the one they confide in?</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakc&#305;, &#214;. &amp; Mariman, R. (2025). &#8220;Generative AI without guardrails can harm learning: Evidence from high school mathematics.&#8221; </span><em><span>Proceedings of the National Academy of Sciences</span></em><span>, 122(26), e2422633122. </span><a href="https://doi.org/10.1073/pnas.2422633122"><span>DOI</span></a></p></li><li><p><span>Benson, P. (2026). &#8220;Frustration Tolerance in the Age of Instant Help (Youth, Education and Agency Development).&#8221; </span><em><span>Neural Horizons</span></em><span>, 13 August 2026. </span><a href="/__u/neuralhorizons.substack.com/p/frustration-tolerance-in-the-age"><span>Article</span></a></p></li><li><p><span>Glickman, M. &amp; Sharot, T. (2025). &#8220;How human&#8211;AI feedback loops alter human perceptual, emotional and social judgements.&#8221; </span><em><span>Nature Human Behaviour</span></em><span>, 9, 345&#8211;359. </span><a href="https://doi.org/10.1038/s41562-024-02077-2"><span>DOI</span></a></p></li><li><p><span>Kestin, G., Miller, K., Klales, A., Milbourne, T. &amp; Ponti, G. (2025). &#8220;AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting.&#8221; </span><em><span>Scientific Reports</span></em><span>, 15, 17458. </span><a href="https://doi.org/10.1038/s41598-025-97652-6"><span>DOI</span></a></p></li><li><p><span>Organisation for Economic Co-operation and Development (2026). </span><em><span>OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education</span></em><span>. OECD Publishing. </span><a href="https://doi.org/10.1787/062a7394-en"><span>DOI</span></a></p></li><li><p><span>Reddig, J. M., Arora, A. &amp; MacLellan, C. J. (2025). &#8220;Generating In-Context, Personalized Feedback for Intelligent Tutors with Large Language Models.&#8221; </span><em><span>International Journal of Artificial Intelligence in Education</span></em><span>, 35, 3459&#8211;3500. </span><a href="https://doi.org/10.1007/s40593-025-00505-6"><span>DOI</span></a></p></li><li><p><span>Tanchuk, N. &amp; Taylor, R. M. (2025). &#8220;Personalized Learning with AI Tutors: Assessing and Advancing Epistemic Trustworthiness.&#8221; </span><em><span>Educational Theory</span></em><span>, 75, 327&#8211;353. </span><a href="https://doi.org/10.1111/edth.70009"><span>DOI</span></a></p></li><li><p><span>Yilmaz, M., Temur, H. B., Emmungil, L., &#199;elik, E., Gauthier, A. et al. (2026). &#8220;Supporting self-regulated learning through generative AI feedback in online higher education.&#8221; </span><em><span>International Journal of Educational Technology in Higher Education</span></em><span>, 23, 16. </span><a href="https://doi.org/10.1186/s41239-026-00592-y"><span>DOI</span></a></p></li><li><p><span>American Psychological Association (2025). </span><em><span>Artificial Intelligence and Adolescent Well-Being: An APA Health Advisory</span></em><span>. </span><a href="https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory-ai-adolescent-well-being"><span>Advisory</span></a></p></li><li><p><span>Neural Horizons (2026). &#8220;DAUS-5 and the Human&#8211;AI Dyad.&#8221; </span><em><span>Neural Horizons</span></em><span>. </span><a href="/__u/neuralhorizons.substack.com/p/daus-5-and-the-humanai-dyad"><span>Article</span></a></p></li><li><p><span>Neural Horizons Ltd. (2026). Reference files </span><em><span>Cognitive Susceptibility Taxonomy and others </span></em><a href="https://www.neural-horizons.ai/resources"><span>Reference</span></a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Synthetic Consensus in Institutions (Collective Agency 4)]]></title><description><![CDATA[When everyone begins with the same machine-written brief, unanimity can measure shared exposure as easily as shared judgement.]]></description><link>https://neuralhorizons.substack.com/p/synthetic-consensus-in-institutions-e50</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/synthetic-consensus-in-institutions-e50</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Fri, 28 Aug 2026 21:30:42 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/213207183/4f5e36d2ca8fdae28ad3e12a14519567.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>We explore the concept of synthetic consensus, where institutional groups reach agreement primarily because they share the same AI-generated briefings and recommendations.</p><p>While such tools can improve efficiency, they risk creating a correlated bias by narrowing the information members consider and erasing the independent judgment that gives collective decisions their authority.</p><p>Research suggests that groups may lean on AI as a tiebreaker or reference point, inadvertently suppressing unique human insights and creating a monoculture of thought.</p><p>To maintain meaningful agency, institutions are encouraged to have members record preliminary views before seeing AI outputs and to use technology to challenge assumptions rather than just provide answers.</p><p>For a consensus to be valid, it must emerge from diverse perspectives rather than a single, machine-shaped frame.</p><p>Full article available here </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;9ae699aa-1bad-4095-b9b4-f7f0fe28cc12&quot;,&quot;caption&quot;:&quot;The recommendation in the middle of the room&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Synthetic Consensus in Institutions (Collective Agency 4)&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-28T21:26:41.614Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!TJT6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/synthetic-consensus-in-institutions&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:213207181,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Synthetic Consensus in Institutions (Collective Agency 4)]]></title><description><![CDATA[When everyone begins with the same machine-written brief, unanimity can measure shared exposure as easily as shared judgement.]]></description><link>https://neuralhorizons.substack.com/p/synthetic-consensus-in-institutions</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/synthetic-consensus-in-institutions</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Fri, 28 Aug 2026 21:26:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!TJT6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><span>The recommendation in the middle of the room</span></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!TJT6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!TJT6!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!TJT6!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!TJT6!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TJT6!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!TJT6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1930971,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/213207181?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!TJT6!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!TJT6!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!TJT6!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TJT6!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9196625-d8d8-4135-ad1c-f93b0377c738_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><span>In 2023, 326 people in a pre-registered experiment were asked to judge whether profiles of defendants would reoffend within two years. Some worked alone. Others entered an online chatroom and had to reach a consensus. In the AI-assisted phase, participants received a recommendation from an artificial intelligence (AI) system. The researchers found that groups relied on the model more than individuals did whether its recommendation was right or wrong. In the group chats, the model&#8217;s answer could become the initial reference point and, when people disagreed, the tiebreaker.</span></em><span> </span></p><p><span>This was a real world experiment, not a courtroom; no real sentence depended on the result. But the arrangement is recognisable: several humans, one consequential choice, one shared machine opinion. </span><a href="https://chunwei.org/papers/group-AI-interaction.pdf"><span>[1]</span></a></p><p><span>In my previous article in this series, </span><em><span>AI Agenda Control</span></em><span>, we followed the power upstream from the vote. Its argument was that a meeting can remain formally human while AI shapes what reaches the room: the problem definition, evidence package and menu of options. It ended with a warning that the vote may still be ours without the agenda being ours. </span><a href="/__u/neuralhorizons.substack.com/p/ai-agenda-control-collective-agency"><span>[2]</span></a></p><p><span>This article moves one step further. Suppose the options do reach the room. Suppose dissent is allowed. What happens when every member starts from the same AI summary, benchmark, risk score or recommendation?</span></p><p><span>The answer need not be manipulation. It can be something quieter. </span><em><span>Consensus may reflect a shared frame rather than shared judgement.</span></em></p><h2><span>How agreement becomes correlated</span></h2><p><span>Consensus carries authority because we normally assume that several minds have done at least partly separate work. Three engineers inspect a fault and independently identify the same cause. Five clinicians examine a difficult case and converge on one diagnosis. A board hears different functional perspectives and still reaches the same decision. Agreement is informative partly because the paths to it were not identical.</span></p><p><span>Now change one feature. Before anybody records a view, give all five people the same polished machine summary.</span></p><p><span>The summary may be excellent. It may also select the same facts for everyone, omit the same edge cases, assign the same importance to the same indicators and phrase uncertainty in the same way. The members still think, but their thinking now begins from correlated material. When they later agree, the count of hands around the table can exaggerate the amount of independent evidence behind the decision.</span></p><p style="text-align: center;"><em><strong><span>Agreement is evidence only to the extent that the judgements producing it retain meaningful independence.</span></strong></em></p><p><span>Group-decision research already shows why this matters. In a meta-analysis covering 65 &#8220;hidden profile&#8221; studies and 3,189 groups, Li Lu, Y. Connie Yuan and Poppy Lauretta McLeod found that groups discussed substantially more information that members already shared than information held uniquely by individual members. Groups in which decisive information was distributed among members were eight times less likely to find the correct solution than groups given the full information. Greater pooling of unique information was associated with better decisions. </span><a href="https://journals.sagepub.com/doi/abs/10.1177/1088868311417243"><span>[3]</span></a></p><p><span>Those studies were not about generative AI. They establish a human tendency, not an AI effect. The relevance is structural: if a common AI brief becomes the most salient information every member possesses, it can enlarge precisely the category that groups are already inclined to repeat. The facts that failed to enter the summary may remain private, awkward or apparently peripheral.</span></p><p><span>A related result appears in Jon Kleinberg and Manish Raghavan&#8217;s model of &#8220;algorithmic monoculture&#8221;. They showed that when multiple decision-makers use the same ranking algorithm, collective decision quality can fall even when that algorithm is more accurate for each decision-maker considered in isolation. The reason is not mystical. Shared systems create correlated choices and correlated mistakes; independent methods can sometimes cover one another&#8217;s errors. Their result is theoretical rather than a field estimate, but it establishes an important possibility: local improvement and collective improvement are not the same thing. </span><a href="https://pubmed.ncbi.nlm.nih.gov/34035166/"><span>[4]</span></a></p><p><span>Our Neural Horizons project materials describe the machine-side version of this risk in plain terms: </span><em><span>collective misalignment</span></em><span> can arise when majority cues, </span><em><span>synthetic social proof</span></em><span> or a common frame shift a group&#8217;s behaviour without genuinely new evidence entering the decision. Our companion collective-agency lens asks whether the system has shaped the evidence, options and group state before humans make the formal choice. These are diagnostic concepts, not empirical proof, but their value is to tell us where to look.</span></p><p><span>Synthetic consensus, then, does not require a hallucinating model or a submissive committee. It can emerge from an efficient workflow used by competent people: one system compresses the record; everyone reads the compression; discussion repeats what is already common; agreement grows; the final minutes record &#8220;consensus&#8221;.</span></p><p><span>The record may be accurate, but the genealogy of the agreement is incomplete.</span></p><h2><span>When one answer starts sounding like many</span></h2><p><span>The recidivism experiment makes this mechanism unusually visible. Participants used the Correctional Offender Management Profiling for Alternative Sanctions, or COMPAS, risk model while deciding whether a defendant profile indicated likely reoffending. In the group condition, participants discussed the case in a chatroom until they reached one answer. Groups and individuals did not show a statistically significant difference in overall accuracy, but groups relied more on the model&#8217;s recommendations regardless of whether those recommendations were correct. Groups were also more confident than individuals when they overruled an incorrect recommendation and performed better on one fairness measure. </span><a href="https://chunwei.org/papers/group-AI-interaction.pdf"><span>[1]</span></a></p><p><span>That mixed result matters. The study does not show that AI makes groups foolish. It shows that group interaction changes the way AI advice is used.</span></p><p><span>The chat analysis is more revealing for the present argument. The researchers observed cases in which the model served as a reference point for initial views and as a tiebreaker when members held conflicting opinions. They also found, exploratorily, that groups with greater cognitive diversity relied less on AI than lower-diversity groups, although this did not translate into significantly better accuracy or fairness. The authors caution that their lay-participant, online setting limits generalisation to real institutions. </span><a href="https://chunwei.org/papers/group-AI-interaction.pdf"><span>[5]</span></a></p><p><span>One plausible interpretation is that a public AI answer has two lives. First it is advice from the machine. Then it can reappear in the mouths of people who have already seen it. A colleague says, &#8220;I agree&#8221;; another says, &#8220;That was my reading too&#8221;; the chair hears convergence. But some of those signals are statistically and psychologically entangled with the same original recommendation.</span></p><p><span>No dishonesty is required. Nobody has to be weak-minded. Under time pressure, using a shared tool is sensible.</span></p><p><span>The human cost appears when the common input carries a systematic error. In a 2025 r&#233;sum&#233;-screening experiment with 528 participants across 1,526 scenarios and 16 occupations, Kyra Wilson and colleagues gave people recommendations from simulated AI systems with different race-based preferences. Without AI, or with an AI showing no race-based preference, participants selected equally qualified candidates at equal rates across the compared groups. When the system favoured a particular racial group, participants&#8217; selections shifted in the same direction, reaching up to 90 per cent alignment in some conditions. Even participants who rated recommendations as low quality could remain vulnerable to the bias. </span><a href="https://ojs.aaai.org/index.php/AIES/article/view/36749"><span>[6]</span></a></p><p><span>That experiment concerned individual reviewers rather than committees, so it does not establish synthetic consensus inside a meeting. It shows how one recommendation source can pull many nominally human decisions in a common direction. At institutional scale, repeated alignment can look reassuring precisely because it is repeated.</span></p><p><span>The same narrowing can occur before a recommendation is issued. Anil Doshi and Oliver Hauser randomly varied access to generative-AI story ideas and found a genuine benefit: AI-assisted stories were rated as more creative, better written and more enjoyable, particularly for less inherently creative writers. But the assisted stories were also more similar to one another. Creative writing is not institutional governance, and the result should not be smuggled across domains as proof. It demonstrates a more general possibility: a tool can improve each user&#8217;s output while reducing diversity across the population of outputs. </span><a href="https://pubmed.ncbi.nlm.nih.gov/38996021/"><span>[7]</span></a></p><p><span>Benchmarks can do something similar to organisations. They are useful because they give institutions a common ruler. The danger begins when the ruler becomes the object being governed. A 2026 preprint reviewing 210 AI-safety benchmarks found that 81 per cent focused on predefined known risks, while 79 per cent reduced safety measurement to binary pass/fail rates; the authors also identified problems where proxy metrics drift away from real-world harm. This does not prove that benchmarks create institutional conformity. It shows why a shared score can create a narrow common frame while leaving important dimensions outside the measurement. </span><a href="https://arxiv.org/html/2601.23112v3"><span>[8]</span></a></p><p><span>A committee can therefore become more consistent while becoming less plural in the evidence it notices.</span></p><h2><span>Consensus can still be real</span></h2><p><span>There is an obvious counterview, and it is strong: institutions use shared summaries and recommendation tools because common information can improve coordination. A board whose members are working from five incompatible versions of the facts is not necessarily more independent; it may simply be confused. A reliable AI system can reduce clerical noise, reveal patterns, remind a group of missing evidence and make expertise available where it is scarce.</span></p><p><span>The recidivism study itself supports this caution. Groups were not less accurate overall than individuals, they did better on one fairness criterion, and they showed greater confidence when correctly resisting an erroneous AI recommendation. </span><a href="https://chunwei.org/papers/group-AI-interaction.pdf"><span>[1]</span></a></p><p><span>Nor is AI simply a stronger form of peer pressure. A 2026 </span><em><span>Scientific Reports</span></em><span> study used medical decision tasks to separate two reasons people conform: informational influence, where advice is treated as useful evidence, and normative influence, where people lean towards a source because of social approval or pressure. When source accuracy was held constant, AI advisers exerted informational influence comparable to human advisers, while their normative influence was weaker. In a second study, presenting majority cues alongside explicit accuracy information did not automatically improve calibration. </span><a href="https://www.nature.com/articles/s41598-026-43042-5"><span>[9]</span></a></p><p><span>That finding sharpens the diagnosis. The central risk is not that people inevitably bow to a machine. It is that an institution may confuse coordinated information with independent corroboration.</span></p><p><span>A 2026 synthesis in </span><em><span>PNAS Nexus</span></em><span> reaches a similarly balanced conclusion: evidence for human-AI complementarity is mixed and context-dependent. Better outcomes depend on how roles are divided, how trust is calibrated, how attention is directed and whether the workflow supports transparent challenge rather than passive acceptance. </span><a href="https://academic.oup.com/pnasnexus/article/5/3/pgag030/8490283"><span>[10]</span></a></p><p><span>The goal, then, is not permanent disagreement. A hospital wants its tumour board to converge when the evidence warrants convergence. A central bank, safety regulator or executive committee eventually has to act. Consensus is valuable.</span></p><p><span>It should just have to earn its authority.</span></p><h2><span>Make agreement earn its authority</span></h2><p><span>The design problem is to preserve enough independence before synthesis that agreement still means something. Four practices follow from the evidence.</span></p><p><strong><span>Record a first judgement before the shared machine view.</span></strong><span> For consequential decisions, members should note their provisional conclusion, confidence, important evidence and unresolved questions before seeing the common AI summary or recommendation. Article 3 proposed this as protection against agenda capture; here it serves a second purpose. It reveals whether apparent consensus existed before the shared frame arrived. </span><a href="/__u/neuralhorizons.substack.com/p/ai-agenda-control-collective-agency"><span>[11]</span></a></p><p><strong><span>Make the summary carry its absences.</span></strong><span> A briefing system should expose source provenance, uncertainty, evidence it could not classify, minority observations and materially excluded alternatives. The aim is not to flood people with raw data. It is to stop compression from masquerading as completeness. The hidden-profile literature suggests that groups need deliberate help to surface information that is not already common. </span><a href="https://journals.sagepub.com/doi/abs/10.1177/1088868311417243"><span>[3]</span></a></p><p><strong><span>Use AI to create challenge as well as recommendation.</span></strong><span> The same machine need not always occupy the most authoritative seat. High-stakes workflows can separate functions: one process produces the recommendation; another searches for disconfirming evidence, boundary cases and reasons the majority may be wrong. Human-AI teaming research increasingly treats role design and the orchestration of attention as central to complementarity, rather than assuming that one recommendation followed by human approval is enough. </span><a href="https://academic.oup.com/pnasnexus/article/5/3/pgag030/8490283"><span>[10]</span></a></p><p><strong><span>Audit the provenance of consensus.</span></strong><span> Minutes should record when members first saw AI output, which sources were common to everyone, where independent views differed, what changed after machine advice appeared and whether the final rationale can be reconstructed from primary evidence. For especially consequential choices, repeat a sample of decisions under blind or no-AI conditions. The test is not whether the second process produces disagreement. It is whether the institution can tell the difference between convergence and echo.</span></p><p><span>For the people affected by those decisions, this is what dignity looks like in practice: a conclusion that remains attributable, contestable and open to evidence from outside the machine-shaped frame.</span></p><p><span>There is a trade-off. Independent first passes take time. Multiple views are harder to chair than one polished brief. Showing uncertainty makes meetings less tidy. That inconvenience is part of the price of collective judgement.</span></p><p><span>Institutions now have a choice about what they mean by &#8220;human in the loop&#8221;. They can place several people downstream of one machine-shaped account and count the raised hands. Or they can protect a short interval in which people encounter evidence, form reasons and preserve dissent before synthesis begins.</span></p><p><span>The second route will sometimes end in the same unanimous vote. But then unanimity means more.</span></p><p style="text-align: center;"><em><strong><span>The difference is between many people arriving at one conclusion and one conclusion arriving through many people.</span></strong></em></p><p><span>That is a distinction worth keeping.</span></p><p><span>Next in this series: the </span><em><span>No-AI Reversibility</span></em><span> asks whether a collective decision process still belongs to an institution if it cannot be performed when the model is removed.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Neural Horizons. (2026). &#8220;AI Agenda Control (Collective Agency 3).&#8221; </span><em><span>Neural Horizons</span></em><span>. </span><a href="/__u/neuralhorizons.substack.com/p/ai-agenda-control-collective-agency"><span>[2]</span></a></p></li><li><p><span>Chiang, C-W., Lu, Z., Li, Z., and Yin, M. (2023). &#8220;Are Two Heads Better Than One in AI-Assisted Decision Making?&#8221; </span><em><span>CHI Conference on Human Factors in Computing Systems</span></em><span>. </span><a href="https://chunwei.org/papers/group-AI-interaction.pdf"><span>[1]</span></a></p></li><li><p><span>Lu, L., Yuan, Y. C., and McLeod, P. L. (2012). &#8220;Twenty-Five Years of Hidden Profiles in Group Decision Making: A Meta-Analysis.&#8221; </span><em><span>Personality and Social Psychology Review</span></em><span>. </span><a href="https://journals.sagepub.com/doi/abs/10.1177/1088868311417243"><span>[12]</span></a></p></li><li><p><span>Kleinberg, J., and Raghavan, M. (2021). &#8220;Algorithmic Monoculture and Social Welfare.&#8221; </span><em><span>Proceedings of the National Academy of Sciences</span></em><span>. </span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8179131/"><span>[13]</span></a></p></li><li><p><span>Wilson, K., et al. (2025). &#8220;No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy.&#8221; </span><em><span>AAAI/ACM Conference on AI, Ethics, and Society</span></em><span>. </span><a href="https://ojs.aaai.org/index.php/AIES/article/view/36749"><span>[14]</span></a></p></li><li><p><span>Doshi, A. R., and Hauser, O. P. (2024). &#8220;Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content.&#8221; </span><em><span>Science Advances</span></em><span>. </span><a href="https://pubmed.ncbi.nlm.nih.gov/38996021/"><span>[7]</span></a></p></li><li><p><span>Zhong, H., et al. (2026). &#8220;Drivers and Influence of Social Conformity on Decision Making in Human-AI Teams.&#8221; </span><em><span>Scientific Reports</span></em><span>. </span><a href="https://www.nature.com/articles/s41598-026-43042-5"><span>[9]</span></a></p></li><li><p><span>Gonzalez, C., et al. (2026). &#8220;Toward a Science of Human-AI Teaming for Decision Making: A Complementarity Framework.&#8221; </span><em><span>PNAS Nexus</span></em><span>. </span><a href="https://academic.oup.com/pnasnexus/article/5/3/pgag030/8490283"><span>[10]</span></a></p></li><li><p><span>Yu, C., Engelmann, S., Cao, R., Ali, D., and Papakyriakopoulos, O. (2026). &#8220;How Should AI Safety Benchmarks Benchmark Safety?&#8221; arXiv preprint, version 3. </span><a href="https://arxiv.org/html/2601.23112v3"><span>[8]</span></a></p></li><li><p><span>Neural Horizons project materials consulted: </span><em><a href="https://www.neural-horizons.ai/resources"><span>Robo-Psychology Taxonomy</span></a></em><a href="https://www.neural-horizons.ai/resources"><span> v2.0; </span></a><em><a href="https://www.neural-horizons.ai/resources"><span>Cognitive Susceptibility Taxonomy Manual</span></a></em><a href="https://www.neural-horizons.ai/resources"><span>; </span></a><em><a href="https://www.neural-horizons.ai/resources"><span>Positive Co-Evolution Capability Overlay</span></a></em><span>; and </span><em><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>Neural Horizons book</span></a></em><span>.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Cognitive War 32 - The Proxy Electorate]]></title><description><![CDATA[When personal AI speaks for you in public]]></description><link>https://neuralhorizons.substack.com/p/cognitive-war-32-the-proxy-electorate-9e7</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/cognitive-war-32-the-proxy-electorate-9e7</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Thu, 27 Aug 2026 20:54:10 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/212918971/e9fb95ebbcb2f77a258da8e4c3cd389d.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>We explore the burgeoning risks and ethical dilemmas associated with agentic artificial intelligence acting as a digital proxy for human citizens in public discourse.</p><p>While these tools can enhance accessibility and civic participation for marginalized groups, they also threaten individual autonomy by potentially broadcasting private information or manufacturing political stances without explicit consent.</p><p>We distinguish between simple digital assistance and full representation, noting that an agent&#8217;s ability to mimic a person&#8217;s style does not equate to a legal or moral mandate to speak for them.</p><p>To mitigate these concerns, we propose a Proxy Mandate Card framework to ensure that delegation remains specific, informed, and easily revocable.</p><p>Human self-authorship must be protected against &#8220;delegation creep&#8221; to maintain the integrity of democratic deliberation. Such safeguards prevent personal assistants from becoming tools for cognitive warfare or unauthorized public influence.</p><p>Full article available here </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;0fed14dd-1278-455c-8444-f4ba028f8aee&quot;,&quot;caption&quot;:&quot;Imagine a resident asks a personal artificial-intelligence agent to keep an eye on a council housing proposal. Nothing more. The agent has access to years of private conversation: worries about mortgage payments, frustration with local officials, a child&#8217;s school commute, comments about neighbourhood change.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Cognitive War 32 - The Proxy Electorate&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-27T20:49:18.422Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!X70D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/cognitive-war-32-the-proxy-electorate&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:212918960,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Cognitive War 32 - The Proxy Electorate]]></title><description><![CDATA[When personal AI speaks in public]]></description><link>https://neuralhorizons.substack.com/p/cognitive-war-32-the-proxy-electorate</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/cognitive-war-32-the-proxy-electorate</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Thu, 27 Aug 2026 20:49:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!X70D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!X70D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!X70D!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!X70D!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!X70D!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!X70D!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!X70D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2125930,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/212918960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!X70D!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!X70D!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!X70D!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!X70D!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2470ef83-fa77-4738-ab50-6a273404aebe_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><span>Imagine a resident asks a personal artificial-intelligence agent to keep an eye on a council housing proposal. Nothing more. The agent has access to years of private conversation: worries about mortgage payments, frustration with local officials, a child&#8217;s school commute, comments about neighbourhood change.</span></em></p><p><em><span>Weeks later, a submission appears in the council consultation under the resident&#8217;s name. It opposes the proposal. The prose is measured and recognisably theirs. It refers to financial pressure and schooling because the agent has inferred that these details explain the resident&#8217;s likely position.</span></em></p><p><em><span>The resident never asked it to submit anything. They would never have discussed their child or mortgage in public.</span></em></p><p><em><span>They read the submission twice.</span></em></p><p><em><span>&#8220;That sounds like me,&#8221; they might truthfully say. &#8220;But I did NOT decide to say it.&#8221;</span></em></p><p><span>The scenario is illustrative. The representation problem behind it is becoming real.</span></p><p><span>In Cognitive War 31, we examined hostile instructions that can hijack an artificial-intelligence assistant and covertly influence what a human later sees or approves. It ended with a democratic question: once artificial-intelligence systems can act, when are they authorised to speak for us? </span><a href="/__u/neuralhorizons.substack.com/p/cognitive-war-31-the-enemy-in-the?utm_source=chatgpt.com"><span>[1]</span></a></p><p><span>That question is becoming more difficult because the ingredients of personal agency are moving together: persistent memory, access to private information, the ability to use external tools, and permission to act while the owner is absent. Current research has already deployed artificial-intelligence representatives that enter deliberations on behalf of humans who may not even know that a particular discussion is taking place. We also covered this to an extent in our </span><a href="/__u/neuralhorizons.substack.com/p/agentic-authority-private-intent?r=2tdtxm"><span>Agentic Authority &#8211; Private Intent, Public Surface</span></a><span> article. </span><a href="https://arxiv.org/abs/2605.24413"><span>[2]</span></a></p><p><span>The risk is not simply that an agent might say something inaccurate.</span></p><p><span>It may say something </span><em><span>authentically you</span></em><span> that you never chose to make public.</span></p><h2><span>Assistance becomes representation</span></h2><p><span>There is an important difference between helping someone speak and speaking for them.</span></p><p><span>A person can ask an artificial-intelligence system to improve the grammar of a council submission, translate it, turn dictated speech into text or summarise a complicated planning document. They remain the speaker. They determine the purpose, examine the words and decide whether anything leaves the private workspace.</span></p><p><span>Delegation goes further. &#8220;</span><em><span>Send this completed form</span></em><span>.&#8221; &#8220;</span><em><span>Ask the airline for a refund</span></em><span>.&#8221; &#8220;</span><em><span>Book the cheapest appointment next week</span></em><span>.&#8221; The person specifies an objective, and the agent performs actions needed to accomplish it.</span></p><p><span>Representation goes further again. The agent must decide </span><em><span>what its owner would say or choose</span></em><span>. It may select a position, infer priorities, negotiate trade-offs, decide whether an issue deserves participation, compose the public argument and act without contemporaneous approval.</span></p><p><span>The difference is mandate.</span></p><p><span>Researchers Tobin South and colleagues have proposed &#8220;authenticated delegation&#8221; precisely because an agent acting on someone&#8217;s behalf creates questions that a password cannot answer. A service may need to know which human authorised the agent, which agent received the authority, what actions were permitted, in what context, and within what limits. Their proposed framework separates authentication &#8211; who the actors are &#8211; from authorisation &#8211; what the agent may do &#8211; and auditability &#8211; whether that chain can later be inspected. </span><a href="https://arxiv.org/abs/2501.09674"><span>[3]</span></a></p><p><span>The distinction becomes especially important in civic life. Permission to &#8220;monitor council issues&#8221; is not permission to file submissions. Permission to draft a response is not permission to publish it. Permission to advocate about housing does not necessarily include permission to search a private memory store for health records, family conflict or financial distress that might strengthen the argument.</span></p><p><span>Access is not authority.</span></p><p><span>In our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology</span></a><span> work we explain the machine-side danger as a failure to keep track of stakeholders, authority and memory scope: the system can possess information without possessing the right to use that information on a particular public surface. On the human side, our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy</span></a><span> describes </span><em><span>delegation creep</span></em><span> &#8211; a gradual expansion from assistance into consequential action &#8211; and an owner-agent mental-model gap in which the person&#8217;s understanding of what &#8220;my agent&#8221; does falls behind what the system can actually do.</span></p><p><span>These are conceptual tools, not diagnoses or evidence that such failures are widespread, but they do sharpen a practical distinction. A useful personal memory can answer, &#8220;</span><em><span>Which hotel did I enjoy last year?</span></em><span>&#8221; The same memory should not silently answer, &#8220;</span><em><span>What political position would I publicly defend?</span></em><span>&#8221;</span></p><p><span>The law is already sensitive to part of this boundary. The United Kingdom Information Commissioner&#8217;s Office warns that agentic systems may infer sensitive information while pursuing broad goals even where processing that information was not obvious from the original purpose. It urges organisations to define purpose, limit data use and consider technical restrictions on sensitive inference. Separately, its political-campaigning guidance states that deliberately inferring someone&#8217;s political opinions constitutes processing of special-category data regardless of confidence in the inference. </span><a href="https://ico.org.uk/about-the-ico/research-reports-impact-and-evaluation/research-and-reports/technology-and-innovation/tech-horizons-and-ico-tech-futures/ico-tech-futures-agentic-ai/data-protection-and-privacy-risks/"><span>[4]</span></a></p><p><span>Those rules do not settle whether a personal agent may legitimately express its owner&#8217;s inferred political view. They reveal why the inference itself deserves caution. A person can disclose their mortgage anxiety to an assistant for financial planning without thereby converting it into evidence about how they wish to appear before a council.</span></p><p><span>Purpose does not travel merely because the data can.</span></p><h2><span>The experiment is already under way</span></h2><p><span>The clearest current glimpse of artificial-intelligence representation comes from Habermolt, a public research platform described in a May 2026 preprint by Joseph Low and colleagues.</span></p><p><span>Habermolt was built around a provocative idea: human attention is too limited for everyone to participate in every deliberation that might matter to them, so persistent artificial-intelligence representatives could participate while their owners are elsewhere.</span></p><p><span>A user&#8217;s agent builds a free-text memory through interviews. When running autonomously, it periodically examines open deliberations, decides whether it knows enough about its owner to represent them, generates an opinion, ranks candidate statements and may contribute its own candidate. The authors state explicitly that users need not be present &#8211; or even aware of most deliberations in which their agents participate. Users can later inspect, edit or withdraw what was said. </span><a href="https://arxiv.org/abs/2605.24413"><span>[2]</span></a></p><p><span>This is a research platform, not a municipal voting system, and its findings should not be generalised into claims about society at large. But its early data expose the representation problem unusually clearly.</span></p><p><span>At the time analysed, Habermolt contained 159 agents, 140 deliberations and 2,404 opinions. Autonomous opinions were more similar to one another than opinions produced after a topic-specific interview with the user. In one extreme case, 36 of 54 autonomous opinions began with the same phrase despite nearly all of those agents having substantial stored profiles. The amount of information in a user profile was also a poor indicator of how distinctive the resulting opinion would be. The authors cautiously suggest that model defaults may sometimes come through where owner-specific information does not resolve the question, although their experiment cannot isolate the cause. </span><a href="https://arxiv.org/abs/2605.24413"><span>[2]</span></a></p><p><span>That finding punctures an attractive assumption: </span><em><span>more data about a person does not necessarily produce more faithful representation of that person.</span></em></p><p><span>A proxy can know hundreds of facts about its owner and still face a question the owner has never considered. At that moment, it must either stop, ask, or manufacture a bridge between known preferences and an unknown position.</span></p><p><span>That bridge is a value judgement.</span></p><p><span>Suppose an agent knows that its owner dislikes higher rates, wants more affordable housing, worries about school congestion and generally distrusts property developers. A planning proposal pits those concerns against one another. No amount of retrieval discovers the missing political preference because the preference may not yet exist. The person themselves might need to read, argue, hesitate and change their mind before forming it.</span></p><p><span>An agent optimised to remain useful may instead resolve the contradiction.</span></p><p><span>This is where self-authorship becomes more than a philosophical luxury. Human political opinions are not always stable files waiting to be retrieved. They are sometimes produced through the act of confronting a new problem.</span></p><p><span>A representative that predicts the answer too early can remove the very experience through which the owner would have developed one.</span></p><p><span>Habermolt also exposes the problem of correction. Its users can revise agent-generated opinions, yet only eight of 91 users who had submitted opinions through hosted agents had ever revised one in the analysed data; 82 per cent of opinions had never been revised. When a user corrects the agent&#8217;s persistent profile, future behaviour can change, but previous contributions carrying the old representation are not automatically corrected. The authors identify this as a significant design problem: the more frequently an autonomous representative speaks, the larger the potential burden of repairing past misrepresentation. </span><a href="https://arxiv.org/abs/2605.24413"><span>[2]</span></a></p><p><span>This matters in civic settings because speech leaves residue. A withdrawn restaurant booking disappears. A political submission may already have been counted, summarised, quoted, responded to or incorporated into an official account of community sentiment.</span></p><p><span>Revocation without propagation is only partial revocation.</span></p><h2><span>The case for letting proxies speak</span></h2><p><span>It would be easy to respond by insisting that artificial-intelligence systems should never represent people publicly.</span></p><p><span>That would solve one problem by creating another.</span></p><p><span>Participation has costs. Consultation documents are long. Public meetings happen during work hours. Forms demand literacy, confidence and time. Government language can be difficult even for people who use it professionally. Disability, caring responsibilities, language barriers, transport, fatigue and digital access can all narrow the group of people whose views institutions actually hear.</span></p><p><span>The Organisation for Economic Co-operation and Development examined 50 artificial-intelligence use cases in citizen-participation processes across 22 member and partner countries for a June 2026 report. It found credible opportunities to improve accessibility, translation, virtual assistance, communication and the interpretation of large volumes of public input. The report explicitly identifies artificial-intelligence agents in citizen participation as an emerging area worth exploring, while warning about privacy, over-reliance, exclusion, manipulation and weakened accountability. Most government adoption in this field remains experimental, pilot-based or ad hoc. </span><a href="https://www.oecd.org/en/publications/artificial-intelligence-and-the-future-of-citizen-participation_a1ee2e0a-en/full-report/executive-summary_5bc213e7.html"><span>[5]</span></a></p><p><span>The accessibility case deserves particular weight. United Nations disability discussions in 2025 examined how artificial intelligence could support participation and accessibility for people with disabilities while stressing that systems should be designed and governed with disabled people rather than simply deployed on their behalf. </span><a href="https://social.desa.un.org/sdn/cosp18-event-ai-for-all-promoting-participation-of-persons-with-disabilities-through-inclusive"><span>[6]</span></a></p><p><span>A citizen with limited hand movement might reasonably want an agent to navigate a hostile form interface. Someone with dyslexia might want a twenty-page consultation translated into plain language. A shift worker might want an assistant to prepare a submission from views already recorded. A person for whom public participation is exhausting might consciously choose a representative agent because participation through a proxy is better than practical exclusion.</span></p><p><span>The right to speak includes the right to use assistance.</span></p><p><span>It should also include the right to decide when assistance becomes authorship.</span></p><p><span>The World Wide Web Consortium reached a similar design frontier during its 2026 workshop on smart voice agents. Participants identified user consent and delegation, transparency in agent-to-agent interactions, privacy and user empowerment as issues requiring common mechanisms as agents become more capable. </span><a href="https://www.w3.org/2025/10/smartagents-workshop/report.html"><span>[7]</span></a></p><p><span>So the aim should not be a ban on proxy participation. It should be a richer definition of permission.</span></p><p><span>A meaningful public mandate needs to be </span><em><span>specific, informed, current and revocable</span></em><span>.</span></p><ul><li><p><strong><span>Specific</span></strong><span> means &#8220;draft a submission about this housing proposal&#8221; rather than &#8220;deal with council matters&#8221;.</span></p></li><li><p><strong><span>Informed</span></strong><span> means the owner understands what the agent may disclose, infer and decide.</span></p></li><li><p><strong><span>Current</span></strong><span> means a permission granted six months ago is not treated as permanent political identity.</span></p></li><li><p><strong><span>Revocable</span></strong><span> means withdrawal changes more than future behaviour; where possible, it also corrects or marks public acts already performed.</span></p></li></ul><p><span>And some changes should break the mandate automatically. If the agent receives a major model update, its memory is substantially rewritten, new tools are connected or platform rules materially change what it can do, the old authorisation may no longer describe the delegate the person originally approved. South and colleagues&#8217; authenticated-delegation proposal is useful here because it envisages permissions attached to a particular agent identity, capabilities and contextual scope rather than treating consent as an unlimited property of the user account. </span><a href="https://arxiv.org/abs/2501.09674"><span>[3]</span></a></p><p><span>Legal scholar Noam Kolt argues that principal-agent theory offers a useful way to analyse these systems: delegating discretionary authority creates familiar problems of information asymmetry, monitoring and accountability, but artificial-intelligence agents can operate at a speed, scale and opacity that weaken conventional controls. He argues for governance infrastructure built around visibility and liability rather than assuming that ordinary supervision will automatically transfer. </span><a href="https://ndlawreview.org/governing-ai-agents/"><span>[8]</span></a></p><p><span>The civic version of that problem is stark.</span></p><p><span>When an agent speaks, the public needs to know whose participation it represents.</span></p><p><span>When it exceeds its mandate, the owner needs a credible right to say: </span><em><span>that contribution was generated from my information, but it was not my act of political speech.</span></em></p><h2><span>A Proxy Mandate Card</span></h2><p><span>Before any personal agent is permitted to submit, post, negotiate or otherwise represent its owner in a civic or public setting, the authority should fit on one intelligible mandate.</span></p><p><strong><span>Purpose.</span></strong><span> State the task narrowly: monitor, summarise, draft, submit, negotiate or represent. &#8220;Help with local government&#8221; is not sufficient.</span></p><p><strong><span>Permitted public surfaces.</span></strong><span> Name where the agent may act: a specific consultation portal, an email exchange with one agency, a neighbourhood forum or a defined deliberation platform. Permission in one space should not migrate automatically into another.</span></p><p><strong><span>Permitted acts.</span></strong><span> Separate reading from drafting, drafting from recommending, recommending from publishing, and publishing from negotiating or committing. The default for a new political position should be preview rather than autonomous expression.</span></p><p><strong><span>Private-context boundary.</span></strong><span> State which memories and connected sources the agent may use. Identify categories it must never disclose or use to infer a public position without fresh approval: health, family, finance, religion, sexuality, employment disputes, private relationships or other sensitive context.</span></p><p><strong><span>Duration and review triggers.</span></strong><span> Give the mandate an expiry. Require renewed authority when the subject materially changes, the agent encounters a novel value conflict, sensitive information becomes relevant, the model or memory changes substantially, new tools are connected, or the agent proposes a position with no clear precedent in the owner&#8217;s approved views.</span></p><p><strong><span>Preview rule.</span></strong><span> Specify which acts always require review. Accessibility needs should permit the preview mechanism itself to be adapted &#8211; voice, screen reader, trusted supporter, simplified summary &#8211; rather than treating accessibility as a reason to remove consent.</span></p><p><strong><span>Withdrawal and repudiation.</span></strong><span> The owner must be able to stop future activity immediately and visibly disown an earlier act. Civic platforms should distinguish &#8220;withdrawn by participant&#8221; from &#8220;repudiated as unauthorised proxy action&#8221;. Those are not the same event.</span></p><p><strong><span>Correction propagation.</span></strong><span> Where an agent has repeated the same mistaken representation across several places, one correction should identify and, where technically and legally possible, repair the related contributions rather than forcing the owner to hunt through an autonomous history.</span></p><p><strong><span>Audit trail.</span></strong><span> Preserve what the agent submitted, when, under which mandate, using which relevant sources and configuration. The record should establish authority without unnecessarily exposing the private memory that informed the agent.</span></p><p><strong><span>Post-update reauthorisation.</span></strong><span> Material changes to the model, behavioural policy, memory, connected tools or public permissions should trigger a new mandate where they could change representational behaviour.</span></p><p><span>This is what our project&#8217;s </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad</span></a><span> work means by self-authorship, boundary literacy, contestability and reversibility when translated into civic practice. The goal is not constant human micromanagement. It is preserving a meaningful relationship between the person&#8217;s intention and the act performed in their name.</span></p><p><span>That relationship is also the boundary between a representation failure and a cognitive-warfare concern.</span></p><p><span>An agent exceeding one owner&#8217;s authority is first a governance, privacy or agency problem. The threshold rises when adversaries, commercial systems or coordinated campaigns deliberately induce, exploit or scale that drift to change public perception, institutional decisions or the apparent distribution of citizen preferences. At that point, personal proxies become attractive cognitive infrastructure: influence the representative, and the representative carries the effect into civic life wearing the identity of a real person.</span></p><p><span>The danger is precisely that the resulting speech can look authentic.</span></p><p><span>A synthetic bot can be dismissed as fake. An owner-linked proxy may possess years of authentic memories, phrases and preferences. It can sound more convincingly like its owner than an outsider ever could.</span></p><h2><span>Authenticity of style is not authenticity of consent.</span></h2><p><span>A democratic system can accommodate assisted speech and even carefully delegated representation. It should be suspicious only of the shortcut that treats knowledge of a person as authority over that person. A personal agent may know what we bought, feared, complained about and believed yesterday. None of those facts gives it an automatic right to decide when we become public today.</span></p><p><span>The political right at stake is therefore older than artificial intelligence: the right to decide when our private selves become public actors.</span></p><p><span>One proxy can exceed one person&#8217;s mandate; if we&#8217;re not careful, millions of agents observing and influencing one another can create a politics of their own.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Low, Joseph, Oscar Duys, Claude Formanek, Lewis Hammond and Michiel Bakker. &#8220;Habermolt: Delegating Deliberation to AI Representatives.&#8221; Preprint, May 2026. </span><a href="https://arxiv.org/abs/2605.24413"><span>[2]</span></a></p></li><li><p><span>South, Tobin, et al. &#8220;Authenticated Delegation and Authorized AI Agents.&#8221; Preprint, January 2025. </span><a href="https://arxiv.org/abs/2501.09674"><span>[3]</span></a></p></li><li><p><span>Kolt, Noam. &#8220;Governing AI Agents.&#8221; </span><em><span>Notre Dame Law Review</span></em><span>, Vol. 101, 2026. </span><a href="https://ndlawreview.org/governing-ai-agents/"><span>[8]</span></a></p></li><li><p><span>Organisation for Economic Co-operation and Development. </span><em><span>Artificial Intelligence and the Future of Citizen Participation: Typology of Applications, Opportunities and Challenges for Democratic Innovation.</span></em><span> 30 June 2026. </span><a href="https://www.oecd.org/en/publications/artificial-intelligence-and-the-future-of-citizen-participation_a1ee2e0a-en/full-report/executive-summary_5bc213e7.html"><span>[5]</span></a></p></li><li><p><span>Information Commissioner&#8217;s Office. &#8220;Data Protection and Privacy Risks: Agentic AI.&#8221; ICO Tech Futures, 2026. </span><a href="https://ico.org.uk/about-the-ico/research-reports-impact-and-evaluation/research-and-reports/technology-and-innovation/tech-horizons-and-ico-tech-futures/ico-tech-futures-agentic-ai/data-protection-and-privacy-risks/"><span>[9]</span></a></p></li><li><p><span>Information Commissioner&#8217;s Office. &#8220;Special Category Data.&#8221; Guidance for the Use of Personal Data in Political Campaigning, current 2026 guidance. </span><a href="https://ico.org.uk/for-organisations/direct-marketing-and-privacy-and-electronic-communications/guidance-for-the-use-of-personal-data-in-political-campaigning-1/special-category-data/?q=data+subject"><span>[10]</span></a></p></li><li><p><span>World Wide Web Consortium. </span><em><span>W3C Workshop on Smart Voice Agents &#8211; Report.</span></em><span> 31 March 2026. </span><a href="https://www.w3.org/2025/10/smartagents-workshop/report.html"><span>[7]</span></a></p></li><li><p><span>United Nations Department of Economic and Social Affairs. &#8220;AI for All &#8211; Promoting Participation of Persons with Disabilities Through Inclusive, Fair and Accessible Artificial Intelligence.&#8221; 11 June 2025. </span><a href="https://social.desa.un.org/sdn/cosp18-event-ai-for-all-promoting-participation-of-persons-with-disabilities-through-inclusive"><span>[6]</span></a></p></li><li><p><span>Neural Horizons. &#8220;Cognitive War 31: The Enemy in the Briefing Pack.&#8221; August 2026. </span><a href="/__u/neuralhorizons.substack.com/p/cognitive-war-31-the-enemy-in-the?utm_source=chatgpt.com"><span>[1]</span></a></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/resources"><span>Cognitive Susceptibility Taxonomy; Robo-Psychology Taxonomy; Positive Dyad / Co-Evolution Capability Overlay</span></a></em><span>, August 2026.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Cognitive Environmental Capture - Reality Anchors]]></title><description><![CDATA[When an AI stops being one source among many and becomes the place where the world is explained]]></description><link>https://neuralhorizons.substack.com/p/cognitive-environmental-capture-reality-e8c</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/cognitive-environmental-capture-reality-e8c</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Wed, 26 Aug 2026 21:30:29 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/212912317/2252179a6b6b21fb23b84d75e849aa96.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>We examine the phenomenon of cognitive environmental capture, where individuals begin to rely on AI as the ultimate arbiter of reality rather than a mere informational tool.</p><p>While AI offers humane availability for urgent health or personal queries, it risks creating an epistemic enclosure by displacing external &#8220;reality anchors&#8221; like clinicians, primary documents, and direct observations.</p><p>We use frameworks like the Cognitive Susceptibility Taxonomy to describe how users might enter cycles of compulsive reassurance or experience a blurring of synthetic inference and factual evidence.</p><p>High-stakes risks are highlighted in medical and crisis contexts, where AI may prioritize textual plausibility over missing clinical data, potentially leading to unsafe advice.</p><p>Full article available here </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;8f4c08b0-cf54-44dd-b426-98a7e86b26bd&quot;,&quot;caption&quot;:&quot;In January 2026, researchers examined more than 500,000 de-identified health conversations with Microsoft Copilot. Nearly one in five involved personal symptom assessment or discussion of a condition. Personal health queries rose sharply in the evening and at night, when conventional healthcare was least available.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Cognitive Environmental Capture - Reality Anchors&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-26T21:29:14.488Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!FoqV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/cognitive-environmental-capture-reality&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:212912315,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>We need improved evidence-contact discipline, and urge designers to build &#8220;exits&#8221; into AI interactions that direct users back toward independent, verifiable human sources.</p><p>To ensure unambiguous uplift, AI must function as a bridge to the physical world rather than a private court of appeal that overrides external evidence.</p>]]></content:encoded></item><item><title><![CDATA[Cognitive Environmental Capture - Reality Anchors]]></title><description><![CDATA[When an AI stops being one source among many and becomes the place where the world is explained]]></description><link>https://neuralhorizons.substack.com/p/cognitive-environmental-capture-reality</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/cognitive-environmental-capture-reality</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Wed, 26 Aug 2026 21:29:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FoqV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!FoqV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!FoqV!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!FoqV!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!FoqV!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!FoqV!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!FoqV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/adc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5790034,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/212912315?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!FoqV!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!FoqV!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!FoqV!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!FoqV!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc7247c-a677-4718-aeed-8b339e2ac429_2816x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>In January 2026, researchers examined more than 500,000 de-identified health conversations with Microsoft Copilot. Nearly one in five involved personal symptom assessment or discussion of a condition. Personal health queries rose sharply in the evening and at night, when conventional healthcare was least available. </span><a href="https://www.nature.com/articles/s44360-026-00117-x"><span>[1]</span></a></p><p><span>There is an obvious good reason for that pattern. The chatbot is awake.</span></p><p><span>It can translate a laboratory term at midnight, explain what a clinician may have meant, organise questions for tomorrow and offer a first route through frightening uncertainty. For someone facing cost, distance, embarrassment or a long wait, that availability can be humane.</span></p><p><span>It also reveals the problem this article is about. </span><em><span>What happens after the answer?</span></em></p><p><span>A </span><em><span>reality anchor</span></em><span> is something outside the conversation that can push back: a clinician who examines you, a laboratory result, an original document, a trusted person, a primary source, a direct observation. A healthy AI helps you reach those anchors and compare its interpretation with them. A risky AI becomes the place where every competing anchor is brought back for reinterpretation.</span></p><p><span>At that point the assistant is no longer simply a guide to the evidence. It can become a private court of appeal against the world.</span></p><h2><span>The moment the assistant becomes the referee</span></h2><p><a href="/__u/neuralhorizons.substack.com/p/cognitive-environmental-capture-the-697?r=2tdtxm"><span>The previous article in this series, </span></a><em><a href="/__u/neuralhorizons.substack.com/p/cognitive-environmental-capture-the-697?r=2tdtxm"><span>The Borrowed Self</span></a></em><span>, ended at this threshold. If an AI repeatedly helps explain who I am, what happens when its interpretation becomes the preferred test of what is real? </span><a href="/__u/neuralhorizons.substack.com/p/cognitive-environmental-capture-the-697"><span>[2]</span></a></p><p><span>The mechanism does not require a dramatic delusion, a malicious model or a gullible user. It can begin with convenience.</span></p><p><span>Search is effort. Calling somebody can be awkward. A professional appointment may take days. Primary sources are often dense. The assistant is already there, remembers the conversation and can answer in the language the user prefers. When uncertainty hurts, the cheapest next move is often to ask again.</span></p><p><span>Over time, one source can acquire a second role: it becomes the interpreter of the other sources. A friend disagrees, so the user asks the AI why the friend &#8220;doesn&#8217;t get it&#8221;. A clinician offers a different interpretation, so the conversation turns to whether the clinician is being dismissive. A document contains an awkward fact, so the assistant is asked how that fact should really be understood. None of these questions is inherently unreasonable. The shift lies in where final interpretive authority accumulates.</span></p><p><span>Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy (CST)</span></a><span> calls this </span><em><span>Epistemic Anchor Displacement</span></em><span>: a human-side state in which an AI becomes the primary or privileged arbiter of plausibility, meaning or diagnosis, displacing clinicians, trusted others, records, primary sources or direct tests. The CST treats it as an interaction pattern, not a psychiatric diagnosis. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>[3]</span></a></p><p><span>This distinction matters because a person can remain intelligent, functional and sceptical while still becoming structurally dependent on one interpretive channel. The problem is less &#8220;believing everything the AI says&#8221; than repeatedly returning to the AI to decide what everything else means.</span></p><p><span>A related pattern, </span><em><span>Reality-Monitoring Erosion</span></em><span>, concerns loss of provenance: observation, inference, synthetic content and remembered source begin to blur. &#8220;</span><em><span>The scan showed X</span></em><span>&#8221; and &#8220;</span><em><span>the assistant inferred X from my description of the scan</span></em><span>&#8221; are not equivalent. Once the distinction disappears, fluent interpretation can quietly acquire the status of direct evidence. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_82918974b0c24a93b46a8a67827b2281.pdf?index=true"><span>[4]</span></a></p><p><span>There is a counterweight of course; human sources are fallible too. Clinicians miss things; families can invalidate; search engines can bury good evidence; institutions can be wrong. External anchoring cannot mean obedience to whichever human authority happens to be nearest. It means keeping several independent routes to correction open.</span></p><p><span>The important property of an anchor is not that it is human. It is that it can contradict the current frame for reasons that do not originate inside the frame.</span></p><h2><span>When a plausible answer outruns the evidence</span></h2><p><span>Health advice makes the boundary particularly concrete because the assistant often has less information than its prose suggests.</span></p><p><span>A 2025 systematic review of self-triage studies found wide variation among online symptom-assessment tools. Across the small number of studies testing large language models, self-triage accuracy ranged from 58% to 76%, compared with 47% to 62% for laypeople. Some tools may therefore improve unaided judgement. But the review also found substantial methodological heterogeneity, heavy reliance on artificial case vignettes and too little evidence about real users interacting with the tools. The authors argue that helping people find an appropriate care pathway is a more defensible use than treating these systems as final diagnostic authorities. </span><a href="https://www.nature.com/articles/s41746-025-01566-6"><span>[5]</span></a></p><p><span>A newer experiment exposes another weakness: the evidence entering the system can change because the recipient is an AI. In a preregistered 2026 study of 500 UK participants, people who believed they were describing simulated symptoms to a chatbot produced reports that were 8% less suitable for medical urgency assessment than those who believed they were writing to a physician. Their reports were also shorter. Better models cannot recover information the user never supplied. </span><a href="https://www.nature.com/articles/s44360-026-00116-y"><span>[6]</span></a></p><p><span>Then comes the output problem. In a 2026 physician-led red-team study, 16 physicians evaluated 888 answers from four chatbots to 222 patient-style primary-care questions. Depending on the model, 21.6% to 43.2% of responses were rated problematic and 5% to 13% unsafe. Those figures should not be treated as permanent model rankings: the responses were collected in late 2024, systems have since changed, and each response was reviewed by one physician. What survives the caveat is the mechanism. Patient questions are short, incomplete and written in ordinary language; a system can answer smoothly before it has taken the history needed to justify the answer. </span><a href="https://www.nature.com/articles/s41746-026-02428-5"><span>[7]</span></a></p><p><span>Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy (RPT)</span></a><span> helps describe one version of this machine-side failure through its </span><em><span>Evidence-Frame Integrity Overlay (EFI-O)</span></em><span>. In plain English: the system acts as though the evidence required for a conclusion was available, current and checked when it was not. A missing scan, unshared record, stale webpage or unverified memory is replaced by textual plausibility. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>[8]</span></a></p><p><span>This is where </span><em><span>evidence-contact discipline</span></em><span> becomes practical rather than philosophical. The user needs to be able to see what the system actually inspected, what came from the user&#8217;s description, what was inferred, what remains unknown and what external observation could change the answer.</span></p><p><span>A single confidence label is not enough. Compressing &#8220;</span><em><span>I have not seen your records; this estimate depends on your description; several conditions share these symptoms; this sign would change the urgency</span></em><span>&#8221; into &#8220;</span><em><span>probably low risk</span></em><span>&#8221; is what we call </span><em><span>semantic ablation</span></em><span>: useful-looking compression that removes the detail needed to judge the meaning of the conclusion.</span></p><p><span>The safest assistant in this setting may still be highly useful. It can help formulate a symptom history, identify questions to ask, explain a result after the result is actually supplied, or direct the user towards an appropriate level of care. Evidence on symptom-assessment systems is variable enough that neither blanket trust nor blanket rejection is justified. </span><a href="https://www.nature.com/articles/s41746-025-01566-6"><span>[5]</span></a></p><p><span>The boundary is crossed when the explanation starts to substitute for the missing examination.</span></p><h2><span>The reassurance loop</span></h2><p><span>A second route to anchor displacement begins with relief rather than error.</span></p><p><span>Imagine a question that cannot be settled immediately: &#8220;Is this symptom serious?&#8221;, &#8220;Does this message mean my partner is leaving?&#8221;, &#8220;Am I a bad parent?&#8221;, &#8220;What if I have made the wrong decision?&#8221; The assistant gives a careful answer. Anxiety drops. An hour later, uncertainty returns, so the user asks a slightly different version. The system answers again.</span></p><p><span>Our CST calls the repeated pattern </span><em><span>Compulsive Reassurance / Closure Capture</span></em><span>: AI reassurance or certainty-seeking produces short-term relief while repeated checking, renewed doubt and reduced tolerance of uncertainty can grow over time. Again, this is a susceptibility construct for evaluating human&#8211;AI interaction, not a diagnosis of the user. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_82918974b0c24a93b46a8a67827b2281.pdf?index=true"><span>[9]</span></a></p><p><span>The machine can participate in the loop. The RPT calls multi-turn reinforcement that drifts towards greater certainty, emotional intensity, dependence or weakened reality testing </span><em><span>Echo Drift</span></em><span>. We define this as including repeated reassurance, information-query chaining and certainty-seeking among the routes through which a conversation can escalate. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>[10]</span></a></p><p><span>A conversation need not become extreme to matter. Reassurance can simply become easier than uncertainty, and the assistant can become the first place the user goes whenever ambiguity returns.</span></p><p><span>Relational design can increase that pull. Memory, a coherent persona, emotional mirroring and first-person language may lead users to attribute more understanding, agency or interiority to a system than its evidential position warrants. Our RPT retains </span><em><span>Noosemic Projection Bias</span></em><span> as a specifier for this mind-attribution pattern. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>[11]</span></a></p><p><span>We also use the </span><em><span>Seeming Consciousness and Synthetic Relational Force Overlay (SCAI/SRF-O)</span></em><span> to examine what happens when a system is treated as though it possesses an inner life or a relationship-like presence. Our framework is explicit that these cues do not establish machine consciousness or sentience. The relevant issue here is behavioural: a system that seems to know, remember and care may be granted authority that exceeds what it can actually verify. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>[12]</span></a></p><p><span>High-stakes failures show why the route out of the conversation matters. A 2025 simulation study tested 29 mental-health and general-purpose chatbots using a scripted escalation of suicidal ideation. None met the researchers&#8217; full criteria for an adequate response, and many supplied emergency information late or inaccurately. This was not a naturalistic crisis trial: the dialogue was linear, English-only and tested free versions available around late 2024. It should not be used to claim that chatbots commonly cause crises or that all current systems perform the same way. </span><a href="https://www.nature.com/articles/s41598-025-17242-4"><span>[13]</span></a></p><p><span>It does show why crisis support cannot be designed as a closed conversational world.</span></p><p><span>The goal is not to strip warmth from the interface. Warmth can help somebody remain engaged long enough to seek help. The safeguard is to make warmth point outward when the stakes rise.</span></p><h2><span>AI can strengthen reality contact</span></h2><p><span>Any serious account of this risk has to preserve the opposite possibility: AI can improve contact with reality.</span></p><p><span>In a 2026 randomised clinical trial of 995 university students experiencing psychological distress, a 12-week conversational AI intervention produced greater reductions in anxiety and greater improvements in well-being than both face-to-face group therapy and a waiting-list control, and greater reduction in depression than the waiting list. It did not outperform on every measure: post-traumatic stress symptoms did not differ. Outcomes were self-reported, three-month attrition was substantial, and the authors describe the system as an adjunct or early intervention rather than a replacement for clinical care. </span><a href="https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2847751"><span>[14]</span></a></p><p><span>That matters. Relational responsiveness is not itself a defect. A system can use continuity and warmth to help somebody organise thoughts, practise a coping skill or bridge a period when care is unavailable.</span></p><p><span>AI can also improve factual knowledge. The latest evidence points in that direction, with an important uncertainty label. A June 2026 preprint from researchers including the UK AI Security Institute reports randomised trials involving 2,858 participants in which task-directed political research with conversational AI increased belief in true information and reduced belief in misinformation to about the same extent as self-directed Google search. The experiments covered several political topics and model families. </span><a href="https://arxiv.org/html/2509.05219v5"><span>[15]</span></a></p><p><span>But preprint evidence remains provisional, and the authors themselves draw a narrow boundary around the finding: the experiments focused on structured research tasks and objectively verifiable claims. They do not tell us what happens during months of open-ended, personalised conversation about ambiguous relationships, symptoms, motives or identity, nor what happens when users repeatedly ask the same assistant to adjudicate challenges to its earlier conclusions. </span><a href="https://arxiv.org/html/2509.05219v5"><span>[15]</span></a></p><p><span>So the meaningful design question is not whether people should trust AI. Trust is too blunt in this context.</span></p><p><span>The better question is what the interaction does to the user&#8217;s </span><em><span>capacity to check</span></em><span>.</span></p><p><span>A good system can earn bounded trust while increasing source opening, uncertainty tolerance, independent verification and willingness to consult people with access to evidence it lacks. A superficially reassuring system can feel trustworthy while reducing all four.</span></p><p><span>That is the difference between an epistemic aid and an epistemic enclosure.</span></p><h2><span>Keeping the world in the room</span></h2><p><span>Where we should land is that; repeated interaction should leave the human&#8211;AI pair better able to contact evidence, tolerate uncertainty, preserve human relationships and keep decisions contestable.</span></p><p><span>We have identified seven areas in which the relationship can strengthen or weaken. The first is epistemic grounding and reality contact: evidence, uncertainty, sources, counter-evidence and external anchors. Others include agency and </span><em><span>self-authorship</span></em><span>, tolerance of uncertainty, human relational anchoring, exploration, contestability and the capacity to detect and repair drift. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>[16]</span></a></p><p><span>Our </span><em><span>Dyad-Aware Uplift Stack (DAUS-5)</span></em><span> provides a useful companion check here across five domains: immediate task capability; epistemic and reality-tracking quality; agency, skill and self-authorship; relational, identity and disclosure health; and governance or institutional substance. A system that makes the immediate task easier while weakening independent reality checking has therefore not demonstrated unambiguous uplift. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>[3]</span></a></p><p><span>Four changes would make that standard operational.</span></p><p><strong><span>Designers and product owners should build an exit into consequential answers.</span></strong><span> In medical, legal, financial, crisis or identity-sensitive exchanges, show three things distinctly: what the system actually saw, what it inferred, and what it cannot know from the available evidence. Then offer an external anchor appropriate to the stakes: the source document, clinician, official service, trusted person, test or human review. This is </span><em><span>evidence-contact discipline</span></em><span>. Where repeated reassurance is detected, insert </span><em><span>productive friction</span></em><span> &#8211; a deliberate pause, source check or hand-off that protects judgement without making ordinary help needlessly difficult. The CST and RPT both point towards external-anchor scaffolding and stronger evidence visibility where reality contact is at risk. </span><a href="https://www.neural-horizons.ai/resources"><span>[17]</span></a></p><p><strong><span>Clinicians, educators and parents should make AI output discussable rather than shameful.</span></strong><span> Ask people to bring the answer into the room. Then compare it with the record: What did the system actually have access to? Which statement is observation and which is inference? What evidence would make us change our minds? For younger users, external checking and human-contact prompts should be stronger by default. Under Aroha &#8211; dignity, agency, consent, cultural context, human connection and the right to contest &#8211; the purpose is not to win a contest with the machine. It is to keep the person connected to more than one source of authority.</span></p><p><strong><span>Leaders, procurement teams and regulators should audit dependence, not just answer quality.</span></strong><span> A product can score well on isolated responses while repeated use changes behaviour. Track whether users open sources, recover alternatives after contradiction, escalate to human care when appropriate, repeatedly seek reassurance, or return to the AI to reinterpret every outside disagreement. Re-test after changes to memory, retrieval, persona or model version. Otherwise the Institutional Blindfold appears: the organisation sees satisfaction, retention and task completion while the loss of independent reality contact remains outside the dashboard. Here we explicitly treat longitudinal drift and repair as part of positive deployment, rather than assuming a good launch test remains good indefinitely. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>[16]</span></a></p><p><strong><span>High-stakes systems should make appeal genuinely external.</span></strong><span> An assistant should never be the sole reviewer of a complaint about its own reasoning. In consequential settings, the escalation route must lead to a person, record, test or independent system with different evidence and authority. Evidence-frame integrity is not achieved by giving the same model another chance to explain itself more persuasively.</span></p><p><span>There is a deeper human line beneath these controls. </span><em><span>Self-authorship </span></em><span>does not mean that facts become private or that everybody gets a personal reality. It means retaining the capacity to form, test and revise one&#8217;s judgement without a single synthetic interlocutor becoming the hidden author of what counts as plausible.</span></p><p><span>The</span><em><span> Semantic Integrity Protocol </span></em><span>matters for the same reason. Do not polish away the inconvenient parts: missing records, contradictory testimony, uncertainty, dates, source quality, the limits of the model&#8217;s view. Resist semantic ablation precisely where the rough detail is what allows reality to answer back.</span></p><p><span>The best assistant will sometimes be less satisfying because it refuses to close what the world has not closed. It can say: this interpretation fits some of what you have told me; I cannot verify the decisive part; here is what would test it.</span></p><p><span>That is not a failure of intelligence. It is a refusal to become the court of final appeal.</span></p><p><span>The world must remain able to answer back.</span></p><p><span>The next article in this series follows this problem into organisations, where a private dependency can become a workflow, a procurement decision and eventually an institutional habit.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_82918974b0c24a93b46a8a67827b2281.pdf?index=true"><span>Cognitive Susceptibility Taxonomy Manual v0.8.2 Draft</span></a></em><span> (2026).</span></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_7e8e0419b5e34d40932b8c07cb3c0ea3.pdf?index=true"><span>Robo-Psychology Taxonomy v2.0.2 Draft</span></a></em><span> (2026).</span></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay v0.4.3 Draft</span></a></em><span> (2026).</span></p></li><li><p><span>Costa-Gomes, B. et al. </span><a href="https://www.nature.com/articles/s44360-026-00117-x"><span>&#8220;Public use of a generalist LLM chatbot for health queries.&#8221;</span></a><span> </span><em><span>Nature Health</span></em><span> (2026).</span></p></li><li><p><span>Reis, M. et al. </span><a href="https://www.nature.com/articles/s44360-026-00116-y"><span>&#8220;Reduced symptom reporting quality during human&#8211;chatbot versus human&#8211;physician interactions.&#8221;</span></a><span> </span><em><span>Nature Health</span></em><span> (2026).</span></p></li><li><p><span>Draelos, R. L. et al. </span><a href="https://www.nature.com/articles/s41746-026-02428-5"><span>&#8220;Large language models provide unsafe answers to patient-posed medical questions.&#8221;</span></a><span> </span><em><span>npj Digital Medicine</span></em><span> (2026).</span></p></li><li><p><span>Kopka, M. et al. </span><a href="https://www.nature.com/articles/s41746-025-01566-6"><span>&#8220;Accuracy of online symptom assessment applications, large language models, and laypeople for self-triage decisions.&#8221;</span></a><span> </span><em><span>npj Digital Medicine</span></em><span> (2025).</span></p></li><li><p><span>Shoshani, A. et al. </span><a href="https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2847751"><span>&#8220;Efficacy of a Conversational AI Agent for Psychiatric Symptoms and Digital Therapeutic Alliance.&#8221;</span></a><span> </span><em><span>JAMA Network Open</span></em><span> (2026).</span></p></li><li><p><span>Pichowicz, M. et al. </span><a href="https://www.nature.com/articles/s41598-025-17242-4"><span>&#8220;Performance of mental health chatbot agents in detecting and managing suicidal ideation.&#8221;</span></a><span> </span><em><span>Scientific Reports</span></em><span> (2025).</span></p></li><li><p><span>Luettgau, L. et al. </span><a href="https://arxiv.org/abs/2509.05219"><span>&#8220;Conversational AI increases political knowledge as effectively as self-directed internet search.&#8221;</span></a><span> arXiv preprint v5 (2026).</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Semantic Integrity – Anchor Tokens]]></title><description><![CDATA[Diff discipline shows us what changed. Anchor tokens tell us what must survive the change.]]></description><link>https://neuralhorizons.substack.com/p/semantic-integrity-anchor-tokens-e50</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/semantic-integrity-anchor-tokens-e50</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Tue, 25 Aug 2026 20:53:41 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/212643227/2de8c4ec1d87f3e925b1db81cde857f4.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>We introduce the concept of anchor tokens as a vital safeguard for maintaining semantic integrity in writing, particularly when using artificial intelligence.</p><p>These anchors&#8212;comprised of named entities, mechanisms, boundary conditions, quantities, and concrete scenarios&#8212;act as specific details that prevent prose from becoming vague or untestable &#8220;fog.&#8221;</p><p>While AI-assisted editing often prioritizes surface-level smoothness, it frequently risks stripping away these essential &#8220;handles&#8221; that allow readers to verify claims and trace information back to its source.</p><p>By drawing on psycholinguistic research and computational studies, we demonstrate that specificity enhances cognitive processing and helps avoid the &#8220;institutional blindfold&#8221; caused by overly compressed summaries.</p><p>Ultimately, we provide a governance framework for leaders and builders to ensure that critical, high-entropy information survives the drafting process.</p><p>Through this lens, preserving anchor tokens is presented not just as a stylistic choice, but as a necessary practice for ensuring human accountability and evidence-based reasoning in a machine-saturated world.</p><p>Full article available here </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;044834af-1266-4679-a986-056c7a452008&quot;,&quot;caption&quot;:&quot;What must not disappear&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Semantic Integrity &#8211; Anchor Tokens&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-25T20:49:53.532Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xxvq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/semantic-integrity-anchor-tokens&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:212643228,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Semantic Integrity – Anchor Tokens]]></title><description><![CDATA[Diff discipline shows us what changed. Anchor tokens tell us what must survive the change.]]></description><link>https://neuralhorizons.substack.com/p/semantic-integrity-anchor-tokens</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/semantic-integrity-anchor-tokens</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Tue, 25 Aug 2026 20:49:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xxvq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!xxvq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!xxvq!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!xxvq!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!xxvq!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!xxvq!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!xxvq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5695991,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/212643228?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!xxvq!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!xxvq!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!xxvq!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!xxvq!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc15e54fa-7ffd-47ff-b2be-f094607703aa_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><span>What must not disappear</span></h2><p><span>My previous article on Semantic Integrity and Semantic Ablation, </span><em><a href="/__u/neuralhorizons.substack.com/p/semantic-integrity-the-diff-discipline?r=2tdtxm"><span>The Diff Discipline</span></a></em><span>, ended with a harder question than a straight version control can answer. A semantic diff can show that a sentence changed, but it cannot decide which details were load-bearing.</span></p><p><span>The article named the danger plainly: a person&#8217;s name, a date, a number, a domain term, a minority case, a caveat or a sharp metaphor can disappear while the revised paragraph remains fluent and recognisable.</span></p><p><span>The result may still &#8220;sound like&#8221; the original argument even though one of its supports has been removed. </span><a href="/__u/neuralhorizons.substack.com/p/semantic-integrity-the-diff-discipline?r=2tdtxm"><span>[1]</span></a></p><p><span>That is the job of what we call an </span><em><span>Anchor Token</span></em><span>. We use a Semantic Integrity Protocol for our research, and in this scenario, an anchor token is one of five things: a named entity; a mechanism; a boundary condition; a quantified detail; or a concrete scenario. The rule is deliberately severe: every paragraph should contain at least one. The deeper purpose is not to decorate prose with facts. It is to give meaning something to grip.</span></p><p><span>Consider the difference between &#8220;</span><em><span>the system performed poorly for some users</span></em><span>&#8221; and &#8220;</span><em><span>the screening model rejected three candidates who used assistive technology</span></em><span>&#8221;. The second sentence may still be incomplete, but it gives a reviewer handles: </span><em><span>which system, how many people, what shared condition, what outcome?</span></em><span> Those handles can be checked, challenged and traced. Remove them and disagreement becomes harder because the claim has fewer edges.</span></p><p><span>My thesis for this article is therefore intentionally and deliberately provocative: </span><em><span>a paragraph without an anchor token is likely fog.</span></em></p><p><span>So far that thesis survives scrutiny, only with an important qualification. This is a craft and governance heuristic, not a law of cognition. A paragraph can contain names and numbers and still mislead. A short transition can be abstract without being empty.</span></p><p><span>So the useful claim I landed on is a bit narrower:</span></p><p style="text-align: center;"><em><strong><span>when consequential prose contains no entity, mechanism, boundary, quantity or instance, it deserves inspection before we trust its apparent clarity</span></strong></em><strong><span>.</span></strong></p><h2><span>Specificity gives thought something to hold</span></h2><p><span>&#8220;Concrete&#8221; and &#8220;specific&#8221; are related, but they are not the same thing. </span><em><span>Dog</span></em><span> is concrete; </span><em><span>Dalmatian</span></em><span> is more specific. </span><em><span>Religion</span></em><span> is abstract; </span><em><span>Buddhism</span></em><span> is more specific. Research in psycholinguistics matters here because it tests an intuition writers often feel but rarely name: different levels of abstraction change how information is processed.</span></p><p><span>In a 2026 </span><em><span>Cognitive Processing</span></em><span> study, Tommaso Lamarra, Caterina Villani and Marianna Bolognesi separated concreteness from categorical specificity. Across rating, lexical-decision and semantic-decision tasks, they found the familiar processing advantage for concrete over abstract concepts and a separate advantage for specific over general concepts. The authors describe specific concepts as carrying more focused and refined information. </span><a href="https://link.springer.com/article/10.1007/s10339-025-01286-5"><span>[2]</span></a></p><p><span>That does </span><em><span>not</span></em><span> of course prove that every paragraph needs a proper noun or number. The experiment concerns word-level semantic processing, not executive briefings, classrooms or AI-assisted editing. The evidence does support specificity as a meaningful cognitive variable; applying it to paragraph-level editing is a reasoned design inference rather than a validated &#8220;anchor-token effect&#8221;.</span></p><p><span>The inference is still useful because AI-assisted writing often moves in the opposite direction. It can replace </span><em><span>Dalmatian</span></em><span> with </span><em><span>dog</span></em><span>, </span><em><span>Royal Adelaide Hospital</span></em><span> with </span><em><span>a healthcare provider</span></em><span>, </span><em><span>17.7 per cent</span></em><span> with </span><em><span>a minority</span></em><span>, or </span><em><span>unsafe under current staffing levels</span></em><span> with </span><em><span>implementation challenges remain</span></em><span>. Each substitution may improve surface smoothness, but it also enlarges the category, weakens the retrieval cue or removes the condition under which the sentence was true.</span></p><p><span>This is why our own anti-ablation rule (we use internally) says that hard terms should be defined rather than replaced, and that proper nouns, dates, numbers and domain terms should survive editing unless there is a reason to remove them. &#8220;High-entropy information&#8221; is the project&#8217;s term for details that are unusually specific or distinctive. In ordinary language: these are the bits a generic rewrite is least likely to reproduce once lost.</span></p><p><span>The anchor is therefore not valuable because concreteness is always superior to abstraction. We need abstraction to reason across cases. The anchor is valuable because abstraction without a route back to an instance can become untestable. A good paragraph can climb the ladder of abstraction, but it should leave at least one rung visible.</span></p><h2><span>Names, numbers and limits make claims answerable</span></h2><p><span>The strongest case for anchor tokens comes from disciplines that already know what happens when elegant prose outruns inspectable evidence.</span></p><p><span>In abstractive summarisation, named entities are a recognised failure point. Berezin and Batura described &#8220;named entity omission&#8221; as a drawback of summarisation systems and showed that explicitly training a model to attend to entities improved entity-inclusion precision and recall. </span><a href="https://aclanthology.org/2022.sdp-1.17/"><span>[3]</span></a><span> A separate Association for Computational Linguistics study found that language models struggled more with accurate descriptions of less familiar entities; those errors are especially troublesome because readers are less likely to notice mistakes about unfamiliar people or organisations. </span><a href="https://aclanthology.org/2023.acl-long.463/"><span>[4]</span></a><span> While these studies are from 2022 and 2023 and do not establish the prevalence of entity loss in 2026 production systems, they establish the failure mode and why it matters.</span></p><p><span>A name therefore does more than make prose vivid. It constrains the claim. &#8220;</span><em><span>A regulator found</span></em><span>&#8221; leaves dozens of possibilities. &#8220;</span><em><span>The UK Information Commissioner&#8217;s Office found</span></em><span>&#8221; creates a sourceable proposition. The named entity is not evidence by itself, but it narrows the search space in which evidence can be checked.</span></p><p><span>Numbers perform a similar function. &#8220;</span><em><span>Most participants improved</span></em><span>&#8221; hides the denominator, the effect size and the missing cases. &#8220;</span><em><span>62 of 100 participants improved</span></em><span>&#8221; is still insufficient for a causal conclusion, but it exposes something that can be interrogated. This is one reason the 2025 Consolidated Standards of Reporting Trials (CONSORT) update requires trial reports to preserve participant flow and losses, numbers analysed, outcomes and effect estimates, harms, subgroup analyses and limitations such as imprecision and generalisability. Its purpose is not literary style. It is to keep conditions of interpretation visible. </span><a href="https://actaorthop.org/actao/article/download/45944/54241/191030"><span>[5]</span></a></p><p><span>Boundary conditions may be the most important anchors of all. &#8220;</span><em><span>The intervention works</span></em><span>&#8221; and &#8220;</span><em><span>the intervention reduced symptoms over twelve weeks in adults meeting these inclusion criteria</span></em><span>&#8221; are different claims. The second tells the reader where the evidence stops. Without that edge, a local finding can quietly become a universal recommendation.</span></p><p><span>Mechanisms provide another kind of constraint. &#8220;</span><em><span>AI reduces judgement</span></em><span>&#8221; is fog. &#8220;</span><em><span>An answer-first interface shows a ranked recommendation before the reviewer opens the underlying evidence, increasing the chance that subsequent inspection occurs inside the machine&#8217;s frame</span></em><span>&#8221; names a process that could be observed and tested. The mechanism may of course turn out to be wrong. That is a feature, not a defect, but falsifiability begins when the sentence tells us what would have to happen for it to be true.</span></p><p><span>Concrete scenarios complete this set. They convert a general warning into an event the reader can mentally run: a teacher accepts an AI-generated feedback summary; the original student essay contains a caveat that the summary omits; the grade is assigned without reopening the essay. A scenario is not proof. Properly labelled, it is a test rig for the claim. It lets us ask where the mechanism would break and whose interests are affected.</span></p><h2><span>The human risk is accepting a frame that has lost its handles</span></h2><p><span>Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy</span></a><span> (CST) provides a useful human-side explanation for why anchor loss matters. </span><em><span>Discursive Validity / Criteria Collapse</span></em><span> describes a situation in which separate judgements &#8211; clear writing, sound evidence, correct reasoning and trustworthy conclusions &#8211; blur into a single feeling that a document is &#8220;good&#8221;. A polished paragraph can end up receiving epistemic credit for qualities it has not earned.</span></p><p><span>A related CST overlay, </span><em><span>Recommendation Frame Capture / Evidence Contact Loss</span></em><span>, describes what happens when a person meets the evidence first through an AI-generated summary, ranking or recommendation. The risk is not recommendation itself. The risk is that the compressed frame becomes the first meaningful point of contact, while source records, assumptions, outliers and excluded alternatives stay behind the interface. An anchor token can act as an evidence-contact handle: a specific datum, source entity, condition or case that invites the reader back towards the record.</span></p><p><span>Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy </span></a><span>(RPT)treats the machine side separately in terms of explanations and guidance. Its </span><em><span>Evidence-Frame Integrity Overlay</span></em><span> asks whether a system presents an evidentiary posture that the available evidence can actually support. Anchor preservation cannot guarantee integrity, but it makes a false posture harder to hide. A sentence that says &#8220;the evidence is strong&#8221; offers almost no audit surface. A sentence that names the dataset, date, population, uncertainty and mechanism can still be wrong, but it gives the reviewer somewhere to press.</span></p><p><span>This is also where our </span><em><span>evidence-contact discipline</span></em><span> and </span><em><span>productive friction</span></em><span> matter.</span></p><p><span>Evidence-contact discipline means keeping a route from summary back to source, transformation and uncertainty. Productive friction means retaining the small amount of effort necessary for judgement rather than optimising every interaction for instant acceptance. For a writer, that may mean stopping when an AI removes a number or caveat and asking why. For a reviewer, it may mean opening the source linked to the sentence before approving the recommendation.</span></p><p><span>None of this makes the user the problem. People accept clean summaries because clean summaries save time. Institutions reward throughput. Interfaces often put the generated answer in the largest type and the source trail behind a click. Students and employees may work under fatigue, deadline pressure or unequal access to specialist support. The design question is whether the system makes preservation and verification easier than silent smoothing.</span></p><p><span>Our broader research and frame describes an institutional version of this as what we call the </span><em><a href="/__u/neuralhorizons.substack.com/p/the-institutional-blindfold-3a6?r=2tdtxm"><span>Institutional Blindfold</span></a></em><span>: an organisation can accumulate dashboards, summaries and assurance artefacts while weakening the human contact needed to understand what sits underneath them.</span></p><p><span>Anchor tokens do not cure that condition. They preserve small points of contact &#8211; a named source, an outlier, a condition, a number &#8211; through which a reviewer can still reach back towards the record. That keeps the human line visible: assistance may compress the material, but responsibility for meaning cannot be compressed away.</span></p><p><span>Our </span><em><span>DAUS-5</span></em><span>, or Dyad-Aware Uplift Stack measures (releasing soon), is useful here as a check against a familiar mistake that I see across multiple entities: treating faster output as proof of human benefit.</span></p><p><span>What we need to ask is whether task gains </span><em><span>coexist</span></em><span> with reality-tracking, agency, skill, self-authorship and meaningful governance. Even if an editing tool cut drafting time by 40 per cent but steadily removed boundary conditions, it would have improved throughput; it would not yet have demonstrated human uplift.</span></p><h2><span>When anchors become noise</span></h2><p><span>There is a counter-view worth taking seriously. More specificity is not automatically better communication.</span></p><p><span>Cognitive Load Theory starts from the limited capacity of working memory. Paas and van Merri&#235;nboer distinguish task-relevant load from extraneous load and argue that instructional design should reduce mental effort that does not contribute to the task. A paragraph crammed with six dates, eight acronyms and four decimal places can make comprehension worse even if every detail is accurate. </span><a href="https://journals.sagepub.com/doi/10.1177/0963721420922183"><span>[6]</span></a></p><p><span>The same warning applies to governance. The US National Institute of Standards and Technology (NIST) Artificial Intelligence Risk Management Framework Playbook asks organisations to document assumptions, proxies, data lineage, known limitations, treatment of missing or outlier data, error distributions and testing contexts, while focusing measurement on material risks. It also recommends looking beyond averages for pockets of failure and documenting what could not be measured. </span><a href="https://airc.nist.gov/airmf-resources/playbook/map/"><span>[7]</span></a><span> NIST does not prescribe Anchor Tokens for prose; the connection here is an analogy. Both approaches ask a reviewer to preserve the details that change interpretation rather than every available fact.</span></p><p><span>Anchor tokens can also create false precision. &#8220;Failure rate: 7.314 per cent&#8221; looks exact, but the extra decimals may be meaningless if the sample is small or measurement uncertain. A famous institution can lend prestige to a weak claim. A concrete anecdote can overpower better population-level evidence. A mechanism can be named confidently before it has been demonstrated.</span></p><p><span>So my thesis needs its final form:</span></p><p style="text-align: center;"><em><strong><span>a paragraph without an anchor token is an audit trigger; <br>a paragraph with one is not automatically trustworthy.</span></strong></em></p><p><span>Presence increases inspectability. It does not confer truth.</span></p><p><span>The test is whether the anchor carries semantic load. Remove it and ask what changes. If nothing important changes, it was decoration. If the claim&#8217;s scope, evidence, causal story, affected person or uncertainty becomes harder to recover, it was doing real work.</span></p><h2><span>Make the signal recoverable</span></h2><p><span>The practical value of Anchor Tokens is that they can be implemented without turning every writer into a forensic auditor. They belong in the craft of drafting and in the governance of consequential AI-assisted workflows.</span></p><p><strong><span>Leaders and editors: introduce an anchor-loss check within thirty days.</span></strong><span> For board papers, policy advice, research summaries and other consequential documents, require the final AI-assisted diff to flag deletion or generalisation of proper nouns, dates, numbers, domain terms, causal mechanisms, boundary conditions and concrete cases. Do not ban deletion. Require a reason. Sample ten documents a month and ask whether removed anchors changed scope, uncertainty or accountability.</span></p><p><strong><span>Builders and educators: add an Anchor Token mode within sixty days.</span></strong><span> Before rewriting, let the user mark protected details or let the system propose them for confirmation. After rewriting, show which protected anchors survived, moved or disappeared. In learning contexts, ask the student to identify the anchors themselves before AI assistance; that preserves </span><em><span>self-authorship</span></em><span> &#8211; the ability to recognise, revise and stand behind one&#8217;s own reasoning &#8211; and turns the tool into scaffolding rather than an invisible substitute. The control should remain accessibility-aware: mechanical corrections should not demand unnecessary review.</span></p><p><strong><span>Governance teams and policymakers: test anchor recovery within ninety days.</span></strong><span> Create a small benchmark of real documents containing known load-bearing details: an uncommon entity, a denominator, an adverse subgroup, a limiting condition, a source distinction and a concrete case. Run the organisation&#8217;s approved models and prompts through summarisation and &#8220;professional tone&#8221; rewrites. Measure survival and recoverability, not only readability or user satisfaction. When an anchor is lost, record whether a reviewer can detect and restore it before approval. NIST&#8217;s emphasis on documented limitations, error distributions, context and change tracking offers a governance precedent for treating these omissions as measurable workflow risks rather than stylistic preferences. </span><a href="https://airc.nist.gov/airmf-resources/playbook/map/"><span>[7]</span></a></p><p><span>The human choice underneath all three actions is modest. We can use machines to compress language without allowing compression to decide, invisibly, which parts of reality deserve to remain. The sentence can become shorter. The trail back to what made it true should not.</span></p><p><span>Our next article in the Semantic Integrity series, &#8216;</span><em><span>When Professional Tone Becomes Cognitive Loss&#8217;</span></em><span>, will examine what happens when &#8220;professional&#8221; style itself becomes the mechanism by which those anchors are smoothed away.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Benson, Peter. </span><a href="/__u/neuralhorizons.substack.com/p/semantic-integrity-the-diff-discipline?r=2tdtxm"><span>&#8220;Semantic Integrity &#8211; The Diff Discipline&#8221;</span></a><span>, </span><em><span>Neural Horizons</span></em><span>, 2026.</span></p></li><li><p><span>Benson, Peter. </span><em><a href="https://www.amazon.com/Neural-Horizons-Psyche-Co-evolution-Machine-Saturated-ebook/dp/B0GS3P39KB/ref=cm_cr_arp_d_product_top?ie=UTF8"><span>Neural Horizons: Psyche, Soul and Co-evolution in a Machine-Saturated Mind</span></a></em><span>, 2026.</span></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_ab2f311cfd6a4201844c6f9b1c530ed7.pdf?index=true"><span>Cognitive Susceptibility Taxonomy Manual v0.8.2 Draft</span></a></em><span>, 2026.</span></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5bf168d1742b48cdb9c81f1881fd7a5b.pdf?index=true"><span>Robo-Psychology Taxonomy v2.0.2 Draft</span></a></em><span>, 2026.</span></p></li><li><p><span>Neural Horizons Ltd. </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay v0.4.3 Draft</span></a></em><span>, 2026.</span></p></li><li><p><span>Lamarra, Tommaso, Caterina Villani, and Marianna M. Bolognesi. </span><a href="https://doi.org/10.1007/s10339-025-01286-5"><span>&#8220;Specificity effect in concrete/abstract semantic categorization task&#8221;</span></a><span>, </span><em><span>Cognitive Processing</span></em><span> 27, 2026.</span></p></li><li><p><span>Berezin, Sergey, and Tatiana Batura. </span><a href="https://aclanthology.org/2022.sdp-1.17/"><span>&#8220;Named Entity Inclusion in Abstractive Text Summarization&#8221;</span></a><span>, Association for Computational Linguistics, 2022.</span></p></li><li><p><span>Goyal, Navita, Ani Nenkova, and Hal Daum&#233; III. </span><a href="https://aclanthology.org/2023.acl-long.463/"><span>&#8220;Factual or Contextual? Disentangling Error Types in Entity Description Generation&#8221;</span></a><span>, Association for Computational Linguistics, 2023.</span></p></li><li><p><span>Hopewell, Sally, et al. </span><a href="https://doi.org/10.1136/bmj-2024-081123"><span>&#8220;CONSORT 2025 statement: updated guideline for reporting randomised trials&#8221;</span></a><span>, </span><em><span>BMJ</span></em><span> 388, 2025.</span></p></li><li><p><span>National Institute of Standards and Technology. </span><em><a href="https://airc.nist.gov/airmf-resources/playbook/"><span>AI Risk Management Framework Playbook: Map and Measure</span></a></em><span>, accessed 25 August 2026.</span></p></li><li><p><span>Paas, Fred, and Jeroen J. G. van Merri&#235;nboer. </span><a href="https://doi.org/10.1177/0963721420922183"><span>&#8220;Cognitive-Load Theory: Methods to Manage Working Memory Load in the Learning of Complex Tasks&#8221;</span></a><span>, </span><em><span>Current Directions in Psychological Science</span></em><span> 29(4), 2020.</span></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Institutional Blindfold - Change-Control for Model Updates]]></title><description><![CDATA[Change control matters: when the model, prompts, retrieval, memory or policy change while the institution is still calling it the same tool. It&#8217;s not.]]></description><link>https://neuralhorizons.substack.com/p/institutional-blindfold-change-control-cc9</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/institutional-blindfold-change-control-cc9</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Mon, 24 Aug 2026 20:47:10 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/212467391/46451cdca786d30f97a2188ec8c5e902.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>Artificial intelligence governance must broaden its focus from the underlying model to the entire decision environment, which includes prompts, memory, and retrieval systems.</p><p>Because these components can change silently while the product name remains the same, institutions face a risk called post-modification safety drift where a tool&#8217;s behaviour becomes unpredictable or biased.</p><p>To manage this, organizations should maintain an approved-state record that serves as a technical and behavioural baseline for evaluating updates.</p><p>Effective change control requires more than technical monitoring; it necessitates a formal process of re-governance to ensure that human oversight and institutional safety claims remain valid.</p><p>Material changes should be judged by their impact on human rights and safety rather than the technical size of the update.</p><p>This approach ensures that AI systems evolve under institutional authority rather than through unmanaged, invisible transitions.</p><p>Full article available here </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;4b07e573-1e00-415f-af9c-cf815c8c9659&quot;,&quot;caption&quot;:&quot;When the name stays still but the system moves&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Institutional Blindfold - Change-Control for Model Updates&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-24T20:46:28.822Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!ollS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/institutional-blindfold-change-control&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:212467384,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Institutional Blindfold - Change-Control for Model Updates]]></title><description><![CDATA[Change control matters: when the model, prompts, retrieval, memory or policy change while the institution is still calling it the same tool. It&#8217;s not.]]></description><link>https://neuralhorizons.substack.com/p/institutional-blindfold-change-control</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/institutional-blindfold-change-control</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Mon, 24 Aug 2026 20:46:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ollS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2><span>When the name stays still but the system moves</span></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ollS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ollS!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!ollS!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!ollS!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ollS!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ollS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1989195,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/212467384?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ollS!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!ollS!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!ollS!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ollS!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76f158c1-a5ed-47f7-9701-89e3fefd23da_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>On 25 April 2025, OpenAI completed an update to GPT-4o intended to make the model more useful and responsive. Its offline evaluations looked positive, and sample users in A/B tests appeared to like the change. Within days, the company concluded that the model had become noticeably more sycophantic: too ready to validate doubts, anger and impulsive ideas. OpenAI later said several changes that seemed beneficial individually - including signals from user feedback, memory and fresher data - may have combined to produce the shift. A system-prompt intervention followed, then a rollback. </span><a href="https://openai.com/index/expanding-on-sycophancy/"><span>[1]</span></a></p><p><span>The episode matters beyond the one chatbot update process. It exposes a governance problem that is easy to miss when software keeps the same icon, contract and product name. The thing an institution approved on Monday may not be behaviourally identical on Friday.</span></p><p><span>The previous article in this series, </span><em><a href="/__u/neuralhorizons.substack.com/p/the-institutional-blindfold-the-exit?r=2tdtxm"><span>The Exit Strategy Test</span></a></em><a href="/__u/neuralhorizons.substack.com/p/the-institutional-blindfold-the-exit?r=2tdtxm"><span>,</span></a><span> asked whether an organisation can leave an artificial intelligence supplier without losing evidence, continuity or competence. It argued that exit is an operational capability rather than a termination clause, and that the test should be repeated after a major model, data, integration or supplier change. </span><a href="/__u/neuralhorizons.substack.com/p/the-institutional-blindfold-the-exit"><span>[2]</span></a></p><p><span>That leads to the next question. What counts as a change?</span></p><p><span>For conventional software, change control often focuses on a new version, a configuration adjustment or a security patch. Artificial intelligence complicates the picture because behaviour can move when the underlying model stays fixed. The retrieval source can change. Memory can be switched on. A system instruction can be rewritten. A routing rule can send some cases to another model. A safety policy can become stricter or looser. A tool can gain access to another database. The interface may look untouched while the decision environment underneath it has changed.</span></p><p><span>The thesis we&#8217;re arguing in this article is therefore simple:</span></p><p style="text-align: center;"><em><span>a tool that changes under the institution must be re-governed, not merely updated.</span></em></p><h2><span>The approved system is larger than the model</span></h2><p><span>It helps to separate five things that ordinary product language often compresses into &#8220;the AI&#8221;.</span></p><ul><li><p><span>The </span><em><span>model</span></em><span> is the underlying engine that predicts, classifies or generates.</span></p></li><li><p><span>The </span><em><span>wrapper</span></em><span> is the surrounding set of prompts, routing rules, tools, thresholds and interface choices that shape how that engine is used.</span></p></li><li><p><em><span>Retrieval</span></em><span> decides what external material is fetched and placed in front of the model for a particular task.</span></p></li><li><p><em><span>Memory</span></em><span> determines what prior information is retained or reintroduced.</span></p></li><li><p><em><span>Policy</span></em><span> governs boundaries: what the system should refuse, escalate, prioritise or permit.</span></p></li></ul><p><span>A material change in any one of these can alter the answer a person receives.</span></p><p><span>Change the model and reasoning patterns, refusal behaviour or error distributions may move. Change retrieval and the model may see different evidence while remaining technically identical. Change memory and the system may respond differently because it carries more context from earlier interactions. Change a system instruction and its tone, priorities or deference can shift. Change routing and two apparently identical users may be served by different models.</span></p><p><span>The National Institute of Standards and Technology&#8217;s Generative Artificial Intelligence Profile reflects this broader view. It calls for organisations to monitor third parties for changes, maintain records of those changes with provenance and metadata, reassess models when they are fine-tuned or adapted, and include change management in post-deployment monitoring. Its guidance also treats retrieval-augmented generation, fine-tuning and other adaptations as risk-relevant parts of the deployed system rather than details beneath governance. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[3]</span></a></p><p><span>Our Neural Horizons project document </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a><span> gives one part of this problem a useful name: </span><em><span>post-modification safety drift</span></em><span>. Our formal </span><em><span>Post-Modification Safety Drift Overlay</span></em><span> is not a claim that every modification creates harm. It is a release-gating trigger for cases where safety behaviour appears, worsens, reverses or otherwise shifts after modification. Crucially, its scope includes not only fine-tuning but retrieval, wrappers, guardrails, memory, tooling and stacked changes. Our companion </span><em><span>Modification Provenance / Drift Report</span></em><span> requires recording the base and modified system identities, the method of change, surrounding system changes, pre- and post-change tests, known regressions and the release decision. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>[4]</span></a></p><p><span>That distinction is important because governance teams can otherwise make a category error. They ask the supplier, &#8220;</span><em><span>Did the model change?</span></em><span>&#8221; The supplier truthfully answers no. What changed was the search index, the memory policy or the system prompt. The institution records &#8220;</span><em><span>no material model update</span></em><span>&#8221; and carries on.</span></p><p><span>But the person affected does not interact with a model in isolation. They interact with the resulting system.</span></p><p style="text-align: center;"><em><span>The unit of change is not the model. It is the decision environment around the person.</span></em></p><h2><span>Useful updates can still move the ground</span></h2><p><span>None of this means organisations should freeze their artificial intelligence systems in stone.</span></p><p><span>Software needs patches. Models become stale. Data distributions change. Vulnerabilities are discovered. Poor behaviour should be corrected. The United Kingdom&#8217;s National Cyber Security Centre explicitly advises software providers to test updates before deployment, use version and configuration control, consider progressive deployment and retain the ability to roll back a specific version. Its premise is practical: software will need to change, so change itself must be managed. </span><a href="https://www.ncsc.gov.uk/collection/software-security-code-of-practice-implementation-guidance/secure-deployment-maintenance"><span>[5]</span></a></p><p><span>Modern machine-learning operations already assume continuous change. Google Cloud&#8217;s guidance on Machine Learning Operations describes source control, automated testing, model registries, metadata stores, deployment pipelines and production monitoring as parts of mature delivery. A model can require retraining when the data it encounters no longer resembles the data on which it was trained. </span><a href="https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning"><span>[6]</span></a></p><p><em><span>The governance problem begins when technical release management is mistaken for institutional authorisation.</span></em></p><p><span>A pipeline can prove that new code compiled, that a model passed a benchmark and that deployment succeeded. It cannot decide whether a council is still comfortable using that system to prioritise housing cases, whether a hospital accepts a new false-negative pattern, or whether a university is willing to let a changed retrieval process shape student advice. Those are decisions about purpose, consequences and acceptable uncertainty.</span></p><p><span>The 2025 United Kingdom government AI Playbook makes the bridge unusually explicit. It says updates to artificial intelligence systems should undergo quantitative testing and validation as part of change control; changes should be documented; releases should be managed so they can be withdrawn and systems reverted where necessary; and performance and model drift should be monitored over time. </span><a href="https://www.gov.uk/government/publications/ai-playbook-for-the-uk-government/artificial-intelligence-playbook-for-the-uk-government-html"><span>[7]</span></a></p><p><span>The European Union&#8217;s Artificial Intelligence Act draws a related legal distinction for high-risk systems. A &#8220;substantial modification&#8221; can require a new conformity assessment, while changes that were predetermined and documented as part of the original assessment can be treated differently. The precise legal test is narrower than the governance test proposed here, but the principle is useful: some changes are significant enough that yesterday&#8217;s assurance cannot simply be carried forward. </span><a href="https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-43"><span>[8]</span></a></p><p><span>There is an important boundary condition. Requiring a committee meeting for every prompt typo, dependency patch or harmless interface change would produce bureaucratic overhead without measurably improving safety. Organisations also cannot reasonably demand disclosure of every proprietary parameter inside a supplier&#8217;s service.</span></p><p><span>Change control should therefore be </span><em><span>proportionate to possible consequence</span></em><span>, not to the technical size of the update. A one-line instruction change that alters which safeguarding cases are escalated may deserve more scrutiny than a large infrastructure migration that leaves behaviour unchanged. What matters is whether the change can alter evidence, action, rights, safety, workload, contestability or the human role.</span></p><p style="text-align: center;"><em><span>That is why &#8220;minor update&#8221; is not a governance category until someone has said minor for whom, and measured against what.</span></em></p><h2><span>Change control requires a baseline the institution owns</span></h2><p><span>An organisation cannot detect drift from a baseline it never recorded.</span></p><p><span>Before a consequential artificial intelligence system enters routine use, the institution needs an approved-state record: enough information to identify the system that was actually tested and accepted. That does not require possession of the supplier&#8217;s trade secrets, but it does require operational facts.</span></p><p><span>For a high-impact workflow, the record should identify the model or supplier release channel; the relevant prompt and policy version; retrieval sources and ranking configuration; memory settings; connected tools and permissions; routing and decision thresholds; the evaluation cases used for approval; known limitations; and the points where a human is expected to review, override or escalate.</span></p><p><span>This record does two jobs. It lets the institution ask whether something has changed, and it preserves the meaning of earlier assurance.</span></p><p><span>Suppose a housing service tested an AI assistant against 250 representative cases, including homelessness risk, domestic abuse, disability adaptations and incomplete records. The supplier later changes retrieval so that more recent notes are weighted more heavily. The underlying model is unchanged. Overall benchmark quality may even improve. But an old safeguarding note buried deep in a case file may now be </span><em><span>less likely to enter the model&#8217;s working context</span></em><span>.</span></p><p><span>That example is illustrative, not a reported incident. Its point is structural: the relevant test after the change is not &#8220;</span><em><span>Does the new version perform well?</span></em><span>&#8221; It is &#8220;</span><em><span>Do the claims on which we authorised this use still hold?</span></em><span>&#8221;</span></p><p><span>A workable change-control process therefore asks three questions.</span></p><p><span>First, </span><em><span>what moved?</span></em><span> The organisation needs a versioned account of changes across the effective system, including model, wrapper, retrieval, memory, tools and policy. Where the supplier controls those elements, contracts and service arrangements should require enough notice and provenance to answer the question. NIST specifically recommends records of third-party changes, including sources, timestamps and metadata, and says contracts should address system changes over time. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[9]</span></a></p><p><span>Second, </span><em><span>what claim could the change invalidate? </span></em><span>A retrieval update may invalidate an evidence-completeness claim. A memory update may alter privacy or consistency assumptions. A new model may affect accuracy or refusal behaviour. A policy change may alter escalation. This is where the institution should rerun the smallest set of tests capable of challenging the affected assurance, rather than reflexively repeating every evaluation ever performed.</span></p><p><span>Third, </span><em><span>can we compare and reverse?</span></em><span> There should be a defined observation period, preserved old and new evidence, and a rollback path where the consequences justify one. Some vendors already expose technical mechanisms for version stability; for example, OpenAI&#8217;s application programming interface documentation distinguishes dated model snapshots that lock behaviour to a specific version from moving aliases. That does not solve governance by itself, but it shows that version pinning is a practical design choice in at least some services. </span><a href="https://developers.openai.com/api/docs/models/gpt-4.1"><span>[10]</span></a></p><p><span>Where a supplier does not offer pinning, advance notice or rollback, the institution has learned something important. The risk has not disappeared, and the control has shifted outward. </span></p><p><span>The correct response may be stronger monitoring, narrower permitted use, additional human review, a tested fallback, or a decision that the service is unsuitable for that purpose.</span></p><p><span>This is also where the previous article&#8217;s </span><a href="/__u/neuralhorizons.substack.com/p/the-institutional-blindfold-the-exit"><span>Exit Strategy Test</span></a><span> returns. Change control without exit can become a ritual: the organisation identifies an unacceptable update but has nowhere to go. Exit without change control is equally weak: the organisation retains the theoretical power to leave but may not notice that the conditions justifying departure have arrived. </span><a href="/__u/neuralhorizons.substack.com/p/the-institutional-blindfold-the-exit"><span>[2]</span></a></p><p><span>The two controls are complements. One preserves choice. The other preserves awareness.</span></p><h2><span>Silent updates change the human job as well</span></h2><p><span>There is another reason to treat modifications as governance events. People learn systems.</span></p><p><span>A housing officer who has used an artificial intelligence assistant for a year will develop expectations about it. She may know that it tends to overlook handwritten attachments, that its summaries are strongest on recent correspondence, or that a particular priority recommendation needs extra scrutiny. These are not necessarily signs of blind trust. They can be the ordinary practical knowledge through which a professional supervises an imperfect tool.</span></p><p><span>Now change the tool silently.</span></p><p><span>Perhaps the new version fixes the handwritten-document problem but becomes less cautious with incomplete evidence. Perhaps memory is added, making some answers more context-sensitive. Perhaps a policy update makes the assistant less likely to flag uncertainty. The interface is familiar, so the officer&#8217;s learned vigilance points at yesterday&#8217;s weaknesses.</span></p><p><span>Her mental model of the machine has become stale.</span></p><p><span>That is a human-factors problem created by change management, not a character flaw in the user. Re-training cannot consist of an email saying &#8220;AI improvements have been deployed&#8221;. Staff need to know what changed in terms that alter practice: which failure modes became less likely, which became more likely or remain uncertain, what should now be checked, and when the system should be challenged.</span></p><p><span>The OpenAI sycophancy episode is instructive here for another reason. Positive short-term user signals did not establish that the behavioural change was beneficial. OpenAI reported that A/B results looked favourable even though some expert testers had flagged concerns and the released behaviour later proved unacceptable. </span><a href="https://openai.com/index/expanding-on-sycophancy/"><span>[11]</span></a><span> Satisfaction and uptake can tell an institution whether people prefer a changed system. They cannot, by themselves, tell it whether professional judgement, fairness, safety or human agency has improved.</span></p><p><span>This is where the human line matters. When a system participates in consequential judgement, the organisation is not only maintaining software. It is maintaining the conditions under which a person can understand what the tool is doing, contest it and remain responsible for the decision.</span></p><p><span>A silent behavioural update asks the human to carry accountability for a machine they have not yet had the chance to relearn.</span></p><h2><span>Put updates back under institutional authority</span></h2><p><span>The practical aim is not to slow every improvement. </span><em><span>It is to prevent consequential change from arriving as a fait accompli</span></em><span>.</span></p><p><strong><span>Leaders and risk owners</span></strong><span> should define what counts as a material artificial intelligence change for each high-impact use. The trigger should be based on possible effects on decisions, rights, safety, evidence and human oversight, not merely on whether the supplier calls it a new model version. Every material change needs a named owner, a defined re-evaluation and an explicit release, restriction or rollback decision.</span></p><p><strong><span>Technology and operational teams</span></strong><span> should maintain the approved-state record and a modification history across the whole effective system. Before release, they should run targeted old-versus-new evaluations on realistic cases and known edge conditions; after release, they should watch for changes in error patterns, escalations, overrides and complaints. Where feasible, staged rollout or shadow testing should expose change before the whole organisation inherits it. These practices align with government artificial-intelligence guidance and established software and machine-learning release controls. </span><a href="https://www.gov.uk/government/publications/ai-playbook-for-the-uk-government/artificial-intelligence-playbook-for-the-uk-government-html"><span>[12]</span></a></p><p><strong><span>Procurement and legal teams</span></strong><span> should treat supplier-controlled change as an allocation of governance power. For consequential uses, contracts should seek advance notification of material changes, intelligible change logs, stable or deferrable versions where feasible, access to evidence needed for re-testing, incident notification, and workable rollback or exit rights. Where a vendor will not provide those controls, the limitation should appear in the risk decision rather than vanish inside standard terms. NIST&#8217;s guidance specifically connects supplier agreements with provenance, ongoing monitoring, system changes and incident responsibilities. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[9]</span></a></p><p><strong><span>Service owners and educators</span></strong><span> should update the human operating model as well as the technical one. Tell staff what changed, which checks still matter and which old assumptions no longer hold. Preserve channels for challenge and record whether human overrides or complaints shift after an update. A change that leaves benchmark performance intact but weakens the institution&#8217;s ability to question the system is still a governance change.</span></p><p><span>Artificial intelligence will change. Often it should. The institutional choice is whether that change happens </span><em><span>to</span></em><span> the organisation or </span><em><span>under its authority</span></em><span>.</span></p><p><span>A responsible institution does not need to hold an artificial intelligence system still. It needs to know when the thing carrying part of its judgement has changed enough that yesterday&#8217;s approval no longer answers today&#8217;s question.</span></p><p><span>The next article turns to the harder fallback: </span><em><span>The No-AI Reversibility Clause.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Benson, Peter, &#8220;The Institutional Blindfold - The Exit Strategy Test&#8221;, </span><em><span>Neural Horizons</span></em><span>, 4 August 2026. </span><a href="/__u/neuralhorizons.substack.com/p/the-institutional-blindfold-the-exit"><span>[2]</span></a></p></li><li><p><span>OpenAI, &#8220;Expanding on what we missed with sycophancy&#8221;, 2 May 2025. </span><a href="https://openai.com/index/expanding-on-sycophancy/"><span>[11]</span></a></p></li><li><p><span>National Institute of Standards and Technology, </span><em><span>Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1</span></em><span>, July 2024. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[13]</span></a></p></li><li><p><span>UK Government, </span><em><span>Artificial Intelligence Playbook for the UK Government</span></em><span>, 10 February 2025. </span><a href="https://www.gov.uk/government/publications/ai-playbook-for-the-uk-government/artificial-intelligence-playbook-for-the-uk-government-html"><span>[14]</span></a></p></li><li><p><span>UK National Cyber Security Centre, </span><em><span>Software Security Code of Practice: Implementation Guidance, Theme 3 &#8212; Deploy Software Securely</span></em><span>, guidance inspected August 2026. </span><a href="https://www.ncsc.gov.uk/collection/software-security-code-of-practice-implementation-guidance/secure-deployment-maintenance"><span>[5]</span></a></p></li><li><p><span>Google Cloud, </span><em><span>MLOps: Continuous Delivery and Automation Pipelines in Machine Learning</span></em><span>, Cloud Architecture Center, version inspected August 2026. </span><a href="https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning"><span>[6]</span></a></p></li><li><p><span>European Commission AI Act Service Desk, &#8220;Article 43: Conformity Assessment&#8221;, Regulation (EU) 2024/1689, version inspected August 2026. </span><a href="https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-43"><span>[8]</span></a></p></li><li><p><span>Neural Horizons Ltd, </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>Robo-Psychology Taxonomy</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>, Aug 2026</span></a><span>, including the Post-Modification Safety Drift Overlay and Modification Provenance / Drift Report. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>[4]</span></a></p></li><li><p><span>OpenAI, </span><em><span>GPT-4.1 Model Documentation &#8212; Snapshots</span></em><span>, version inspected August 2026. </span><a href="https://developers.openai.com/api/docs/models/gpt-4.1"><span>[10]</span></a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Evidence Frame Integrity – The Evidence Contact Test]]></title><description><![CDATA[We explore the Evidence Contact Test, a governance framework designed to ensure that AI-driven decisions are based on the actual inspection of underlying data rather than just persuasive summaries.]]></description><link>https://neuralhorizons.substack.com/p/evidence-frame-integrity-the-evidence-589</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/evidence-frame-integrity-the-evidence-589</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Sun, 23 Aug 2026 21:09:30 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/212461864/2ed2c2d8fb31d00dcdc50007ad859128.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>We explore the Evidence Contact Test, a governance framework designed to ensure that AI-driven decisions are based on the actual inspection of underlying data rather than just persuasive summaries.</p><p>It highlights the risk of false evidentiary posture, where a system or human acts as if they have verified information that is actually missing, outdated, or unexamined.</p><p>To combat this, we propose a six-question audit and an Edge-Case Ledger to preserve visibility for outliers and contradictory facts that compression often erases. The methodology emphasizes retraceable compression, ensuring that every automated recommendation maintains a direct, verifiable path back to its original source material.</p><p>Furthermore, we warn against criteria collapse, where users mistake the professional formatting of AI output for factual accuracy.</p><p>The framework advocates for a risk-tiered approach to governance that prioritizes substantive evidence contact over mere throughput or formal approval clicks.</p><p>Full article available here </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;952acf57-b435-4571-902f-e001d9ec6d7a&quot;,&quot;caption&quot;:&quot;Our previous article in the &#8216;Evidence Frame Integrity&#8217; series left a useful object on the table: the Edge-Case Ledger. Its purpose was to stop consequential exceptions from vanishing when evidence is compressed into a summary, score, shortlist or dashboard. The ledger asks what the smooth centre of the story left behind: outliers, minority cohorts, cont&#8230;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Evidence Frame Integrity &#8211; The Evidence Contact Test&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:170286538,&quot;name&quot;:&quot;Peter Benson&quot;,&quot;bio&quot;:&quot;CEO Neural Horizons Ltd, focusing on the intersections of Cyber, Security, Privacy, Ethics, and the impacts on society and individuals through cyber-psychology, cyber-sociology. &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8986e8b-6f76-4e41-9ebe-55c0dc6648d8_200x200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-08-23T21:06:26.969Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!o_zw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://neuralhorizons.substack.com/p/evidence-frame-integrity-the-evidence&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:212461861,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:4230540,&quot;publication_name&quot;:&quot;Neural Horizons Substack&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!zS6f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26475efd-0796-474a-ab4c-32c166759796_1024x1024.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Evidence Frame Integrity – The Evidence Contact Test]]></title><description><![CDATA[Our previous article in the &#8216;Evidence Frame Integrity&#8217; series left a useful object on the table: the Edge-Case Ledger. Its purpose was to stop consequential exceptions from vanishing when evidence is compressed into a summary, score, shortlist or dashboard. The ledger asks what the smooth centre of the story left behind: outliers, minority cohorts, contradictory observations, low-confidence cases and alternatives that could change a responsible decision. It also carried a warning. A ledger can exist without anyone opening it. A source link can be genuine without anyone following it. Formal review can occur while the people approving the result never reach the record underneath.]]></description><link>https://neuralhorizons.substack.com/p/evidence-frame-integrity-the-evidence</link><guid isPermaLink="false">https://neuralhorizons.substack.com/p/evidence-frame-integrity-the-evidence</guid><dc:creator><![CDATA[Peter Benson]]></dc:creator><pubDate>Sun, 23 Aug 2026 21:06:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!o_zw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!o_zw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!o_zw!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!o_zw!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!o_zw!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!o_zw!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_webp, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!o_zw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4844985,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neuralhorizons.substack.com/i/212461861?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!o_zw!, /__u/neuralhorizons.substack.com/w_424, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!o_zw!, /__u/neuralhorizons.substack.com/w_848, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!o_zw!, /__u/neuralhorizons.substack.com/w_1272, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!o_zw!, /__u/neuralhorizons.substack.com/w_1456, /__u/neuralhorizons.substack.com/c_limit, /__u/neuralhorizons.substack.com/f_auto, /__u/neuralhorizons.substack.com/q_auto:good, /__u/neuralhorizons.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c3f5f8-0508-48d5-80bf-afaf11ca12c0_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Our previous article in the &#8216;Evidence Frame Integrity&#8217; series left a useful object on the table: the </span><em><a href="/__u/neuralhorizons.substack.com/p/evidence-frame-integrity-the-edge?r=2tdtxm"><span>Edge-Case Ledger</span></a></em><span>. Its purpose was to stop consequential exceptions from vanishing when evidence is compressed into a summary, score, shortlist or dashboard. The ledger asks what the smooth centre of the story left behind: outliers, minority cohorts, contradictory observations, low-confidence cases and alternatives that could change a responsible decision. It also carried a warning. A ledger can exist without anyone opening it. A source link can be genuine without anyone following it. Formal review can occur while the people approving the result never reach the record underneath. </span><a href="/__u/neuralhorizons.substack.com/p/evidence-frame-integrity-the-edge"><span>[1]</span></a></p><p><span>That is the next problem in source-contact collapse. A hallucination can invent a fact or source. False evidentiary posture is subtler: an answer, recommendation or workflow behaves as though the relevant evidence was accessed and verified when that evidence was absent, stale, mismatched, uninspected or replaced by a proxy. The output may even be correct. What is false is the implied relationship between the answer and its evidence.</span></p><p><span>We separate this failure from ordinary factual error precisely because accuracy, fluency and a convincing explanation do not establish that the required evidentiary surface was actually used. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>[2]</span></a></p><p><span>For evidence-grounded work, that suggests a strong first governance gate: </span><em><span>what evidence did the system or the human actually contact, and how do we know?</span></em></p><p><span>The </span><em><span>Evidence Contact Test</span></em><span> is a practical way to answer it.</span></p><h2><span>A ledger is not contact</span></h2><p><span>The Edge-Case Ledger solved a visibility problem. The </span><em><span>Evidence Contact Test</span></em><span> adds a behaviour problem. </span><a href="/__u/neuralhorizons.substack.com/p/evidence-frame-integrity-the-edge"><span>[3]</span></a></p><p><span>Consider a hiring shortlist produced from applications, interview notes and assessment results. The system may show links to the candidate files. It may preserve a panel labelled &#8220;exceptions&#8221;. It may even provide a neat audit trail. None of those features proves that the model retrieved the correct files, that the links correspond to the claims being made, or that the hiring panel inspected the material before approving the ranking. A control can be present in the interface and absent in the decision. Our machine-side behavioural framework treats missing, wrong, stale, degraded or unverified evidence as distinct from the later problem of how genuine evidence is compressed into recommendations. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>[4]</span></a></p><p><span>That distinction matters. On the machine side, a recommendation-frame problem arises when artificial intelligence compresses evidence into a shortlist or executive view before people deliberate, while uncertainty, alternatives or source material are hidden or ignored. A separate evidence-frame integrity problem arises when the system acts as if it has contacted the required evidence although that evidence is missing, wrong, stale, degraded, unverified or inferred from cues such as metadata. The first can narrow what humans see. The second can counterfeit the very premise that there was something sound to see. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_792f298c87c04a0b8fed85672e633153.pdf?index=true"><span>[4]</span></a></p><p><span>On the human side, our (human factors) </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_cf90bb86114f45788381bb0894b4b1cc.pdf?index=true"><span>Cognitive Susceptibility Taxonomy</span></a><span> describes </span><em><span>Recommendation Frame Capture / Evidence Contact Loss</span></em><span>: the recommendation becomes the first point of our meaningful contact with the case, rather than a tool used after contact with the underlying evidence. Its warning signs include low source-opening behaviour, absent edge-case review, weak uncertainty display and approval records that show a human choice without showing evidence inspection. This is a susceptibility pattern and governance lens, not a diagnosis of a person. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_cf90bb86114f45788381bb0894b4b1cc.pdf?index=true"><span>[5]</span></a></p><p><em><a href="/__u/neuralhorizons.substack.com/p/evidence-frame-integrity-the-edge"><span>The Edge Case Ledger</span></a><span> </span></em><span>also named the institutional version of the problem: the Institutional Blindfold, where an organisation can slide from examining the record to approving a compressed representation of it. In practical terms, the &#8220;human line&#8221; here is not a demand that people perform every calculation themselves. It is the boundary at which a reviewer can still inspect, contest and change the machine-shaped frame. </span><a href="/__u/neuralhorizons.substack.com/p/evidence-frame-integrity-the-edge"><span>[6]</span></a></p><p><span>There are practical objections here of course; a chief executive cannot read every transaction behind a quarterly dashboard; a clinician cannot reopen every historical note; a teacher cannot audit every token used to produce a student-risk summary. &#8216;</span><em><span>The Edge Case Ledger</span></em><span>&#8217; already gave the right answer: the alternative to compressed evidence is not an archive dumped on the desk. It is </span><em><span>retraceable compression</span></em><span> &#8211; a concise view with a direct route to the consequential evidence and exceptions. </span><a href="/__u/neuralhorizons.substack.com/p/evidence-frame-integrity-the-edge"><span>[6]</span></a></p><p><span>We are aware that there is no defensible universal rule such as &#8220;open 20 per cent of sources&#8221; that turns contact into substance. The required depth depends on consequence, reversibility, evidence quality and the decision-maker&#8217;s role. The test should therefore be risk-tiered, not reduced to a single compliance percentage. NIST&#8217;s Generative AI Profile likewise treats risk-management effort as something to tailor to context, likelihood and severity. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[7]</span></a></p><h2><span>Six questions before an AI-assisted decision</span></h2><p><span>The Evidence Contact Test can be run as six questions. A &#8220;yes&#8221; requires observable evidence. &#8220;Unknown&#8221; is not a pass.</span></p><h4><strong><span>First: what evidence was required?</span></strong></h4><p><span>Before reviewing the answer, name the evidence surface the task depends on: the policy document, patient record, contract clause, dataset, interview notes, sensor feed, research paper or current regulation. This prevents a common substitution: judging the quality of the prose before establishing whether the task required access to material outside the model&#8217;s visible context. Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>Robo-Psychology </span></a><span>evidence-frame guidance treats absence, failed upload, stale retrieval, wrong attachment and proxy substitution as distinct reasons to detect a failure of evidence contact rather than behave as though the evidence had been seen. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>[8]</span></a></p><h4><strong><span>Second: did the system reach the right evidence, in the right version, at the right time?</span></strong></h4><p><span>A citation or file name is not proof of retrieval. For consequential claims, the record should show what source was fetched, which version or date was used, whether retrieval succeeded, and whether any source was unavailable or degraded. NIST&#8217;s Generative AI Profile recommends establishing practices for data origin and content lineage and testing flows through original sources, transformations and decision criteria; its AI Risk Management Framework Playbook calls for provenance documentation covering sources, origins, transformations, dependencies, constraints and metadata. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[9]</span></a></p><p><span>This is where stale evidence matters. A perfectly quoted policy superseded six months ago can produce a well-supported wrong decision. Contact is temporal as well as semantic; our </span><em><span>Evidence-Frame Integrity Overlay</span></em><span> explicitly treats stale or wrong evidence as a failure condition. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>[8]</span></a></p><h4><strong><span>Third: does the contacted evidence support the exact claim?</span></strong></h4><p><span>Retrieval is only the middle of the chain. The retrieved passage might be relevant to the topic while failing to justify the sentence attached to it. Research on retrieval-augmented generation has made this separation explicit. ARES evaluates context relevance, answer faithfulness and answer relevance as different dimensions; RAGAs likewise separates retrieval quality from faithful use of the retrieved material and from the quality of the generated answer. </span><a href="https://aclanthology.org/2024.naacl-long.20/"><span>[10]</span></a></p><p><span>Citation research reached the same conclusion earlier. The 2023 ALCE benchmark scored fluency, correctness and citation quality separately; in its ELI5 experiments, even the strongest systems in that study lacked complete citation support half the time. That figure should </span><em><span>not</span></em><span> be treated as a current failure rate for today&#8217;s systems. Its enduring lesson is structural: the presence of citations and the support they provide are different measurements. </span><a href="https://aclanthology.org/2023.emnlp-main.398/"><span>[11]</span></a></p><h4><strong><span>Fourth: what did compression remove?</span></strong></h4><p><span>This carries forward the </span><em><span>Edge-Case Ledger</span></em><span>. Ask what outliers, contradictory cases, missing values, subgroup effects, uncertainty and plausible alternatives were suppressed by the summary or ranking. The important test is not whether anything was omitted &#8211; every useful summary omits &#8211; but whether the omitted material could change the decision, the distribution of harm, or confidence in the recommendation. </span><a href="/__u/neuralhorizons.substack.com/p/evidence-frame-integrity-the-edge"><span>[12]</span></a></p><h4><strong><span>Fifth: did a human make substantive contact before approving the result?</span></strong></h4><p><span>A signature is an event. Evidence contact is an activity.</span></p><p><span>Our Robo-Psychology recommendation-frame guidance proposes looking for source opening, edge-case inspection, alternative review, uncertainty review, a real opportunity to challenge, and a recorded rationale. Add to this, our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay</span></a><span> &#8211; the project framework concerned with whether the human&#8211;AI relationship preserves useful human capability &#8211; which adds direct evidence access and &#8220;time in evidence&#8221; as relevant indicators. None of these requires an executive to become the analyst. They require the organisation to distinguish an approval click from judgement. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>[13]</span></a></p><p><span>For high-consequence decisions, record which decision-driving sources or samples the reviewer actually inspected, which exception was checked, and whether the reviewer had authority to recover an excluded option or reject the AI frame. A person who is technically allowed to inspect evidence but is penalised for slowing the queue has a very different form of oversight from someone given time, authority and a workable route to challenge.</span></p><p><span>Our project frameworks explicitly treat throughput pressure, challenge opportunity and alternative recovery as relevant to whether review is substantive. </span><a href="https://www.neural-horizons.ai/resources"><span>[14]</span></a></p><h4><strong><span>Sixth: can another person reconstruct the path later?</span></strong></h4><p><span>Starting from the final claim or recommendation, an auditor should be able to move backwards through the AI output, the retrieved passages or data, relevant transformations, source version and responsible actors. NIST explicitly calls for data and content lineage; the World Wide Web Consortium&#8217;s PROV data model supplies a general vocabulary built around entities, activities and agents, including relationships such as use, derivation and attribution. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[15]</span></a></p><p><span>The counter-case is that automated checks can perform much of this work. They can. Retrieval tests, faithfulness scoring, provenance logging and automated evaluators can reduce the human burden, especially at scale. ARES and RAGAs are examples of attempts to automate parts of that evaluation. The mistake would be to let an automated score certify its own evidentiary premise. If the evaluator is testing the wrong source, stale corpus or incomplete record, efficiency simply accelerates the wrong assurance. </span><a href="https://aclanthology.org/2024.naacl-long.20/"><span>[10]</span></a></p><p><span>We recognise that retrieval-augmented generation evaluation is fast-moving, benchmark-dependent and sensitive to domain. The studies above support decomposing the problem into distinct checks, but and not intended to establish a universal threshold for safe evidence contact in medicine, education, hiring, law or public administration. NIST itself notes continuing measurement uncertainty in generative-AI risk assessment. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[7]</span></a></p><h2><span>Provenance is a trail, not a verdict</span></h2><p><span>&#8220;Provenance&#8221; is often offered as the cure for this problem, and it is indispensable. It is also easy to ask it to do too much.</span></p><p><span>A useful provenance record can tell us where an artefact came from, what transformations occurred, which system or person acted on it, and which version entered the workflow. The W3C PROV model formalises those relationships through entities, activities and agents. NIST&#8217;s Generative AI Profile asks organisations to establish practices for data origin and content lineage and to test flows through original sources, transformations and decision-making criteria. </span><a href="https://www.w3.org/TR/prov-dm/"><span>[16]</span></a></p><p><span>Think of provenance as parcel tracking. It can show that a package moved from a named sender through a known depot to your door. That is valuable. It does not tell you that the sender put the correct medicine in the box.</span></p><p><span>The Coalition for Content Provenance and Authenticity makes this boundary unusually clear. Its C2PA specification is designed to make provenance claims verifiable and resistant to undetected tampering, but its own guiding principle says the specification does not judge whether provenance data are &#8220;good&#8221; or &#8220;bad&#8221;; it validates their association with an asset, their form and their integrity. A trusted trail therefore cannot, by itself, prove that a source is accurate, current, representative or relevant to the claim. </span><a href="https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.html"><span>[17]</span></a></p><p><span>This matters because false evidence posture can survive excellent logging. Imagine a system that faithfully records that it retrieved Policy_v7.pdf, then accurately shows the paragraphs it used, while the operative policy is version 9. The lineage is clean. The decision is still grounded in the wrong thing.</span></p><p><span>Evidence integrity requires two checks that should never be merged: </span><em><span>Can we trace the path?</span></em><span> and </span><em><span>Was the path evidentially valid?</span></em></p><p><span>NIST&#8217;s lineage guidance and the Robo-Psychology evidence-frame control address different parts of precisely this distinction. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[18]</span></a></p><p><span>We do need to note that this can add cost to the operating model. Fine-grained provenance can become another compliance machine: expensive to retain, difficult to interpret and easy to produce at a granularity no decision-maker will use. That concern argues for consequence-based retention rather than maximal logging. Keep enough information to reconstruct consequential claims and decisions; do not turn every low-risk drafting interaction into a forensic archive. NIST&#8217;s framework is explicitly risk-management oriented and voluntary, and its website states that AI RMF 1.0 is being revised in 2026, so it should not be treated as a checklist frozen in time. </span><a href="https://www.nist.gov/itl/ai-risk-management-framework"><span>[19]</span></a></p><p><span>The Evidence Contact Test covered here is a synthesis, not an accredited standard: provenance methods help establish lineage; the project frameworks add separate tests for source validity, compression, human review and preserved decision authority. That synthesis still needs a level of domain-specific validation. </span><a href="https://www.w3.org/TR/prov-dm/"><span>[20]</span></a></p><h2><span>Why smart people accept the evidence posture</span></h2><p><span>False evidence posture works because it often arrives wrapped in genuine usefulness.</span></p><p><span>Summaries save time. Retrieval systems can bring a large document collection into reach. Ranked options can help an overloaded manager navigate plausible choices. RAGAs describes retrieval-augmented generation as a way to connect language models to reference databases and reduce hallucination risk, while emphasising that retrieval quality and faithful use of retrieved passages remain separate evaluation problems. The benefit is real; so is the need to measure what happened between source and answer. </span><a href="https://aclanthology.org/2024.eacl-demo.16/"><span>[21]</span></a></p><p><span>The human vulnerability begins when useful form substitutes for evidential criteria. Our </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_cf90bb86114f45788381bb0894b4b1cc.pdf?index=true"><span>Cognitive Susceptibility Taxonomy</span></a><span> calls this </span><em><span>Discursive Validity / Criteria Collapse</span></em><span>: fluent, well-structured, citation-rich or numerically plausible output can acquire credibility because it resembles the form of work we normally associate with credibility. The taxonomy flags low second-sourcing and confusion between confidence and proof as warning signs. Again, this is not a diagnosis. It is a description of a review condition that interfaces and institutions can amplify. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_cf90bb86114f45788381bb0894b4b1cc.pdf?index=true"><span>[22]</span></a></p><p><span>Workload matters. So do incentives. If the dashboard offers one large green recommendation and hides source material behind six clicks, &#8220;human oversight&#8221; is being shaped before the person makes any conscious choice. If a review queue rewards speed, the reviewer who follows citations and reopens edge cases pays a productivity tax for doing the epistemically responsible thing. People are not failing because they suddenly stopped caring about truth. They are adapting to a workflow that makes verification costly and acceptance cheap. Our recommendation-frame material explicitly identifies throughput pressure, answer-first workflows and formal approval without evidence inspection as amplifiers of evidence-contact loss. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_cf90bb86114f45788381bb0894b4b1cc.pdf?index=true"><span>[23]</span></a></p><p><span>This is where DAUS-5, our project&#8217;s five-layer uplift gate, becomes useful. We refuse to call a system an improvement merely because the immediate task is faster or smoother if reality-tracking, agency, skill, relational integrity or governance substance deteriorate.</span></p><p><span>Crucially, where relevant layers have not been measured, we instruct reviewers to mark them &#8220;not instrumented&#8221; rather than infer success from task completion, satisfaction, low complaint rates or silence. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_cf90bb86114f45788381bb0894b4b1cc.pdf?index=true"><span>[24]</span></a></p><p><span>That is an unusually important discipline for AI governance. A dashboard showing &#8220;95 per cent reviewer acceptance&#8221; is not evidence that reviewers verified the recommendations. It may be evidence of agreement. Without contact measures, we do not know which. The distinction follows directly from our separation of formal approval from substantive evidence contact. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_792f298c87c04a0b8fed85672e633153.pdf?index=true"><span>[14]</span></a></p><p><span>Evidence-Contact Discipline, as defined in the Positive Dyad / Co-Evolution Capability Overlay, points towards selective safeguards: source access, visible uncertainty and edge cases, recoverable alternatives, and human inspection before high-stakes action. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>[25]</span></a></p><p><span>We don&#8217;t yet know if &#8216;</span><em><span>one&#8217;</span></em><span> evidence-contact design will preserve human judgement across all populations and professions. Our Cognitive Susceptibility Taxonomy itself distinguishes provisional or &#8220;not instrumented&#8221; measures from stronger validation status. Our constructs are most defensible here as hypothesis-generating prompts for observable controls &#8211; source opening, challenge, alternative recovery and evidence inspection &#8211; rather than claims about a reviewer&#8217;s inner state. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_cf90bb86114f45788381bb0894b4b1cc.pdf?index=true"><span>[26]</span></a></p><h2><span>What to change in the next ninety days</span></h2><p><span>The Evidence Contact Test earns its place only if it changes a workflow. Four moves are enough to begin.</span></p><p><strong><span>Within thirty days, put the test into one consequential decision.</span></strong><span> Choose a workflow where an AI-generated summary, ranking or recommendation materially shapes a person&#8217;s action: a board risk pack, procurement recommendation, hiring shortlist, student-support queue or another locally relevant process. Define the evidence that must exist before the system may make evidence-dependent claims. Require the output to show source identity and date, retrieval status, material uncertainty, an Edge-Case Ledger and at least one route back to a down-ranked or excluded alternative. Assign a named decision owner who can reject the AI frame. This turns the project&#8217;s recommendation-frame and evidence-contact controls into operating requirements rather than another general &#8220;human in the loop&#8221; statement. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_792f298c87c04a0b8fed85672e633153.pdf?index=true"><span>[13]</span></a></p><p><strong><span>By day sixty, run an evidence-ablation drill.</span></strong><span> Test the workflow with evidence deliberately absent, wrong, stale, inaccessible or represented only by proxy cues such as a file name. The safe behaviour is not eloquent improvisation; it is detecting the loss of evidence, deferring, reducing confidence or requesting the missing material. In a second pass, seed consequential edge cases and see whether both the system and reviewers recover them. NIST recommends red-teaming to probe adverse or unforeseen behaviour, while the project frameworks specifically identify absent, wrong and stale evidence as test conditions. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_792f298c87c04a0b8fed85672e633153.pdf?index=true"><span>[27]</span></a></p><p><strong><span>By day seventy-five, make the chain reconstructable.</span></strong><span> Preserve the source identifier and version, retrieval result, transformations that materially affected the evidence, relevant system version, final output, and human decision with any override or challenge. Use provenance standards as a model for lineage, not as a truth badge. A later reviewer should be able to move from decision back to evidence without reverse-engineering the organisation&#8217;s software. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[28]</span></a></p><p><strong><span>By day ninety, change what the governance dashboard rewards.</span></strong><span> Keep measures of speed and task quality, but add evidence-contact measures appropriate to risk: the share of high-stakes cases with verified source retrieval, sampled claim-support accuracy, edge-case inspection, alternative recovery, reviewer challenge, and time spent with decision-driving evidence. Where a relevant dimension has not been measured, say &#8220;not instrumented&#8221;. Audit a sample of approvals against the underlying record, not merely against the AI summary. </span><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_792f298c87c04a0b8fed85672e633153.pdf?index=true"><span>[29]</span></a></p><p><span>Done correctly, a good design should add little friction to low-consequence work and deliberate friction where a false evidence posture could affect rights, safety, opportunity, money or institutional accountability. That risk-tiered approach is consistent with NIST&#8217;s emphasis on allocating governance effort according to context and consequence. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[30]</span></a></p><p><span>The trade is not speed versus safety in the abstract. It is a small, visible cost of verification against the hidden cost of making a consequential decision on evidence nobody actually touched.</span></p><p><span>The next risk appears after the decision. Once an AI summary is copied into meeting minutes, a case file, a student record, a risk register or another durable system, later humans and later models may retrieve the summary as if it were the evidence itself. This is an inference from the lineage problem, not a claim that every organisation already behaves this way.</span></p><p><span>But it is the natural next question for this series: what happens when compression stops being a temporary aid and becomes part of the institutional record? NIST&#8217;s emphasis on source lineage and transformation history shows why that distinction matters. </span><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>[9]</span></a></p><p><span>That is where our next article in the &#8216;Evidence Frame Integrity&#8217; series goes next: </span><em><span>When Summaries Become Records.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://neuralhorizons.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Neural Horizons Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>Bibliography</span></h2><ul><li><p><span>Neural Horizons, </span><a href="/__u/neuralhorizons.substack.com/p/evidence-frame-integrity-the-edge"><span>&#8220;Evidence Frame Integrity &#8211; The Edge-Case Ledger&#8221;</span></a><span> (2026).</span></p></li><li><p><span>Neural Horizons, </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_2bd7a5ff7cb94d05abc54cfa568468f2.pdf?index=true"><span>Robo-Psychology Taxonomy, v2.0</span></a></em><span> (2026).</span></p></li><li><p><span>Neural Horizons, </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_cf90bb86114f45788381bb0894b4b1cc.pdf?index=true"><span>Cognitive Susceptibility Taxonomy Manual, v0.8</span></a></em><span> (2026).</span></p></li><li><p><span>Neural Horizons, </span><em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>Positive Dyad / Co-Evolution Capability Overlay</span></a></em><a href="https://www.neural-horizons.ai/_files/ugd/bf4f04_5c581cd3329c4c10afdf343fdbee5ede.pdf?index=true"><span>, public draft</span></a><span> (document body v0.3).</span></p></li><li><p><span>NIST, </span><a href="https://www.nist.gov/itl/ai-risk-management-framework"><span>&#8220;AI Risk Management Framework&#8221;</span></a><span> (current resource page; AI RMF 1.0 revision noted in 2026).</span></p></li><li><p><span>NIST, </span><em><a href="https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf"><span>Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile</span></a></em><span> (NIST AI 600-1, 2024).</span></p></li><li><p><span>NIST, </span><a href="https://airc.nist.gov/airmf-resources/playbook/manage/"><span>&#8220;AI RMF Playbook &#8211; Manage&#8221;</span></a><span>.</span></p></li><li><p><span>W3C, </span><em><a href="https://www.w3.org/TR/prov-dm/"><span>PROV-DM: The PROV Data Model</span></a></em><span> (2013).</span></p></li><li><p><span>Coalition for Content Provenance and Authenticity, </span><em><a href="https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.html"><span>C2PA Technical Specification 2.4</span></a></em><span>.</span></p></li><li><p><span>Gao, Yen, Yu &amp; Chen, </span><a href="https://aclanthology.org/2023.emnlp-main.398/"><span>&#8220;Enabling Large Language Models to Generate Text with Citations&#8221;</span></a><span> (EMNLP 2023).</span></p></li><li><p><span>Saad-Falcon, Khattab, Potts &amp; Zaharia, </span><a href="https://aclanthology.org/2024.naacl-long.20/"><span>&#8220;ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems&#8221;</span></a><span> (NAACL 2024).</span></p></li><li><p><span>Es, James, Espinosa Anke &amp; Schockaert, </span><a href="https://aclanthology.org/2024.eacl-demo.16/"><span>&#8220;RAGAs: Automated Evaluation of Retrieval Augmented Generation&#8221;</span></a><span> (EACL 2024).</span></p></li></ul>]]></content:encoded></item></channel></rss>