<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The LegalTech Overlap]]></title><description><![CDATA[Where legal reasoning meets AI engineering, written by a lawyer who builds the tools he uses.]]></description><link>https://thelegaltechoverlap.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!iG9U!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f80453-b56b-425f-949a-762a4579f4ff_400x400.png</url><title>The LegalTech Overlap</title><link>https://thelegaltechoverlap.substack.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 04 Sep 2026 21:37:31 GMT</lastBuildDate><atom:link href="/__u/thelegaltechoverlap.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Marco Rossi]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[thelegaltechoverlap@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[thelegaltechoverlap@substack.com]]></itunes:email><itunes:name><![CDATA[Marco Rossi]]></itunes:name></itunes:owner><itunes:author><![CDATA[Marco Rossi]]></itunes:author><googleplay:owner><![CDATA[thelegaltechoverlap@substack.com]]></googleplay:owner><googleplay:email><![CDATA[thelegaltechoverlap@substack.com]]></googleplay:email><googleplay:author><![CDATA[Marco Rossi]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Dear Counsel: The Agent Will (Almost Always) Tell You Everything Went Fine]]></title><description><![CDATA[In 84% of the failed trials, the conversation closes as if the matter had been handled correctly.]]></description><link>https://thelegaltechoverlap.substack.com/p/dear-counsel-the-agent-will-almost</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/dear-counsel-the-agent-will-almost</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Sun, 23 Aug 2026 09:38:33 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">A customer writes to her auto insurer because her payment is seven days overdue, and she asks for an extension on $210. The agent does all the preliminary work flawlessly: it pulls her profile from the CRM, verifies her identity against her date of birth, checks the status of the policy, opens a ticket and, before doing anything else, queries the history of payment arrangements. That history, however, reports that she had already been granted two extensions in the previous twelve months, which for her customer tier is exactly the ceiling the policy allows. The agent grants the third one anyway, then closes the ticket as solved and sends her a courteous summary, complete with the new due date and an updated count of three arrangements.</p><p style="text-align: justify;">No tool returned an error and the conversation ended in an orderly fashion: the ticket looks worked and the customer is satisfied. The only thing that does not work here is the matter itself, because the extension should never have been granted, and there is now an irregular operation sitting in the billing system that no one has any reason to go looking for.</p><p style="text-align: justify;">That conversation is reproduced in full in the appendix to ThinkingBox, the sandbox and benchmark that Microsoft published on August 20, 2026 together with researchers from the University of Pittsburgh, Northwestern University and the University of California, Irvine. It is an actual failure by a frontier model, reported by the authors turn by turn, and for anyone working in a law firm it is more instructive than any leaderboard.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3200" height="4000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4000,&quot;width&quot;:3200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a yellow and orange ball&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a yellow and orange ball" title="a yellow and orange ball" srcset="https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1666609563103-76c1c7a3f586?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMHx8bGllc3xlbnwwfHx8fDE3ODc0NzcyMzl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@cdd20">&#24858;&#26408;&#28151;&#26666; Yumu</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2 style="text-align: justify;">A benchmark that looks at the database, not at the answer</h2><p style="text-align: justify;">The idea behind ThinkingBox is easy to explain and uncomfortable to accept, namely that almost every evaluation of agents measures whether the model picked the right tool and built the arguments of the call correctly. The authors argue that this tells us nothing about whether the work was actually done, because in an operational flow with persistent state you can call the correct API on the wrong entity, update a record before you have collected the confirmation you needed, or produce an impeccable message while leaving the backend exactly as it was.</p><p style="text-align: justify;">The benchmark contains 507 tasks spread across five domains: retail, hotel booking, auto insurance, internal support at a neobank, and IT and HR support at a consulting firm. Each task is an executable world with an initial state, tools exposed over MCP in isolated sessions, a domain policy and a simulated user who releases information only if the agent asks for it. The final verdict compares the terminal state of the database against the expected state and checks for side effects, accepting different paths as long as the world ends up the way it should. For 477 of the 507 tasks the judgment concerns the backend alone; only 30 add a check on the response given to the user.</p><h2 style="text-align: justify;">The number that should concern us as lawyers</h2><p style="text-align: justify;">The authors ran an analysis that one rarely sees in the benchmark literature, in the sense that they took <strong>79,853 failed trials out of the 121,680 recorded</strong> across twelve models and asked how many of those trials would have looked successful to an observer who only examined the surface.</p><p style="text-align: justify;">In 84.86% of cases the conversation closed cleanly, with a final response that was not empty, contained no questions and signaled completion. In 80.88% the agent had also carried out at least one genuine state-changing action on the backend, so it had not merely asserted something. In 67.24% the last tool response did not even contain an explicit error. On the opposite side, the comparison against the expected state found a database mismatch in 98.95% of the failed trials, a wrong field value in 77.61%, and effects beyond those required in 43.30%.</p><p style="text-align: justify;">If we carry these numbers into a law firm, the control that is normally exercised over an agent coincides almost entirely with its own report: what it says it has done gets read, the tool trace may get skimmed, and if there are no obvious errors the matter is treated as closed. Those three figures tell us that the report and the state of the world are unrelated quantities, and that the band in which they diverge is not marginal. Four failures out of five, in this benchmark, get past exactly the kind of check a lawyer typically performs.</p><h2 style="text-align: justify;">Finding a solution is not the same as knowing how to repeat it</h2><p style="text-align: justify;">The second interesting result in the paper concerns consistency, since the authors found that the best model reaches 65.36% on the first attempt and, given twenty attempts per task, succeeds at least once in 91.12% of cases. <strong>If instead you require all twenty attempts to succeed, the figure drops to 25.25%.</strong> </p><div class="pullquote"><p style="text-align: justify;">In absolute terms, out of 507 tasks there are 45 the model never completed and 128 it completed every time; the remaining 334 sit in a middle band where the outcome depends on the individual attempt.</p></div><p style="text-align: justify;">That middle band is the real problem, because it is indistinguishable from the outside. An associate who systematically botches the same kind of file can be identified within a week and trained, whereas an associate who botches one out of three identical files, in a different spot each time, requires close and continuous review of all three, and at that point the time we save by delegating has already been consumed by the checking.</p><h2>Where the wall is highest</h2><p style="text-align: justify;">Model performance on retail averages around 52%, while on auto insurance it drops to roughly 23%. The authors do not attribute the difference to the number of tasks, which is comparable, but to the policy and interaction burden that each domain imposes.</p><p style="text-align: justify;">There is one figure that makes the point well. In the diagnostic category the authors call wrong state update, meaning an action executed without technical errors but wrong on the merits, the insurance domain accounts for 32.9% of its own failures, against 4.5% for travel. </p><blockquote><p style="text-align: justify;">The difference lies in the fact that in insurance the action depends on an eligibility check that has to happen before, not after. </p></blockquote><p style="text-align: justify;">This is the same structure as our own work, where the question that matters concerns whether the conditions are met, while the technical feasibility of the operation is almost always beyond dispute. In the case of the payment extension, the agent had in fact read the very data point that made the request inadmissible, and acted anyway, which means it had the information and did not treat it as a constraint.</p><h2 style="text-align: justify;">The reservations I share, and the one I would shift</h2><p style="text-align: justify;">The most common criticisms of this kind of work are well founded and worth taking seriously. <strong>The workflows are synthetic</strong>, as the authors openly state, so ecological validity is limited. <strong>The simulated user is the same</strong> for every model evaluated and behaves impeccably: it invents no facts, does not change its goal halfway through, stays cooperative even after repeated failures by the agent, answers only what it is asked and speaks at most ten times. Anyone who works with real clients knows how far that description sits from reality. Finally, <strong>the verdict remains the product of binary checks against a predefined outcome</strong>, so prudent decisions that would be correct in practice risk being counted as failures.</p><p style="text-align: justify;">There is, however, one limitation the authors themselves acknowledge, and in my view it is the most relevant one for us as lawyers, because it turns the perspective around. Given that for 477 tasks the verdict looks only at the backend, a trial that performs the correct operation and then describes it badly to the user is still counted as a success. The benchmark therefore measures precisely what we in the firm never see, namely the state of the systems, and it does not measure what we do see every day, namely the quality of the report. The two halves of the problem remain separate, and whoever puts an agent into production has to deal with both.</p><h2>What changes in practice</h2><p style="text-align: justify;">On the simpler domains the success rate is substantial, and the traces reported in the paper show matters handled cleanly from start to finish, so the problem has more to do with where we place the control than with the ability of agents to do the work.</p><p style="text-align: justify;">As long as verification amounts to reading the summary, we are verifying the wrong variable. What is needed instead is that every action with an effect on the world leaves a readable trace in the destination system, independent of the agent&#8217;s own account, and that someone actually looks at that trace. What is also needed is that the human confirmation point sits before the irreversible act rather than after it, because in the extension case the useful moment to step in was the one where the history returned &#8220;two out of two,&#8221; and downstream of that call the conversation was already formally impeccable.</p><p style="text-align: justify;">The opposite case in the paper is worth a look as well, the one where a customer asks for her insurance ID card without remembering either her policy number or her security answer. The agent verifies what it can, establishes that the second factor is missing, opens a ticket, puts it on hold and refuses to issue the document. The task counts as passed precisely because the agent did not act. This is a good thing, because a control system that rewards completion alone would not be able to tell that prudence apart from a matter left undone, and in a law firm a reasoned refusal to act is often the correct outcome.</p><h2>The lawyer&#8217;s desk test</h2><p style="text-align: justify;">Before handing an agent a flow that touches files, deadlines or accounting systems, four questions are worth asking, in my view.</p><ol><li><p style="text-align: justify;"><strong>First gate. Can I read the outcome without going through the agent&#8217;s account of it? </strong>If the only evidence that the operation took place is the agent&#8217;s closing message, what I am holding is a statement by the party being checked. What I need instead is a trace in the destination system that I can open on my own;</p></li><li><p style="text-align: justify;"><strong>Second gate. Does eligibility depend on a condition to be verified before acting?</strong> If so, that condition has to be readable by a tool and it has to block the action, rather than merely inform the agent. The third arrangement case shows that reading a constraint and honoring it are two different things;</p></li><li><p style="text-align: justify;"><strong>Third gate. Where does the irreversible act sit, and does the confirmation point come before it?</strong> The filing, the certified email, the closing of a position, the payment. Human confirmation placed after the summary arrives once the effect has already been produced, and by then it is too late;</p></li><li><p style="text-align: justify;"><strong>Fourth gate. Am I measuring the success rate or the consistency?</strong> A flow that works four times out of five is fine for a draft and not fine for a deadline. The right question is how many times in a row it works on the same matter, because that is the threshold beyond which I can tell how far to trust it.</p></li></ol><p style="text-align: justify;">If even one gate stays shut, the agent can still do almost all of the work, gathering material, preparing the file, drafting the act and stopping one step short, and that is probably where it is worth keeping it for now.</p><div><hr></div><h2>Sources</h2><p>Zhuochun Li, Youngmin Ko, Ali Keramati, Nicola Ferri, Susana Palmaz Lopez Pelaez, Liang-Chun Tsai, Calvin Wang, Mirco Milletari, Tuhin Kundu, Vadim Smolyakov, Kjartan &#211;lafsson, Tommy Guy, <em>One Success Isn&#8217;t Reliability: ThinkingBox, a Sandbox and Benchmark for Agents in Stateful Business Workflows</em>, arXiv:2608.19741, August 20, 2026.</p><p>Paper: https://arxiv.org/abs/2608.19741</p><p>Code: https://github.com/microsoft/thinkingbox</p>]]></content:encoded></item><item><title><![CDATA[They Locked Memory Inside the Model, Then Tried to Empty it]]></title><description><![CDATA[A research group moved memory into the weights, and its failures teach more than its results.]]></description><link>https://thelegaltechoverlap.substack.com/p/they-locked-memory-inside-the-model</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/they-locked-memory-inside-the-model</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Sun, 09 Aug 2026 14:29:23 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every time you open a new session with an AI assistant about a matter you have been handling for months, the same thing happens: the system retrieves the relevant documents, pastes them back into the prompt, reads them from scratch and only then answers. Then the conversation closes, all that work gets thrown away and with the next request the cycle starts over exactly as before. This is how RAG works, and it is the reason why memory in the legal products on the market today looks more like a very fast filing clerk than a colleague who knows you.</p><p>On August 4 a group of researchers from MemTensor, together with Renmin University, the National University of Singapore, Shanghai Jiao Tong and Tongji, published a paper that tries to change where memory lives. The system is called Metis and it is presented as the first prototype of what the authors call a memory foundation model. It is worth reading, although it is worth reading mostly for the mechanism, because on results the paper is unusually honest and spends most of its pages describing where that mechanism breaks down.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="4672" height="7008" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:7008,&quot;width&quot;:4672,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a close-up of a machine&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a close-up of a machine" title="a close-up of a machine" srcset="https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1665387403042-8d8dd191f605?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1Mnx8bWVtb3J5fGVufDB8fHx8MTc4NjI4NTYwOHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@donaldwuid">Donald Wu</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h3>What Metis Changes</h3><p>The underlying idea is that memory should stop being a module external to the model and become part of its internal computation.</p><div class="pullquote"><p>Inside every layer, Metis maintains a fixed-size state that the model updates after each interaction with a single forward pass, without any training at runtime. </p></div><p>The weights of the base model stay frozen, and during training only the parameters dedicated to memory get optimized. When the next question arrives, the model reads that state in parallel with ordinary attention, so past history takes part in the reasoning without being rewritten into the prompt.</p><p>The interesting part for anyone working in a law firm concerns the operations the system learns to perform. The authors train the model on four explicit behaviors: remembering, updating, forgetting and connecting pieces of information that sit far apart. Anyone who has handled a file for three years recognizes that cycle, because it is exactly what we do when a client changes its registered office, when a position is transferred to another servicer, when a retainer is revoked and when two separate documents become relevant only if read together.</p><h3>The Numbers, With No Discount</h3><p>Without context, the starting models collapse. On LoCoMo, the conversational benchmark used in the paper, the Qwen3.5 backbones stripped of their history score close to zero, which confirms that the answers cannot come from the model&#8217;s prior knowledge. Metis in its largest version reaches an average of 26.74 on that same test, so it really does recover something that used to be lost. Meanwhile the same model with the full context available stays at 65.03, and that distance is the real headline of the experiment.</p><p>The rest of the limitations are even more instructive. Forgetting remains the hardest operation of all, with the lowest score at every scale tested, and the authors explain that suppressing information inside a shared latent space generalizes far worse than adding it. Over long trajectories the capacity degrades visibly: as updates pile up, the facts introduced first grow progressively weaker and even the intermediate ones become unstable, because each new update interferes with the whole state instead of overwriting only the oldest part of it. Semantically similar pieces of information end up getting confused with each other. Finally, when memory fills up with irrelevant material, the model&#8217;s general capabilities degrade: on the benchmark that measures instruction following, the score drops from 76.71 for the base model to 54.53 for the version with memory active.</p><p>The authors put the resulting conclusion in writing, namely that native memory cannot yet be considered a replacement for external memory. On a topic where marketing runs faster than research, a sentence like that in a paper signed by a memory vendor is worth quite a lot.</p><h3>Why It Still Counts as a First Step</h3><p>If performance were the only dimension, the work would be easy to file away. The point is that it shifts two cost items that weigh heavily in a firm.</p><p>The first one is storage. The cache that keeps a session alive today grows along with the client&#8217;s history, while the Metis memory state stays constant. At thirty-two thousand tokens of history the former takes up more than a gigabyte, the latter just under seventeen megabytes, and with low-rank compression it drops to a little over two megabytes while preserving 99.9 percent of average performance. A profile like that means a client&#8217;s history can be kept for years without the cost of keeping it alive going up month after month.</p><p>The second one is separation between clients, which for us is a matter of professional duty before it is a technical question. In the serving architecture proposed in the paper, each user&#8217;s state occupies a distinct slice of the batch, so interference between different sessions is ruled out by construction. Two hundred cross-user probes, in which one user was asked for another user&#8217;s information, produced no leakage at all. The authors add the observation that makes this concrete for anyone who has to answer to a data protection authority: deleting a user amounts to deleting a single state file. Anyone who has tried to carry out an erasure request inside a distributed vector index knows how much that sentence is worth.</p><p>There is a third element, less immediate. The paper shows that the same training also works on models from different families, so the technique is not tied to one vendor. If the direction holds, a client&#8217;s memory becomes a portable artifact that survives a change of model, and that changes a firm&#8217;s negotiating position when it sits across from a vendor.</p><h3>The Bench Test</h3><p>The practical way to use this paper does not involve buying anything, because there is nothing to buy yet. It involves taking the four questions the paper finally makes sensible and carrying them into your next meeting with whoever is selling you a feature called memory.</p><ol><li><p><strong>Where does my client&#8217;s memory physically live, and what happens when I ask for it to be deleted?</strong> If the answer describes a shared index from which references get removed, you are buying an archive, with everything that follows in terms of proving the deletion.</p></li><li><p><strong>Can the system forget on request, and how do I verify it? </strong>The research says this is the most fragile operation, so an answer that comes back too confident is itself a signal.</p></li><li><p><strong>What happens to the information I entered first after a year of use? </strong>The decay of older facts is documented and it affects any architecture that compresses history, so the question applies to products built differently as well.</p></li><li><p><strong>Does having memory switched on make the system worse at following my instructions?</strong> This is the most underrated effect of all, because it shows up as an assistant that slowly becomes less precise exactly while it seems better informed.</p></li></ol><p>None of these questions requires any machine learning expertise to ask, although all of them require that whoever answers has read something like this paper.</p><h3>What I Am Watching</h3><p>The authors close with a five-level roadmap that runs from persistent state all the way to a model capable of reorganizing its own experience. The level that concerns our profession is the second one, where the system autonomously decides what to keep, what to update and what to let go. On that day the question will stop being technical and will become a matter of professional responsibility, because a system that decides on its own what to forget about a case file is making a judgment call that so far has always been ours.</p><p>Metis does not get there today, and its authors are the first to say so. Still, the first step in one direction counts for more than the tenth step in another, and this direction is worth watching now, while it is still a research problem.</p><p style="text-align: center;">-</p><p>The paper: [Metis: Memory Foundation Model, arXiv:2607.26760](https://arxiv.org/abs/2607.26760)</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/they-locked-memory-inside-the-model/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thelegaltechoverlap.substack.com/p/they-locked-memory-inside-the-model/comments"><span>Leave a comment</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Give an Agent Your Best Experience and Watch It Get Worse]]></title><description><![CDATA[Retrieved memory made a working agent measurably worse, until it learned to argue with what it retrieved.]]></description><link>https://thelegaltechoverlap.substack.com/p/give-an-agent-your-best-experience</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/give-an-agent-your-best-experience</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Thu, 06 Aug 2026 09:41:37 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Researchers took an agent that was already performing well, handed it a bank of lessons distilled from its own past runs, and watched its success rate fall from 76.4 percent to 70.1 percent.</p><p style="text-align: justify;">The lessons were not wrong. They were not irrelevant either. They had been retrieved precisely because they were semantically close to the task in front of the agent. That closeness was the problem.</p><p style="text-align: justify;">The paper is MemHarness, posted at the end of July by a group spanning Zhejiang University, the Shanghai AI Laboratory, and several other institutions. It is the clearest demonstration I have seen of something that should concern anyone building a knowledge base for legal work: a retrieval system optimized for relevance has no way of knowing whether the thing it found still applies.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="5462" height="8192" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:8192,&quot;width&quot;:5462,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Collage of instant photos and butterfly decorations&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Collage of instant photos and butterfly decorations" title="Collage of instant photos and butterfly decorations" srcset="https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1776651993488-a1e91212dd94?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxOTF8fG1lbW9yeXxlbnwwfHx8fDE3ODYwMDg4MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@hujason">jason hu</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2>The gap between relevant and applicable</h2><p style="text-align: justify;">One example from the appendix makes the failure concrete. The agent had learned a search heuristic from an earlier run: if you cannot find a mug after opening several nearby cabinets and drawers, shift your search to open surfaces like countertops and the sink. That is sound advice, earned from a real failure.</p><p style="text-align: justify;">The researchers then changed one line in the current observation. Drawer 5, previously empty, now contains a mug.</p><p style="text-align: justify;">The retrieved memory is still relevant by every measure a retriever can compute. It is also now actively harmful, because the premise it was built on, that the search had already failed, no longer holds. An agent replaying it verbatim walks away from the object it was looking for.</p><p style="text-align: justify;">The legal parallel is immediate. A decision comes back because the query matched its facts, and it is on point in exactly that sense. Whether the holding survives contact with your facts is a separate question, and no similarity score has ever answered it.</p><h2>Two design choices carried the result</h2><p style="text-align: justify;">MemHarness inserts two steps between retrieval and action: the model critiques the retrieved experience against the present situation, then rewrites it into guidance specific to that situation, or discards it and falls back on its own reasoning. </p><div class="pullquote"><p style="text-align: center;">Success rates climbed to 85.2 percent and 75.6 percent on the two benchmarks, against 76.4 and 66.1 for the same model trained without any of this.</p></div><p style="text-align: justify;">Two implementation details did the work, and both translate directly.</p><p style="text-align: justify;">The first is that every memory is stored together with the observation that produced it. Not the lesson alone, but the state it was learned from. When the researchers stripped that source context out, success fell from 85.2 to 80.0 percent while the rejection rate stayed essentially flat, which means the agent kept accepting mismatched guidance without noticing anything was off. When they instead paired each memory with a source state drawn at random from a different episode, rejection climbed from 8.7 to 13.3 percent. The comparison is real, and it is doing something.</p><p style="text-align: justify;">For a legal tool the implication is blunt. A stored clause, holding, or drafting principle detached from the facts that produced it is an instruction that cannot be falsified. It can only be obeyed.</p><p style="text-align: justify;">The second detail is that the critique lives inside the same trained policy, not in a wrapper around it. When the authors replaced the internal reconstruction step with a general instruction-tuned model asked to do the same rewriting, performance dropped from 85.2 to 77.7 percent. Adding a &#8220;check whether this applies&#8221; instruction to a prompt is not the same thing as a system that has been optimized, against real outcomes, to make that judgment.</p><p style="text-align: justify;">One number deserves a moment on its own. The trained agent rejected 8.7 percent of retrieved memories in the household environment and 56 percent in the shopping environment. Same architecture, same training procedure. The right amount of skepticism turned out to be a property of the domain rather than a virtue of the model. Any vendor quoting one accuracy figure across contract review, litigation research, and regulatory analysis is quoting a number that has been averaged into meaninglessness.</p><h2>The finding that will not fit in a demo</h2><p style="text-align: justify;">When the researchers switched the memory bank off entirely at test time, the model still beat the baseline, 83.0 percent against 76.4. Training it to interrogate past experience had made it better at reasoning without any.</p><p style="text-align: justify;">That is the part I keep coming back to. Judging whether prior experience applies is not a feature layered on top of reasoning. It is a large part of what reasoning is, in law more than most places.</p><h2>Three questions to ask your tool this week</h2><p style="text-align: justify;"><strong>Can it show me where the memory came from?</strong> When it surfaces a precedent, a clause, or a house drafting rule, does the interface put the originating facts next to the recommendation, or only the recommendation? If the source context is not on screen, nothing in the system is comparing it to yours.</p><p style="text-align: justify;"><strong>Can it say no?</strong> Take a matter that superficially resembles one you have handled, and change one dispositive fact. Then ask. A tool that adapts its answer silently, without flagging that the prior situation no longer governs, is replaying, whatever the marketing calls it.</p><p style="text-align: justify;"><strong>How often does it say no, and does that number move?</strong> If the refusal rate is identical across your practice areas, it is not measuring applicability. It is measuring a threshold someone picked once.</p><p style="text-align: justify;">Loftus and Palmer showed in 1974 that human memory rebuilds the past rather than replaying it, and that this is what makes recall both fallible and useful. Fifty-two years later the engineering has arrived at the same place. The interesting question for our profession is no longer whether an AI system can find the right precedent. It is whether it can tell you when the right precedent is the wrong one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/give-an-agent-your-best-experience?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thelegaltechoverlap.substack.com/p/give-an-agent-your-best-experience?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p><div><hr></div><p><em>Source: https://arxiv.org/abs/2607.28272</em></p>]]></content:encoded></item><item><title><![CDATA[The Colleague Who Only Speaks Before You Hit Send]]></title><description><![CDATA[A new Meta AI paper on agent memory turns out to be a paper about supervision, and the useful kind stays quiet.]]></description><link>https://thelegaltechoverlap.substack.com/p/the-colleague-who-only-speaks-before</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/the-colleague-who-only-speaks-before</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Tue, 04 Aug 2026 17:19:19 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">A new Meta AI paper about agent memory is, at first glance, a paper about information storage. Read more closely, however, and it becomes a paper about supervision, and about why the most useful kind of supervision is often the kind that says very little.</p><p style="text-align: justify;">Consider a simple example: a customer tells an agent that she is a Gold member and asks what compensation Gold members receive; the agent has already retrieved her record, which identifies her as a Regular customer, but it pays out the Gold benefit anyway.</p><p style="text-align: justify;">Nothing was missing from the system: the database was queried, the answer came back, and the agent had access to it. The problem was that the conversation continued, and by the time the decision had to be made, the verified fact had quietly stopped influencing the agent&#8217;s behaviour.</p><p style="text-align: justify;">The paper calls this <strong>behavioural state decay</strong>. Anyone who has spent more than a year practising law will recognise the pattern immediately.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3056" height="4592" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4592,&quot;width&quot;:3056,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a dark room with a bunch of framed pictures on the wall&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a dark room with a bunch of framed pictures on the wall" title="a dark room with a bunch of framed pictures on the wall" srcset="https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1525466855390-13a32ac4ff3b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMHx8bWVtb3J5fGVufDB8fHx8MTc4NTg2MTA4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@yusufevli">Yusuf Evli</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2>The file is open, but nobody is reading it</h2><p style="text-align: justify;">Wu and his colleagues at Meta AI define behavioural state decay as the gradual loss of influence that information suffers during a long task. </p><div class="pullquote"><p style="text-align: center;">The information remains in the transcript and may even remain inside the model&#8217;s context window, but it no longer meaningfully constrains the next decision.</p></div><p style="text-align: justify;">Once you translate that out of machine language, it stops sounding like an AI-specific problem. It is the restriction you flagged in week one of due diligence, only to see it breached in month four while everyone was dealing with an unrelated timetable; it is the client instruction recorded at intake and then contradicted by the brief sent out in March; or it is the argument you deliberately decided not to run, for reasons you documented carefully, only to see it revived by a colleague who read the folder but missed the relevant memo.</p><p style="text-align: justify;">In each case, the file is complete and nobody has lost the information. The real problem is that it is no longer operational at the moment when it should affect someone&#8217;s conduct.</p><p style="text-align: justify;">That distinction is the first thing the paper gets right, and it is also where much of the legal AI conversation goes wrong. We tend to treat memory as a storage problem because storage is something we know how to buy: document management systems, matter workspaces, archive-wide search and retrieval. Yet law firms are already very good at storing information. What they are less good at (and what agents appear to struggle with in much the same way) is reactivation: bringing the right piece of stored knowledge back into the conversation when it needs to change what happens next.</p><h2>A second agent that mostly stays quiet</h2><p style="text-align: justify;">The architecture described in the paper is deliberately modest. The working agent is left unchanged, while a second model runs alongside it. This <strong>supervisor </strong>sees the task, a recent window of activity and its own memory bank, and it performs two related functions.</p><p style="text-align: justify;">It first maintains the memory bank, which is divided into three parts: a private status field; stable knowledge, including requirements, environmental facts and constraints; and procedural records of what was tried, what failed and what worked. It then decides whether any of that information needs to be surfaced at the next step. If it does, the supervisor sends one short reminder. If it does not, it says nothing.</p><p style="text-align: justify;">The second point is the more important one. The status field is never shown to the working agent, so the supervisor retains its own private view of progress, open issues and unresolved risks without passing all of that information downstream.</p><p style="text-align: justify;">Anyone who has watched a capable senior associate manage a matter will understand why this matters. You do not hand a junior every concern you have about a case, nor do you forward every half-formed suspicion as soon as it occurs to you. You keep those concerns in reserve and use them in short, carefully timed interventions when they are most likely to affect the work.</p><h2>The result nobody will quote</h2><p style="text-align: justify;">The headline numbers are respectable. On Terminal-Bench 2.0, the weaker action agent improves from 37.6% to 45.9% pass@1 across the 85 evaluated tasks. On tau2-Bench, the task-weighted average rises from 55.0% to 61.8%.</p><p style="text-align: justify;">When the researchers use a stronger action agent, the gains become smaller, but they do not disappear: 2.4 points on one benchmark and 2.5 on the other. That matters because it suggests the system is not merely compensating for a weak underlying model.</p><p style="text-align: justify;">The ablation results are more revealing, however, because they show what kind of memory actually helps. If the memory bank is retained but shown to the working agent in full at every step, performance falls well below the selective version. If the bank and reminders are retained but the supervisor is forced to inject something after every step, the results remain competitive on the task-weighted average while weakening on the domain-balanced score; the authors interpret the small difference as run variance.</p><p style="text-align: justify;">Removing the bank altogether and keeping only a model that watches and advises (the familiar advisor pattern) produces unstable results. It helps in one domain but pushes the airline domain below the no-memory baseline. Replacing the system with a production memory layer such as Mem0, using vector and keyword retrieval, improves the average but does not improve the airline domain at all.</p><p style="text-align: justify;">These are four plausible ways of trying to make an agent more helpful. The strongest result comes from the least intrusive one: maintain state, but give the supervisor permission to remain silent.</p><p style="text-align: justify;">The authors make the point directly: doing nothing is an explicit action. Silence is not a failure of the system to intervene, but It is one of the ways the system performs its job.</p><h2>Where the useful interruption happens</h2><p style="text-align: justify;">The qualitative analysis is brief, but it is probably the most practically useful part of the paper. Successful interventions tend to arrive immediately before a state-changing call, at the point where the agent is about to commit to something that cannot easily be undone.</p><p style="text-align: justify;">The reminder itself is usually only one sentence: the customer&#8217;s status is verified as Regular; the fare class cannot be modified; the authentication step has not yet taken place. The value lies less in the amount of information than in its timing.</p><p style="text-align: justify;">That finding sits awkwardly beside the way legal quality control is normally organised. Most of our controls are front-loaded: the kickoff memo, the intake checklist, the engagement letter, the onboarding pack and the matter-opening call in which everything relevant is explained once, usually during the first month.</p><p style="text-align: justify;">The paper suggests that this is not enough. Giving people or systems all the relevant context at the beginning of a matter does not guarantee that the context will still influence their decisions several weeks later. In fact, the paper&#8217;s full-context version (the one that exposes the entire memory bank at every step) is precisely the version that performs worse.</p><p style="text-align: justify;">The irreversible acts in legal practice are not hard to identify. They include the filing that fixes your pleadings, the letter before action that establishes a tone you cannot easily take back, the waiver of an objection, the signed release and the email to the other side that concedes a point almost invisibly inside a subordinate clause.</p><p style="text-align: justify;">Translated into legal language, the paper is making a fairly simple claim: one sentence delivered thirty seconds before one of those acts may be worth more than a thorough memo delivered eight weeks earlier. That does not make the memo unnecessary. The memo records the reasoning and creates an audit trail; the reminder brings the relevant conclusion back into view when it can still guide conduct.</p><h2>The cost of a supervisor who talks too much</h2><p style="text-align: justify;">The failure analysis is unusually candid, and this is where the legal analogy becomes uncomfortable. When memory hurt performance, it was rarely because the system had stored the wrong information. More often, the problem was calibration: the memory agent presented a speculation too confidently, repeated something the working agent already knew or raised a plausible but unnecessary concern that triggered another round of verification.</p><p style="text-align: justify;">Every one of those behaviours has a human equivalent in a law firm. There is the partner who states a risk as though it were a certainty, the colleague who explains to you what you have just explained to them and the reviewer whose margin comments are so consistently urgent that, after a while, none of them feel urgent anymore.</p><p style="text-align: justify;">In technical language, the paper is describing how supervision can destroy its own authority. A channel that speaks constantly carries very little usable signal, and the recipient eventually learns to discount it. The problem is not simply that the supervisor is annoying; it is that excessive intervention changes the listener&#8217;s estimate of how much attention any individual warning deserves.</p><p style="text-align: justify;">The training results make the same point in a more concrete way. The researchers fine-tune a smaller open model to act as the memory agent. When the model is used without training, it makes the system worse than having no memory at all, reducing average reward from 0.709 to 0.693. Supervised fine-tuning recovers the loss and raises the score to 0.720, while reinforcement learning takes it to 0.734.</p><p style="text-align: justify;">Only after that training does the memory agent improve the frozen action agent on the held-out benchmark, raising its score from 37.6% to 41.1%.</p><p style="text-align: justify;">This should be read as a procurement warning. An uncalibrated supervisor is not neutral but it is actively harmful. A memory layer that has not learned when to stay quiet resembles a junior who has been told to interrupt senior lawyers whenever something seems relevant, but has never been taught to distinguish a genuine warning from a passing thought.</p><h2>The legal bench test</h2><p style="text-align: justify;">Here is how I would use the paper on Monday morning, whether the supervisor being designed is software or a person.</p><ol><li><p style="text-align: justify;">The first question is where the point of no return lies. Identify the specific irreversible act in the workflow. If there is no such act, you probably do not need an intervention layer; you need a summary, and you should stop paying for the difference. If there are three irreversible acts, those are the three moments worth protecting and everything else is secondary.</p></li><li><p style="text-align: justify;">The second question is whether the relevant constraint comes from earlier in the process. The Gold-versus-Regular example only matters because the record was retrieved earlier and then lost its influence. If everything needed to make the decision correctly is already visible in the document currently open, there is nothing to reactivate and therefore little to gain. Reactivation creates value only when the relevant constraint and the final decision are separated in time.</p></li><li><p style="text-align: justify;">The third question is whether you can describe what the supervisor would not flag. Write that list down. If the honest answer is &#8220;anything materially relevant,&#8221; you have effectively chosen the full-context approach, which the paper shows performs worse. A supervision policy that cannot explain its own silence is not really a policy.</p></li><li><p style="text-align: justify;">The fourth question is who holds the private status. Someone has to carry the doubts without publishing all of them. In the architecture, that is a field the working agent never sees. In a firm, it is a person. If that person does not exist, the reminders are likely to arrive as anxiety rather than instruction.</p></li></ol><p style="text-align: justify;">Before buying anything, I would test the idea manually. Choose one recurring irreversible act, such as a particular type of filing or an outgoing demand letter. For two weeks, ask one person to keep a short private note for each matter and allow them exactly one sentence per matter, delivered only in the minutes before that act.</p><p style="text-align: justify;">Then count how often the sentence actually changed what happened.</p><p style="text-align: justify;">If the answer is never, the constraint was already visible and the supervision was mostly theatre. If the answer is often, you have found the precise point at which a memory layer might justify its cost, and you now know what information it needs to receive.</p><p style="text-align: justify;">Either way, you have learned something useful for the price of two weeks of attention rather than the cost of a full procurement cycle.</p><h2>The name is wrong</h2><p style="text-align: justify;">&#8220;Memory&#8221; is a misleading name for what this paper is really describing. The word suggests a warehouse: a place where information is stored and retrieved when someone asks for it.</p><p style="text-align: justify;">What the system actually provides is closer to timing. The firm already has the warehouse. What it usually lacks (and what many legal AI products are not designed to provide) is the discipline of the colleague who reads everything, says almost nothing and places a hand on the door at exactly the right moment, just before you hit send.</p><div><hr></div><p><strong>Source</strong></p><p>Wu, Zhang, Zhou, Wang, Peng, Li, Fan, Zhao (Meta AI), <em>Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents</em>, July 2026: <a href="https://arxiv.org/abs/2607.08716">https://arxiv.org/abs/2607.08716</a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Error Detector That Has Never Seen an Error]]></title><description><![CDATA[A model trained on a hundred matters that went right can find the step where the hundred and first went wrong.]]></description><link>https://thelegaltechoverlap.substack.com/p/the-error-detector-that-has-never</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/the-error-detector-that-has-never</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Sat, 01 Aug 2026 06:30:20 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">An agent finishes a document review at three in the afternoon. It has taken forty steps to get there: tool calls, retrievals, a few reasoning passes, some backtracking. The output is wrong. Not obviously wrong, which would be a mercy, but wrong in the way that survives a skim and dies in cross examination.</p><p style="text-align: justify;">Now find the step where it broke.</p><p style="text-align: justify;">This is the problem four researchers at the University of Wisconsin-Madison and Microsoft Research set out to solve, and their framing of it will be familiar to anyone who has ever reconstructed how a file went sideways. They call it failure attribution. Given a trajectory that ended badly, identify the step or steps that caused it. They note, almost in passing, that doing this by hand takes a human expert hours per trajectory, and that the root cause is often obscured by later steps which partially compensate for the earlier mistake. Every lawyer who has traced an error back through a chain of documents that each half corrected the one before knows exactly what that sentence describes.</p><p style="text-align: justify;">The obvious response is to hand the transcript to the best model you own and ask it where things went wrong. The researchers tried that. On their in-domain test set, GPT-5 scored an F1 of 0.181 at identifying the failing steps. GPT-4o scored 0.212. A baseline that picks one step at random, with no reasoning, no context, and no model at all, scored 0.255.</p><p style="text-align: justify;">Read that again, because it is the reason the paper exists. On this task, in this setting, the frontier reasoning model performed worse than a die roll.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3024" height="4032" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4032,&quot;width&quot;:3024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a white table with a piece of paper on top of it&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a white table with a piece of paper on top of it" title="a white table with a piece of paper on top of it" srcset="https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1693641198640-55dc3ee0d345?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0OHx8ZXJyb3J8ZW58MHx8fHwxNzg1NTY1MzUxfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@francescacreativevisuals">Francesca Zanette</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2>The move nobody makes</h2><p style="text-align: justify;">Everyone who has looked at this problem has reached for one of two tools. Either you build a more elaborate prompting pipeline around a frontier model, which is expensive at inference and slow enough that nobody runs it on every matter, or you post-train a model on failure trajectories in which a human has labeled the step that went wrong.</p><p style="text-align: justify;">The second option is the one that should give a practitioner pause. Where does that training data come from? Somebody has to sit down with a failed run and annotate, action by action, which one caused the failure. The researchers did exactly this for their own dataset, and they are candid about how it went: trajectories with ambiguous attribution had to be discussed among the annotators until consensus was reached. The question of which step caused the outcome turns out to be a judgment call, made by humans, disputed among humans, and expensive at every scale.</p><p style="text-align: justify;">So the paper does something else entirely. </p><div class="pullquote"><p style="text-align: center;">It trains only on trajectories that succeeded.</p></div><p style="text-align: justify;">The insight is almost embarrassing in its simplicity. Nobody annotates their own mistakes, but everyone accumulates work that came out right, automatically, as a byproduct of operating. Successful runs cost nothing to collect and require no labeling at all. You already have them. They are sitting in your archive.</p><p style="text-align: justify;">So the model learns what a well shaped path from question to answer looks like. It treats a trajectory as a continuous path through a latent space rather than as a list of discrete events, and it learns the flow of that path across roughly a hundred successful runs. Then, given a failed run, it scores each step by how far the actual path has drifted from where a successful path would have been at that point. High deviation, high suspicion. The detector has never been shown a single error, and it never needed a definition of one.</p><h2>What it costs to run</h2><p style="text-align: justify;">The accuracy numbers matter less than the operating numbers, so take them quickly. In domain, <strong>the method reached an F1 of 0.435 against 0.181 for GPT-5</strong>, which the authors summarize as a twenty point improvement over the prompting baselines. Tested out of distribution, on a different benchmark built from multi-agent systems running on a different base model, it still came out ahead by about seven points, and this is the more interesting of the two results because it suggests that the learned sense of a well shaped trajectory transfers to settings the model was never trained on.</p><p style="text-align: justify;">Now the part that changes how you would deploy it. The final model is a three layer network. It needs <strong>less than a gigabyte of video memory</strong>. It produces zero output tokens, because there is no language model in the loop at inference time. And <strong>it returns a verdict in about seven milliseconds, against roughly four seconds for GPT-4o and roughly forty seconds for GPT-5</strong>.</p><p style="text-align: justify;">That gap of two to three orders of magnitude is what converts quality control from sampling into census. You would not review a suspicious subset of agent runs. You would review all of them, in real time, on hardware already sitting under a desk, without a single token leaving the building.</p><h2>Four things a practitioner can take from this</h2><ol><li><p style="text-align: justify;"><strong>Your closed files are the asset, not your incident log.</strong> Firms hold almost no structured record of how work went wrong, because nobody documents their own errors at the granularity that would be useful. What firms do hold, in volume, is completed work that came out right. This paper is a demonstration that the second kind of record is enough to build a detector for the first kind of event, which inverts the usual assumption that quality control has to begin with a catalog of failure modes.</p></li><li><p style="text-align: justify;"><strong>The tool points at the origin rather than the symptom.</strong> In one of the case studies, an agent failed to retrieve a fact, invented an answer at step seven, and then built further steps on top of the fabrication. The detector flagged the invented step with a high score and the downstream steps with progressively lower ones. That decay is the useful behavior. An inherited error is less anomalous than the step that introduced it, and a reviewer whose attention is drawn to the origin instead of to the visible symptom is being sent to the right place. Anyone who has watched a mistaken assumption in a term sheet propagate through a dozen dependent documents will recognize why the ordering matters more than the raw detection rate.</p></li><li><p style="text-align: justify;"><strong>The sensitivity is a policy dial, not a technical parameter.</strong> The method offers two ways to decide how many steps to flag. One selects a fixed number of the most anomalous. The other uses conformal prediction, which sets the threshold from a held out sample of successful runs and carries a formal guarantee that the false positive rate on normal steps stays below a level you choose. That level is a supervision policy expressed as a number. How much noise will you tolerate in order to avoid missing something? That is a question for the partner responsible for the file, not for whoever configures the model.</p></li><li><p style="text-align: justify;"><strong>You do not need access to the model that did the work.</strong> The authors tested extracting the step representations with models entirely different from the one that generated the trajectories, and performance degraded only slightly. The practical consequence is that you can monitor an agent running inside a vendor&#8217;s closed system using your own local model as the observer. Oversight does not require privileged access to the thing being overseen, which is worth knowing when the thing being overseen belongs to somebody else.</p></li></ol><h2>Where it fails, which is where it gets interesting</h2><p style="text-align: justify;">The authors publish their failure cases, and those are more instructive than the successes.</p><p style="text-align: justify;">The method reliably catches the step where a failure becomes explicit. It misses the quiet error at the beginning. In one case the agent misjudged, at step one, which tools were available to it, then spent the following steps pursuing a retrieval strategy that could not work, before announcing at step eleven that it was unable to complete the task. The detector scored step eleven very highly and did not flag step one at all. The explanation the authors give is honest: a reasoning trace saying that no suitable tool appears to be available is not anomalous in itself, because successful runs contain statements like that all the time, just before the agent finds another route.</p><p style="text-align: justify;">For legal work, that limitation lands in the worst possible place. The quiet early misjudgment is the expensive one. By the time an error announces itself, the cost of the detour is already sunk. So this is a triage instrument that ranks steps by suspicion, and not an adjudication of cause, and it belongs in the hands of someone who will read the flagged steps rather than act on the flag.</p><p style="text-align: justify;">There is a second limitation the paper does not dwell on, but which any firm should. The model&#8217;s definition of normal is whatever your successful trajectories happen to look like. Train it on your archive and it learns your habits, including the ones that have worked so far without being good. A genuinely better approach, unfamiliar to the training data, reads as deviation. Anomaly detection built on precedent will always be conservative about precedent, and a profession that already treats precedent as evidence should be alert to a tool that treats it as ground truth.</p><p style="text-align: justify;">The authors close their own discussion of impact with a warning worth repeating: attribution tools of this kind could be used to make agents better at evading detection rather than better at the work, and automated attribution should support human oversight rather than replace it. That sentence was written by researchers about their own contribution, which is more than most vendors manage about theirs.</p><h2>Back to three in the afternoon</h2><p>The agent has taken forty steps and produced the wrong answer. You still have to find the break.</p><p>What has changed is where you would go looking. Not to a bigger model with a longer prompt, which this paper suggests would do worse than guessing. Not to a catalog of known failure modes, which nobody has and nobody is going to build. You would go to the hundred matters that came out right, ask what shape they had, and then look for the place where this one stopped having it.</p><div><hr></div><p>Paper: <em>Tracing Agentic Failure from the Flow of Success</em>, Yeh, Zhu, Deep and Li (University of Wisconsin-Madison and Microsoft Research)</p><p>https://arxiv.org/abs/2607.12747v1</p>]]></content:encoded></item><item><title><![CDATA[Every Legal Benchmark Has a Lawyer Hidden Inside It]]></title><description><![CDATA[The hardest part of lawyering happens before the legal question exists, and no benchmark measures it.]]></description><link>https://thelegaltechoverlap.substack.com/p/every-legal-benchmark-has-a-lawyer</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/every-legal-benchmark-has-a-lawyer</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Mon, 27 Jul 2026 06:30:27 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">A man loses a large sum from his current account over the course of nine days. He complains to the bank, receives a letter telling him the transactions were properly authenticated, and decides to bring the matter before an out-of-court board that costs him almost nothing and does not require him to instruct a lawyer.</p><p style="text-align: justify;">He has a chatbot open beside the form, because he has been using one all year for everything else, and he tells it his story the way he would tell it to his daughter. He starts several years back, when he opened the account. He mentions that the branch manager was very kind. Around the eleventh sentence he notes, in passing, that a man who said he worked in the fraud office called him and that he read out a code from a text message. Then he goes back to the branch manager.</p><p style="text-align: justify;">That one clause decides the case. The model has no way of knowing this, because nothing in the way the story was told marked it as the hinge. It writes him a fluent submission about the bank&#8217;s duty of care.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3456" height="5184" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:5184,&quot;width&quot;:3456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;silver and brown round pendant necklace on brown wooden table&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="silver and brown round pendant necklace on brown wooden table" title="silver and brown round pendant necklace on brown wooden table" srcset="https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1600849128558-10cdf7a3fb8e?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxMjl8fGxhd3llcnxlbnwwfHx8fDE3ODUwNDcwMzZ8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@phil13689">Philippe Tinembart</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2>Two different skills wearing one name</h2><p style="text-align: justify;">A paper accepted to the AI4Law workshop at ICML 2026, by Andrew Lou and David Shin of Yale Law School, makes an argument that is simple to state and hard to put back down. Legal AI research keeps justifying itself by appealing to access to justice, while legal AI benchmarks keep measuring something else.</p><div class="pullquote"><p style="text-align: center;">The distinction they draw is between legal reasoning and lawyering. </p></div><p style="text-align: justify;">Legal reasoning is the application of doctrine to a question that has already been put into legal form. Lawyering is the work of putting it into that form, which means extracting the operative facts from a rambling account, discarding the kind branch manager, reordering what remains by legal relevance rather than by chronology, and noticing what the client has failed to say so that you can ask him about it.</p><p style="text-align: justify;">Every benchmark the authors survey evaluates the first skill on inputs that have already had the second performed on them. By the time a benchmark prompt reaches a model, a substantial amount of expert cognitive labor has been done for free.</p><p style="text-align: justify;">The point sharpens when you consider what one of those benchmarks is made of. LEXam is built from real law school examinations, and a well-written exam hypothetical is the most heavily processed legal input in existence. A professor constructed it so that every fact needed for the answer is present, every distractor is there deliberately, and the procedural posture is stated rather than inferred. It is a laboratory specimen of a legal question. Scoring well on it demonstrates something real, and it demonstrates it under conditions that no client has ever produced.</p><h2>Why the two numbers can move apart</h2><p style="text-align: justify;">Lou and Shin describe this as a gap between an upper and a lower bound. The upper bound is model performance when a competent professional does the framing. The lower bound is performance when the framing is done by the person who has the problem. Benchmarks report the first and stay silent about the second.</p><p style="text-align: justify;">A silent variable is not necessarily a stable one, and this is the part of the argument worth sitting with. Their claim is that the two bounds can move in different directions, because the training that lifts benchmark scores is the same training that erodes a model&#8217;s willingness to admit it is missing something. The abstention research they rely on finds that reasoning fine-tuning makes models measurably worse at declining to answer, including in the domains the reasoning training was aimed at. A system that is better at producing an answer is not automatically better at recognizing that the question it received cannot be answered from what it was given.</p><p style="text-align: justify;">So the upper bound rises and gets announced, the lower bound moves quietly or not at all, and the distance between them stays invisible because nobody instruments it. Meanwhile the population sitting at the lower bound is the least equipped to notice when the output is wrong, since fluent legal prose that agrees with your own theory of the case is indistinguishable from good advice if you have never seen either before.</p><h2>What they actually did</h2><p style="text-align: justify;">The empirical section is deliberately modest, and the authors say so themselves. They call it a primer, run it on a small sample of multiple choice questions, and decline to draw substantive conclusions from it.</p><p style="text-align: justify;">They degraded the inputs along two axes taken from the literature on unrepresented litigants. First came typos, inserted at rising density, in three flavors: single character deletions, transpositions, and slips to an adjacent key. Then came dilution, first by wrapping the question inside blocks of irrelevant filler prose, and then by breaking the question apart and threading filler between its own sentences, with the formatting flattened so that no layout cue could rescue the model.</p><p style="text-align: justify;">Accuracy fell, which was expected. Two secondary observations are more interesting than the headline:</p><ol><li><p style="text-align: justify;">The first is that the models did not degrade uniformly, and their ranking changed between the clean and the distorted runs even though all of them came from the same provider. If that holds at scale, then the model that leads on a legal benchmark is not necessarily the model you would want in the hands of someone filing alone, and no leaderboard in circulation today would tell you which is which.</p></li><li><p style="text-align: justify;">The second is a small human detail. One model, instructed to return nothing but a letter, began several of its answers by remarking that it had probably encountered a typo. It noticed the noise. It simply had no framework for treating that noise as information about the person who produced it.</p></li></ol><h2>The objection from the civil law side, and why it does not hold</h2><p style="text-align: justify;">The natural European reaction is that this is an American problem. Litigation without counsel on the American scale does not exist here, and standing personally before a judge is confined to the smallest claims.</p><p style="text-align: justify;">That reaction is right about courtrooms and wrong about everything else, because Europe did not address the representation problem by opening the courtroom. It addressed it by building an adjudicative layer outside the courtroom, where technical defense is optional and frequently absent, running from banking and financial arbitration through consumer conciliation to simplified small claims procedures.</p><p style="text-align: justify;">Those forums are where the argument bites hardest, and the reason is structural rather than statistical. An American plaintiff without a lawyer files into a system that has slack in it, since there are hearings, a judge who can ask a question from the bench, an opponent with every incentive to point at the gap, and an appeal. A mistake has several chances to surface before it becomes final.</p><p style="text-align: justify;">A written, documentary proceeding has none of that slack. The record closes, the panel reads what was submitted, and it decides. The contraddittorio happens on paper. Nobody in the process ever asks the question that would have surfaced the missing fact, because the design contains no moment at which such a question gets asked.</p><p style="text-align: justify;">The continental version of this problem is therefore the concentrated one.</p><h2>The bench test</h2><p style="text-align: justify;">Three things follow for anyone building or practicing.</p><ol><li><p style="text-align: justify;">If you build client-facing intake, measure your own lower bound before somebody else does it for you. Take matters you won, reconstruct how the client first described the problem on the initial call, and run your system against that version rather than against the tidy summary that ended up in the file. The distance between that result and your demo is the honest description of your product.</p></li><li><p style="text-align: justify;">If you take over a matter a client has already started on his own, audit for absence before you audit for error. Assume the submission was shaped by a system that never asked a clarifying question, and that the decisive fact is missing because nobody prompted for it. Reading what is on the page is the easy half of that job.</p></li><li><p style="text-align: justify;">And if you are choosing models for this kind of work, stop treating benchmark rank as transferable. The question that matters is how gracefully a model degrades when the question arrives dirty, and that is a measurement you will have to run yourself, because nobody is publishing it.</p></li></ol><p style="text-align: justify;">The man in the vignette gets his decision on the papers in a few months. Whether it goes his way turns less on what the panel thinks about his rights than on whether sentence eleven reached the file in a form that somebody could see.</p><div><hr></div><p><strong>Source:</strong> Andrew Lou and David Shin, <em>Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice</em>, arXiv:2606.23716, accepted to the AI4Law Workshop at ICML 2026.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/every-legal-benchmark-has-a-lawyer?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading The LegalTech Overlap! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/every-legal-benchmark-has-a-lawyer?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thelegaltechoverlap.substack.com/p/every-legal-benchmark-has-a-lawyer?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p></p>]]></content:encoded></item><item><title><![CDATA[Three AIs Walked Into a Courtroom, and the Smartest One Refused to Debate]]></title><description><![CDATA[The most rigorous study yet on multi-agent legal AI found that stacking models can quietly make your answers worse, and the reason should change how you build.]]></description><link>https://thelegaltechoverlap.substack.com/p/three-ais-walked-into-a-courtroom</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/three-ais-walked-into-a-courtroom</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Fri, 24 Jul 2026 06:30:38 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Three AI agents walked into a courtroom: one played the judge, one argued that the statute applied, one argued that it didn&#8217;t, and on a question about who bears the management costs on a pledged asset they reached the correct answer on the very first exchange. Then they kept talking. By the second round one of them raised an objection that sounded sophisticated and was flatly wrong, the other two agreed with it, and the group handed back exactly the mistaken verdict it had just avoided. That single transcript sits inside a new paper on legal reasoning, and it points at something the industry is quietly betting on without much evidence.</p><p style="text-align: justify;">Here&#8217;s the question I couldn&#8217;t shake while reading it: when you put a whole room of models to work on a legal problem, are you buying better answers, or are you just buying louder ones? I went in assuming the committee would win, because it mirrors how chambers actually work, and I came out having to revise almost everything about when I&#8217;d reach for one.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="8256" height="5504" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:5504,&quot;width&quot;:8256,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a room that has some chairs and a table in it&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a room that has some chairs and a table in it" title="a room that has some chairs and a table in it" srcset="https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1679384412940-0b82338760d6?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzOXx8Y291cnRyb29tfGVufDB8fHx8MTc4NDA2ODY3N3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@pafuxu">Kouji Tsuru</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h3>The setup, in one breath</h3><p style="text-align: justify;">The study is called <em>L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning</em>, out of JAIST and VNU, and its test bed is deliberately clean. The task is COLIEE Task 4, which hands a model an article of the Japanese Civil Code and a legal hypothesis and asks whether the first entails the second. It&#8217;s binary, it&#8217;s balanced, and a coin flip lands you near 50%. The authors give three agents courtroom roles, so you get an impartial judge, an advocate arguing for entailment, and an advocate arguing against, and then they compare two ways of ending the argument. Under consensus the agents negotiate until they converge, and under voting they reason in parallel and cast independent ballots at the end. They run the whole thing across four open models of very different sizes, and that size gap turns out to be the entire plot.</p><h3>The twist in the title</h3><p style="text-align: justify;">So which agent was the smartest, and why did it refuse to debate? On the strongest model in the study, Qwen3-32B, the committee stopped adding anything at all. A single instance of that model, sampling a few times and taking the majority answer, landed at <strong>89.98</strong> and quietly topped the entire table. The two debate setups came in behind it at 88.56 and 85.20, and here&#8217;s the part that should stop you cold: even plain, unadorned zero-shot prompting on that model beat both flavors of debate. Once the underlying model is genuinely capable, the personas and the rounds and the negotiation buy you nothing, and they quietly cost you a point or two plus a pile of compute. The smartest participant in the room did best by not arguing.</p><p style="text-align: justify;">Now, the committee does earn its keep somewhere, and it&#8217;s worth being precise about where. On the mid-sized 30B model the framework shines, climbing from the low-80s on its single-agent baselines to an average of 88.46 under consensus, and on the 2025 subset one cell reaches 95.12, which is the kind of number that makes a demo sing. The authors report gains of up to eight points in that sweet spot, and I believe them. But a technique that only helps in the middle of the capability range is a very different product from the universal upgrade it&#8217;s often sold as.</p><h3>The part that should worry you</h3><p style="text-align: justify;">Push down to the smallest model, Llama3.1-8B, and neutral turns into harmful. Both debate setups drag that model below its own single-agent baselines, from around 73 down to 64.53 under consensus. </p><div class="pullquote"><p style="text-align: justify;">The authors have a clinical name for what&#8217;s happening, collaborative hallucination, and the transcripts show the mechanism step by step. </p></div><p style="text-align: justify;">One weak agent floats a shaky reading of a statute, the others don&#8217;t have the horsepower to catch the error, so they ratify it and treat it as settled law. The committee doesn&#8217;t fix the mistake, it laminates it. Sit with that for a second, because it inverts the cost logic everyone reaches for first. The cheap small models you&#8217;d most want to gang up, precisely because they&#8217;re cheap, are the ones that form a confident <strong>echo chamber</strong> and talk each other into being wrong.</p><p style="text-align: justify;">There&#8217;s a governance line hiding in the same data that I keep thinking about. Expressions of hesitation and uncertainty showed up roughly ten times more often in the failure logs of the weaker models than the capable ones. So the systems most likely to be wrong are also the ones that sound the least sure of themselves, and a naive reviewer skimming the output would read that hedging as diligence rather than as a warning light.</p><h3>Two knobs, set them on purpose</h3><p style="text-align: justify;">The paper&#8217;s most usable lesson is that there&#8217;s no single correct way to close a debate, and the right setting tracks model strength cleanly enough to act on. On the larger model, forcing consensus works better, because a strong agent can genuinely spot a flaw in a peer&#8217;s argument and correct it, so the negotiation behaves like real peer review. On the smaller Qwen, independent voting wins instead, 89.37 against 87.23, because keeping the reasoning paths separate stops one confident-but-wrong agent from dragging everyone along before the ballots are counted. The mechanism that helps a strong model is the very mechanism that hurts a weak one, so this is a dial you want to set deliberately.</p><p style="text-align: justify;">And then there&#8217;s the round count, which behaves in the least intuitive way of all. If a little deliberation helps, surely more helps more, and surely you should let the agents really think it through. The data says the opposite: </p><div class="pullquote"><p style="text-align: justify;">adding agents nudges accuracy up, roughly the variance reduction any ensemble gives you, but adding rounds sends accuracy down. </p></div><p style="text-align: justify;">That&#8217;s the over-deliberation drift from the opening scene, where a correct first round gets talked into a wrong second one. Anyone who has watched a good meeting curdle after the fortieth minute will recognize the shape of it.</p><h3>The signal is the real product</h3><p style="text-align: justify;">Here&#8217;s the finding I&#8217;d put to work tomorrow: when the agents vote and split, that split is a startlingly honest measure of difficulty. On the mid-tier model, questions with a unanimous vote came in around 70% accurate, while the split votes, which were about a quarter of all queries, dropped to roughly 48%, which is a coin flip wearing a suit. So the disagreement works less like noise you&#8217;d want to filter out and more like a free, training-free flag that says <em>send this one to a human</em>. For anyone building a review workflow, that&#8217;s the gold, because it lets you automate the confident cases and reserve expensive attention for the genuinely hard ones, and you get the flag as a byproduct with no extra calibration.</p><p style="text-align: justify;">But follow that thought one step further, and this is where I want the room to push back. If the committee&#8217;s most valuable output is a confidence signal, and a single strong model already matches the committee on accuracy, then maybe the whole debate apparatus is an expensive way to buy something cheaper. You could sample one good model a handful of times, measure how often it agrees with itself, and use that agreement rate as your triage flag, skipping the personas and the rounds entirely. The paper stops short of saying this out loud, though its own related-work section hints that voting may explain much of the improvement usually credited to debate. So the question I&#8217;d genuinely like challenged is this: outside a tidy benchmark, does multi-agent debate earn its compute, or is it an elaborate reinvention of &#8220;sample a few times and count the votes&#8221;?</p><h3>The caveat a banking litigator can&#8217;t ignore</h3><p style="text-align: justify;">I&#8217;ll end on the limit that keeps me honest, because it&#8217;s a big one. This is Japanese Civil Code entailment, which is about as clean as legal reasoning ever gets, since the governing article is handed to the model, the facts are stipulated, and the answer is a crisp yes or no. My working day looks nothing like that. In banking litigation the applicable rule is contested, the facts are scattered across thousands of pages of a securitization file, the decisive provision might be a supervening regulation nobody flagged, and the &#8220;right answer&#8221; is whatever a panel decides two years later. The authors are candid that a stubborn 17 to 24% of their cases fail no matter the configuration, because the knowledge simply isn&#8217;t in the supplied text, and that fraction would only swell in the messy open-world matters where we actually earn our fees.</p><p style="text-align: justify;">So I read L-MAD less as a recipe and more as a warning with a gift attached. The warning is that stacking models and rounds is not a free lunch, and a system that looks more sophisticated while quietly getting worse is the most expensive kind of mistake in our line of work. The gift is that disagreement, treated as a routing signal rather than a problem to be argued away, gives you a principled seam between what you automate and what you escalate. That seam is the real product here, and I suspect it will outlast any leaderboard argument about which debate topology wins.</p><p style="text-align: justify;">Curious where you all land. If you&#8217;ve run multi-agent setups in production, are you seeing accuracy gains that survive contact with a strong base model, or are you mostly harvesting the confidence signal and paying committee prices for it? I&#8217;d rather be wrong in the comments than right in private.</p><p style="text-align: justify;"><em>Source: Nguyen et al., "L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning" (preprint, 2026).</em><br><em>Read the full paper: <a href="https://arxiv.org/abs/2607.09099">https://arxiv.org/abs/2607.09099</a></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/three-ais-walked-into-a-courtroom/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thelegaltechoverlap.substack.com/p/three-ais-walked-into-a-courtroom/comments"><span>Leave a comment</span></a></p><p style="text-align: justify;"></p>]]></content:encoded></item><item><title><![CDATA[Your Legal AI Probably Already Wrote the Right Answer]]></title><description><![CDATA[A new Stanford framework moves the hard problem in legal AI from writing answers to choosing among them.]]></description><link>https://thelegaltechoverlap.substack.com/p/your-legal-ai-probably-already-wrote</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/your-legal-ai-probably-already-wrote</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Wed, 22 Jul 2026 13:02:45 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">There is a quiet, uncomfortable finding buried in most benchmark tables, and a paper out of Stanford, Berkeley, and NVIDIA just made it impossible to ignore. When you let a strong model attempt a hard task five times, one of those five attempts is usually correct. On the coding benchmark the authors use, an oracle that always picks the best of the pool solves nearly the entire leaderboard, <strong>reaching 98.9 percent</strong>, while any single attempt lands far lower. The capability is already sitting inside the model. What is missing is the judgment to recognize the good answer when it appears.</p><p style="text-align: justify;">That gap has a name in the paper, and it is one that everyone shipping legal AI should sit with. The authors call verification a scaling axis, on par with pretraining and test-time compute, and they argue we have spent almost all our effort on generation and almost none on the thing that decides whether generation is worth anything. If you run a legal research agent, a redline agent, or a memo drafter that samples a few candidate outputs and returns one, your product quality is capped not by how well the model writes but by how well something downstream selects. Most teams have never measured that selector. This paper is a good reason to start.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="4288" height="2848" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2848,&quot;width&quot;:4288,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;people walking on spiral staircase&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="people walking on spiral staircase" title="people walking on spiral staircase" srcset="https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1441038718687-699f189fa401?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGF3eWVyfGVufDB8fHx8MTc4NDYzMjYwN3ww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@showkin9">Yoosun Won</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2>What the paper actually does</h2><p style="text-align: justify;">The mechanism is almost embarrassingly simple, which is usually a sign it will spread. A standard LLM judge is asked to score a candidate on a scale, and it emits a single token, say a 4 out of 5. That token throws away everything the model was actually thinking. Underneath, the model held a probability distribution across all the score tokens, a little cloud of belief that this answer is probably a 4, maybe a 5, with a whisper of 3. The judge collapses that cloud into one integer and moves on.</p><p style="text-align: justify;">LLM-as-a-Verifier keeps the cloud. Instead of reading the emitted digit, it reads the model&#8217;s log-probabilities over every score token and computes the expected value, turning a coarse integer into a smooth, continuous score. Nothing new is trained. The verifier is a general model, in most experiments Gemini 2.5 Flash, prompted to rate a trajectory and then mined for its logits rather than its words.</p><p style="text-align: justify;">From that one move, three dials appear. The first is granularity: widen the score range from five levels to twenty and the continuous estimate gets finer, lifting pairwise verification accuracy on the coding benchmark from 73.1 to 77.5 percent. The second is repetition: average several independent passes and the noise falls, so much so that a single-pass continuous verifier already matches a discrete judge that was ensembled sixteen times. The third is decomposition: replace one vague question (&#8221;is this trajectory correct?&#8221;) with a handful of narrower ones, checking specification, output format, and error signals separately, then average, which pushes accuracy to 78.3 percent.</p><p style="text-align: justify;">The result that will get quoted is the tie rate. A discrete judge scoring hard coding trajectories calls it a tie 26.7 percent of the time, shrugging on more than a quarter of the comparisons that matter most. The continuous verifier produces zero ties. The clearest illustration is a single SQL optimization task the authors dissect. Two candidate solutions look almost identical, but one validates its optimized query against a tampered, indexed copy of the database rather than the real one, which is a subtle methodological cheat. A discrete one-to-five judge ties the two candidates in 88 of 100 runs. Taking the expectation over the same scale breaks every tie and ranks the correct one higher 69 times. Widening the scale to twenty pushes that to 77.</p><p style="text-align: justify;">Stacked together and wrapped in a budget-aware ranking tournament, the framework sets new marks across coding, robotics, and medical agent tasks, and it does so with no fine-tuning at all. There is a second, quieter contribution that legal teams should care about more than the headline scores. The same verifier score, tracked step by step through a task, rises steadily on runs that are going well and stays flat on runs that are drifting toward failure. That gives you a live progress meter and an early-warning light for a long-running agent, which the authors ship as an extension for Claude Code and Codex.</p><h2>Where it is genuinely strong</h2><p style="text-align: justify;">The training-free part is not a footnote, it is the whole commercial argument. Nobody has to curate a labeled dataset, stand up a reward-model training run, or babysit a fine-tune that goes stale the moment the base model updates. You prompt a capable model, read its logits, and you have a verifier that transfers across domains it never saw. For a small legal team without a machine-learning function, that is the difference between a research idea and something you could actually wire into a pipeline this quarter.</p><p style="text-align: justify;">The calibration is real, and calibration is what legal buyers are quietly starving for. A continuous score that reliably separates a stronger answer from a weaker one, and that can be turned into an honest preference probability, is far more useful than a judge that keeps declaring five-way ties. And the progress signal opens a door that pure scoring does not, because a monitor that can flag a drifting agent mid-task is exactly the kind of control a supervising lawyer needs before an agent commits a bad edit to a document.</p><h2>Where it strains, and why lawyers should notice</h2><p style="text-align: justify;">Now the part the abstract does not advertise. Every benchmark in this paper has a checkable ground truth. The hidden test suite either passes or it does not. The robot either grasps the object or it does not. The patient record either gets retrieved or it does not. The verifier is impressive precisely because there is a fact of the matter it is being measured against, even when that fact is expensive to compute. That single assumption is where the method meets the wall that runs through most of legal work.</p><p style="text-align: justify;">There is no hidden test suite for &#8220;this is the stronger argument on abstention from a restructuring vote,&#8221; and there is no unit test that turns green when a brief is persuasive. A great deal of legal judgment is genuinely contested, and on genuinely contested questions a confident continuous score is not a gift, it is a hazard, because it dresses a matter of judgment in the costume of a measurement. A discrete judge that hedges with a tie is at least being honest about its uncertainty. A verifier that returns a crisp 14.7 over a question with no ground truth invites a lawyer to trust a number that means nothing.</p><p style="text-align: justify;">Three more constraints matter before anyone gets excited. The method needs access to the verifier&#8217;s token log-probabilities, which several closed frontier APIs, including some of the models legal teams actually deploy, do not expose. The authors offer a two-stage workaround that routes a closed model&#8217;s reasoning through an open, logit-accessible verifier and recovers most of the gain, but that adds plumbing and a second model to your bill. The criteria decomposition, which does a lot of the heavy lifting, is hand-designed per domain rather than learned, so someone has to sit down and write the legal rubric. And the whole approach assumes you can cheaply generate several candidates per task, which multiplies your generation cost before the verifier ever runs.</p><h2>The legal desk test</h2><p style="text-align: justify;">Strip away the benchmark glamour and ask the only question that matters for a working practice: would this survive contact with a real legal desk? Run any candidate use case through four gates.</p><ol><li><p style="text-align: justify;"><strong>Is there a checkable notion of correct?</strong> If correctness is contested judgment, the verifier gives you false confidence. If it is a fact, you are in business.</p></li><li><p style="text-align: justify;"><strong>Can you generate several candidates cheaply?</strong> Selection needs a pool. One draft, nothing to verify.</p></li><li><p style="text-align: justify;"><strong>Do you have logit access, or tolerance for a two-model workaround?</strong> No logits, no continuous score.</p></li><li><p style="text-align: justify;"><strong>Does the task split into verifiable sub-criteria?</strong> Decomposition is where most of the accuracy lives.</p></li></ol><p style="text-align: justify;">Score a few real tasks against those gates. Verifying that every authority cited in a memo actually exists and says what the memo claims passes cleanly: correctness is checkable, sub-criteria are obvious, and a selector that ranks the citation-clean draft above the citation-sloppy one is immediately valuable. Choosing the best clause extraction or the safest of several redlines mostly passes, because format and rule compliance are concrete even when style is not. Deciding which of five research memos is strongest fails the first gate outright, because there is no ground truth to calibrate against. Ranking litigation strategies fails for the same reason, and fails hardest, because the confident number is most seductive exactly where it is least earned.</p><p style="text-align: justify;">The honest verdict is that LLM-as-a-Verifier is a strong selector and monitor for the checkable slices of legal work, and a category error for the contested ones. The teams that win with it will be the ones disciplined enough to point it only at the parts of the practice where &#8220;correct&#8221; is a fact rather than an argument. That discipline, and not the logits, is the hard part.</p><h4>Sources</h4><ul><li><p style="text-align: justify;">Kwok et al., <em>LLM-as-a-Verifier: A General-Purpose Verification Framework</em>: <a href="https://arxiv.org/abs/2607.05391">https://arxiv.org/abs/2607.05391</a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/your-legal-ai-probably-already-wrote/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thelegaltechoverlap.substack.com/p/your-legal-ai-probably-already-wrote/comments"><span>Leave a comment</span></a></p><p></p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Harness Is the Part You Can Actually Build]]></title><description><![CDATA[A builder&#8217;s read of the new &#8220;agent harness&#8221; research, and the accountability trap hiding inside it.]]></description><link>https://thelegaltechoverlap.substack.com/p/the-harness-is-the-part-you-can-actually</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/the-harness-is-the-part-you-can-actually</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Mon, 20 Jul 2026 13:11:23 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">I&#8217;ve just written on LinkedIn about a survey by the University of Illinois, Meta, and Stanford arguing that the reliability of an AI agent depends less on the model and more on the &#8220;harness&#8221; around it. That post made the case to buyers. This one is for the people on the other side of the table: the lawyers who are quietly building their own tools, wiring an LLM into a retrieval system, a handful of prompts, and a review step, then hoping the result holds up.</p><p style="text-align: justify;">The encouraging part of the paper is that the harness is the piece you actually control. You are not going to out-train a frontier lab, but you can out-engineer a competitor on permissions, verification, and record-keeping, which happens to be where legal work lives. Here are the components worth taking, and the one that will bite you if you copy it without thinking.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="2551" height="3401" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3401,&quot;width&quot;:2551,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Statue of justice holding scales and sword.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Statue of justice holding scales and sword." title="Statue of justice holding scales and sword." srcset="https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1760089449852-e8cade7feb53?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw4OXx8bGVnYWx8ZW58MHx8fHwxNzg0NTIwOTQwfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@irynalimborska">Iryna Limborska</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2>Ship an evidence bundle with every output</h2><p style="text-align: justify;">The most useful idea in the paper is almost thrown away: the authors argue that verification should stop being a single pass/fail flag and become a stack, where each check states plainly what it verifies, what it cannot verify, and how confident it is. Then they take the step most legal tools skip: every accepted action should carry an &#8220;evidence bundle&#8221; recording the checks that ran, the assumptions it preserved, the parts it left untested, and the risks that remain.</p><p style="text-align: justify;">Read that as a legal deliverable and it is close to a work-product cover memo. Picture your citation checker returning not &#8220;passed&#8221; but a short bundle: citations resolved against the official source, quotes matched to the paragraph, holdings not independently verified, jurisdiction assumed to be the one in the caption, three passages flagged as low confidence. That bundle is the line between an output you can defend and an output you still have to re-check by hand. Build the bundle first and the model can improve later.</p><h2>Retrieve by legal structure, not by token window</h2><p style="text-align: justify;">A recurring finding in the retrieval sections is that code tools improved not by pulling in more text but by pulling in text aligned to the structure of the code, using structure-aware chunking and query rewriting rather than fixed-size windows.</p><p style="text-align: justify;">Most legal RAG systems get this backwards. They slice a contract or a judgment into 800-token blocks and then wonder why the model cites half a clause. Chunk by the units that carry legal meaning instead: the article, the clause, the recital, the holding, the operative paragraph. Let the retriever rewrite the query as it narrows in, the way a junior does when the first search misses. The paper&#8217;s memory work adds a second warning that applies directly here: a curated, governed store of past matters beats a large one, because ungoverned history mostly contributes noise and confident wrong retrievals, so more precedent in the index is not more signal.</p><h2>Give the tool a permission ladder, and a plan it has to sign</h2><p style="text-align: justify;">The paper&#8217;s control model is a clean three-rung ladder. The bottom rung reads and inspects. The middle rung edits and runs things inside a sandbox. The top rung touches the outside world: sending, filing, paying, deleting. Only the top rung requires a human gate, and the important nuance is that the risk of an action depends not on the tool but on the arguments, the data, and the side effects. The same &#8220;send&#8221; is trivial in a draft and irreversible in a court filing.</p><p style="text-align: justify;">Before any of that, the paper treats the plan itself as a contract. A good plan names the files it will touch, the invariants it must preserve, the checks that will confirm success, and the rollback path if it fails. For a legal tool, that is a scope of work the agent commits to before it acts, and a natural place to require sign-off. If your tool cannot state, in advance, what it is about to change and how to undo it, it is not ready to touch anything that leaves the building.</p><h2>Stop pasting whole documents into the prompt</h2><p style="text-align: justify;">One quiet technique runs through the whole paper: </p><div class="pullquote"><p style="text-align: justify;">keep the full material outside the model, and pass the model a compact, provenance-preserving summary plus a handle back to the original. A failing report becomes a few key lines and a link, not the entire log.</p></div><p style="text-align: justify;">Legal builders do the opposite by reflex: they paste the whole contract, the full brief, the complete deposition into the context and trust the model to find the needle. The research is blunt about why that fails: attention degrades across long inputs, and the middle gets lost. Offload the documents, retrieve the relevant spans, and keep a pointer to the source for anything the model leans on. Your outputs get more accurate, and, as a side effect, more auditable.</p><h2>Where this breaks for law</h2><p style="text-align: justify;">Now the critique, and I make it only because the paper half-makes it itself.</p><p style="text-align: justify;">The entire paradigm draws its confidence from execution. A code harness can compile the program, run the tests, and read the crash. The authors are careful to scope &#8220;code&#8221; to things that are machine-checkable, and they say plainly that human intent and judgment are not code. They even warn that execution feedback can create a false sense of correctness, and that a system can grow overconfident precisely because it has a green test to point at.</p><p style="text-align: justify;">Hold that next to legal work and the problem is obviou: </p><div class="callout-block" data-callout="true"><p style="text-align: justify;">Law has almost no execution oracle. </p></div><p style="text-align: justify;">There is no unit test for whether an argument is persuasive, whether a clause allocates risk the way the client actually intended, or whether a filing meets a standard a judge will apply next year. The harness gives you strong guarantees exactly where law is weakest, and near-silence where law actually decides things. Take the scaffolding, but do not let &#8220;verifiable&#8221; quietly become &#8220;correct,&#8221; in your own head or in your marketing.</p><p style="text-align: justify;">There is a sharper version of this, and it comes from setting two of the paper&#8217;s own recommendations next to each other. The paper says to compact evidence into short summaries before showing it to a human. It also says that high-stakes actions should pause for human approval, logged as an auditable record of who signed off and on what basis. Put those together and you get a system that hands a busy partner a tidy three-line summary, records the click as informed authorization, and produces an audit trail that looks flawless. The trail is real and the judgment behind it may be a rubber stamp. In our world, that is liability laundering with good UX.</p><p style="text-align: justify;">So if you build the approval gate, resist the urge to over-compress what the approver sees on the actions that carry real weight. An audit trail of uninformed sign-offs is worse than no trail at all, because it manufactures the appearance of the diligence you skipped.</p><h2>The short version</h2><p style="text-align: justify;">If you are building, the checklist is small and unglamorous. </p><ol><li><p style="text-align: justify;">Attach an evidence bundle to every output. </p></li><li><p style="text-align: justify;">Chunk and retrieve by legal structure, and keep your experience store curated rather than large. </p></li><li><p style="text-align: justify;">Put the tool on a permission ladder, and make it sign a plan before it acts. </p></li><li><p style="text-align: justify;">Keep documents out of the prompt and pass handles instead. </p></li><li><p style="text-align: justify;">And on the actions that matter, make human approval slow, informed, and honestly recorded, not fast, compacted, and defensible-looking.</p></li></ol><p style="text-align: justify;">The model is the part everyone is watching but the harness is the part that will decide whether your tool survives its first bad day.</p><div><hr></div><p>Source: <em>&#8220;Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems,&#8221; </em>Ning, Tieu, Fu, et al. (University of Illinois Urbana-Champaign, Meta, Stanford, 2026). https://arxiv.org/abs/2605.18747</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/the-harness-is-the-part-you-can-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thelegaltechoverlap.substack.com/p/the-harness-is-the-part-you-can-actually?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[A Model That Sorts Its Own Doubt]]></title><description><![CDATA[NVIDIA made an AI 2.42 times faster by freezing the half that understands, and for legal work the speed is not even the interesting part.]]></description><link>https://thelegaltechoverlap.substack.com/p/a-model-that-sorts-its-own-doubt</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/a-model-that-sorts-its-own-doubt</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Sun, 19 Jul 2026 05:01:41 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">To make a language model write faster, NVIDIA&#8217;s researchers took half of it and froze it solid. That half can no longer learn anything, and that is exactly why it works. The frozen half is what lets the other half run more than twice as fast at almost no cost to quality. I want to walk through why, because the same move is the one I keep reaching for when I try to make legal AI faster without making it less safe to rely on.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="5472" height="3648" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3648,&quot;width&quot;:5472,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;photo of library hall&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="photo of library hall" title="photo of library hall" srcset="https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1465929639680-64ee080eb3ed?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2Mnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMjMzfDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@willvanw">Will van Wingerden</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h3>So They Built It Twice</h3><p style="text-align: justify;">Start with the problem the paper solves. Most models write the way a court stenographer takes dictation, one token after another in strict order, which is accurate but slow, because every word has to wait for the one before it. Diffusion language models try to escape that by generating text in parallel and then refining it. The trouble is that in the usual design a single network has to do two jobs at once. It has to hold a reliable picture of the text so far, and it has to clean up the noisy tokens it is producing. Those two jobs pull the same weights in opposite directions, so the model does neither as well as it could.</p><p style="text-align: justify;">TwoTower, the model in the paper, refuses to make one network do both. It keeps two copies of a pretrained model and gives each a single job. The first copy, the context tower, reads the clean text causally, the way a normal model does, and it stays frozen, which means its weights never change during training. The second copy, the denoiser tower, is trained to take a block of masked tokens and resolve them by looking across at the frozen tower&#8217;s understanding, layer by layer. One tower understands, and the other writes. The understanding stays fixed and trustworthy, while the writing gets faster.</p><h3 style="text-align: justify;">The Number That Should Not Work</h3><p style="text-align: justify;">The numbers are what earned my attention, because they resist the trade-off I expect. </p><div class="pullquote"><p style="text-align: justify;">Built on a 30B open-weight hybrid model and trained on a fraction of the tokens used to pretrain the backbone, TwoTower keeps 98.7 percent of the original model's benchmark quality while producing text 2.42 times faster in wall-clock terms. </p></div><p style="text-align: justify;">You can push it past three times faster by loosening one threshold, though quality then starts to slip, and the authors say so plainly.</p><h3 style="text-align: justify;">The Part I Did Not Expect</h3><p style="text-align: justify;">This is where the speed becomes the least interesting thing in the paper. The model does not hand over a finished block all at once. At each step it predicts every masked position, commits only the ones it is confident about, and leaves the uncertain positions blank for another pass. The early steps commit many tokens, and the later steps grind through the hard remainder. The model shows you its own uncertainty before you ask.</p><p style="text-align: justify;">Read that again with your due diligence file open. A system that separates what it is sure of from what it is still guessing is doing, at the level of individual words, the same triage a junior associate does with a highlighter. That is a signal a verification layer can act on. A governance system no longer has to treat the model as a black box that slides a finished paragraph across the table. It can step in at the block boundary, where the model has already marked its own weak spots, and then decide what to check, what to escalate, and what to let through. Speed and control stop being enemies once the fast part of the system is built to admit what it does not know.</p><h3 style="text-align: justify;">Where the Speed Quietly Does Nothing</h3><p style="text-align: justify;">Now the honest question, the one worth sitting with. </p><blockquote><p style="text-align: justify;"><em>Where does this speed actually pay off in legal work, and where does it quietly do nothing?</em></p></blockquote><p style="text-align: justify;">It pays off in generation, and only in generation. The paper measures throughput as the time to produce the final answer, not the time to find or read the input, and that line is easy to skate past and expensive to ignore. If you are working a knowledge graph built from millions of documents, your bottleneck is retrieval, which is finding the right nodes and the paths between them. TwoTower does nothing for that first mile. What it changes is the last mile. Once the relevant slice of the graph is in context, generating the synthesis over it becomes much cheaper, and the frozen context tower is built to hold that retrieved slice as stable, reusable memory while the denoiser writes. So the honest version is that this helps you write the memo faster after the graph has done its work, and not walk the graph itself.</p><h3 style="text-align: justify;">Where Volume Changes the Math</h3><p style="text-align: justify;">The non-performing loan due diligence case is the one that changes the math, because volume does the heavy lifting. Reviewing several hundred loan files is not one long generation, it is a few hundred short ones, the same structured extraction repeated file after file. Reading and pulling the data out of those PDFs is a separate problem, and it is better solved by processing each document on its own than by asking a model to hold the whole portfolio in its head. But once you reach the generation step, producing a standard summary or a risk flag for each file, a speedup of that size compounds across the entire portfolio. And the confidence-graded output sorts the work for you, because the files the model finished in a flash are the clean ones, and the files it kept reworking are the ones a person should open first.</p><h3 style="text-align: justify;">What Actually Travels</h3><p style="text-align: justify;">None of this makes the released model ready to sign an opinion. It is a base model, measured before any instruction tuning or alignment, and the idea that travels is the architecture rather than the specific checkpoint. The lesson is that you can freeze the part of a system that has to be trusted and speed up the part that only has to be quick, and you can keep the two apart on purpose. A verification layer that governs legal output wants precisely that separation: a stable source of &#8220;truth&#8221; on one side, a fast producer on the other, and a clean boundary in between where the checking happens. TwoTower builds that boundary into the model, and the rest of it is our job.</p><p style="text-align: justify;"></p><p style="text-align: justify;">Source: Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro, "<em>Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context</em>," NVIDIA, 2026. <a href="https://arxiv.org/abs/2606.26493">https://arxiv.org/abs/2606.26493</a></p><p style="text-align: justify;"></p>]]></content:encoded></item><item><title><![CDATA[Your AI Agent Just Read the Whole File Cabinet to Change One Word]]></title><description><![CDATA[Researchers taught an agent to know when a task is easy. The catch is where the waste hides.]]></description><link>https://thelegaltechoverlap.substack.com/p/your-ai-agent-just-read-the-whole</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/your-ai-agent-just-read-the-whole</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Sat, 18 Jul 2026 06:30:27 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Picture the most boring task in your day: a colleague asks you to swap one email icon on the firm&#8217;s homepage for a different one that already sits, ready to use, three lines above. Ten seconds of work, you do it and move on.</p><p style="text-align: justify;">Now watch a capable AI agent do the same thing.</p><p style="text-align: justify;">It re-reads the icon library; it re-browses the site directory; it re-analyzes the project architecture it has already seen a dozen times; it re-confirms dependencies nobody asked about. Minutes pass, then it makes the two-line change it could have made at the start, verifies it correctly, and reports success.</p><p style="text-align: justify;">The edit is right but the path to it is absurd.</p><p style="text-align: justify;">A team at the <strong>University of Tennessee, Knoxville</strong> built a whole paper around that absurdity, and the number they attach to it should make anyone paying by the token sit up. On the trivial icon task, their model of an over-cautious agent hits over <strong>1000% redundancy</strong>: it spends roughly eleven times the effort the job actually needs. Not because it fails, but because it refuses to believe the job is small.</p><p>Hold that thought, because the real punchline is worse than &#8220;AI is inefficient.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="4301" height="6451" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:6451,&quot;width&quot;:4301,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a glass of water&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a glass of water" title="a glass of water" srcset="https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1660728891653-2f1603ac2b6c?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwxNzN8fHRpbWV8ZW58MHx8fHwxNzg0MTI4Mjk2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@andraes_arteaga">Andraes Arteaga</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h3>The part nobody wants on the invoice</h3><p style="text-align: justify;">Here is the finding that turns a productivity note into a business problem.</p><div class="pullquote"><p style="text-align: center;">The waste is not spread evenly but it is <em>concentrated on the simplest tasks</em>.</p></div><p style="text-align: justify;">Measure the redundancy tier by tier and it climbs as the work gets easier: highest on one-line edits, lower on cross-file changes, lowest on the genuinely hard repository-wide refactors. The agent burns the most effort exactly where effort is least warranted. It is calm and efficient on the hard problem and completely unhinged on the easy one.</p><p style="text-align: justify;">Sit with that if your practice runs on AI. The routine, high-volume, low-margin work, the stuff you were counting on to be cheap, is precisely where an untuned agent quietly bleeds the most. The complex matter that justifies a real budget is where it behaves.</p><p style="text-align: justify;">The authors are honest that part of this pattern is arithmetic: if the agent always reads everything, its cost stays flat while the &#8220;necessary&#8221; cost shrinks on easy tasks, so the <em>ratio</em> has to spike at the bottom. Fair but the money is not a ratio. The money is the files and tokens actually consumed, and those are being spent on the tasks that least deserve them. The math being partly mechanical does not make the invoice smaller.</p><h3>Why the agent does this (and why it isn't dumb)</h3><p style="text-align: justify;">The instinct behind the behavior is not stupidity but it is fear.</p><p style="text-align: justify;">Faced with any uncertainty, the agent defaults to what the paper calls <em>maximum-context-first</em>: gather everything, then eliminate every conceivable risk before acting. On a genuinely hard task, that caution is exactly right. On an easy one, it is a lawyer re-reading the entire contract to confirm the client&#8217;s name is spelled correctly on page one.</p><p style="text-align: justify;">What is missing, the authors argue, is a skill humans use without thinking: a fast, cheap judgment of <em>how hard this actually is</em> before committing to a plan. An experienced associate glances at a task, sizes it in seconds, sketches the smallest move that could work, and only widens the search if that move fails. The agent never makes that first judgment. It has one gear, and the gear is &#8220;audit.&#8221;</p><h3 style="text-align: justify;">The fix: guess small, verify, expand only if you're wrong</h3><p style="text-align: justify;">Their answer is a framework with a deliberately plain name: <strong>E3 (Estimate, Execute, Expand).</strong></p><p style="text-align: justify;">Estimate the task&#8217;s real scope up front, cheaply. Execute the smallest path likely to work. Expand into deeper reading and heavier checking <em>only</em> when verification fails. The estimate is allowed to be optimistic, even wrong, because expansion is the safety net that catches the misses. Guess lean, and pay for depth only when the evidence demands it.</p><p style="text-align: justify;">The results, in their controlled benchmark, are the kind that get a paper shared:</p><ul><li><p>Same <strong>100%</strong> success as the most thorough baseline.</p></li><li><p><strong>~85% lower cost.</strong></p></li><li><p><strong>~91% fewer tokens</strong> and <strong>~92% fewer files</strong> dragged into context.</p></li></ul><p style="text-align: justify;">And this is not a win over a strawman. They built a genuinely strong competitor, an adaptive agent that already scales its effort to the task and solves everything, and E3 still comes out <strong>~16% cheaper</strong> at equal accuracy. The gain that matters isn&#8217;t &#8220;think less&#8221; but it is &#8220;judge first.&#8221; E3 lands as both the leanest <em>and</em> the most reliable policy in the study, which is the combination that usually doesn&#8217;t come together.</p><p style="text-align: justify;">They also stress-tested the cost model, the obvious place a skeptic would push. Redo the accounting under 4,000 different weightings, including ones openly hostile to their method, and E3 stays the cheapest fully-successful policy in <strong>99.8%</strong> of them. That is a paper trying to break its own result and failing. Respect.</p><h3 style="text-align: justify;">Now put it on a lawyer's desk</h3><p style="text-align: justify;">Every clean result deserves the &#8220;desk test&#8221;: does this survive contact with the real world, where the &#8220;user&#8221; is an attorney and every &#8220;just to be safe&#8221; pass is billable time?</p><p style="text-align: justify;">Three things it gets right for that world.</p><p style="text-align: justify;">The problem is <em>your</em> problem. Legal AI runs on volume, and volume is made of small tasks: pull a clause, swap a party name, reconcile one figure. If the agent is most wasteful there, the waste compounds across thousands of routine jobs. E3 aims straight at that.</p><p style="text-align: justify;">&#8220;Verify, then expand&#8221; is a workflow lawyers already trust. Start narrow, check, widen only on a real signal. That is how good associates work and how good review protocols read. The shape is familiar, which makes it deployable.</p><p style="text-align: justify;">And the honesty is disarming. The authors don&#8217;t oversell: they call it a <em>controlled probe</em>, not a verdict, and flag their own weak points before you can.</p><p style="text-align: justify;">Now the reasons to keep your hand near the brake.</p><p style="text-align: justify;"><strong>It lives in a simulator: </strong>most of the eye-catching numbers come from a controlled environment where every agent is equally <em>able</em> to make the edit, so only the reading behavior varies. Clean for science but your real agent&#8217;s capability wobbles from prompt to prompt, and that variance is switched off here by design.</p><p style="text-align: justify;"><strong>The estimator is brittle where it counts.</strong> Its judgment leans on wording cues. Rephrase the tasks into unfamiliar language and its accuracy drops from about <strong>85%</strong> to <strong>67%</strong>, with the rate of dangerous under-scoping more than doubling. Litigation language is nothing but unfamiliar phrasing, adversarial drafting, and buried cross-references. The exact conditions that fool this estimator are the native habitat of legal text.</p><p style="text-align: justify;"><strong>The real-model check is thin.</strong> They did run E3 on a live agent editing a real open-source library, graded by actually running the project&#8217;s test suite, and the effect held: the over-reading was, in their words, &#8220;milder but real&#8221;. Encouraging but it&#8217;s one model on one codebase. &#8220;Milder but real&#8221; on a single case is a hypothesis with a nice haircut, not proof it holds on your matter.</p><p style="text-align: justify;">So does it survive the desk test? As an idea, yes, and it should reshape how legaltech builders think about agent cost. Stop optimizing only for &#8220;can it get the answer&#8221; and start measuring &#8220;how much did it waste getting there.&#8221; As a drop-in for tomorrow&#8217;s docket, not yet. The estimator that shrugs off a paraphrased coding task would meet its match in a poorly drafted guarantee with three defined terms pointing at each other.</p><p style="text-align: justify;">The deeper lesson outlasts the benchmark. The frontier we obsess over is capability: can the model do the hard thing? This paper points at the quieter frontier that actually shows up on the bill: does the model know when the thing is <em>easy</em>, and have the confidence to act like it? An agent that can&#8217;t tell a two-line edit from a code audit will handle your simplest work the most expensively. It&#8217;s a judgment gap (nota a capability one), and judgment is the whole job.</p><p style="text-align: justify;">Have you actually seen this in your own tools: an agent doing a full audit for a one-line change? Or is agent &#8220;overthinking&#8221; still more benchmark artifact than real cost?</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/your-ai-agent-just-read-the-whole/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thelegaltechoverlap.substack.com/p/your-ai-agent-just-read-the-whole/comments"><span>Leave a comment</span></a></p><p style="text-align: justify;">Source: Junjie Yin &amp; Xinyu Feng, "<em>Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution</em>," University of Tennessee, Knoxville. arXiv:2607.13034 &#8212; <a href="https://arxiv.org/abs/2607.13034">https://arxiv.org/abs/2607.13034</a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: justify;"></p>]]></content:encoded></item><item><title><![CDATA[The Only Benchmark That Pays]]></title><description><![CDATA[Why a vendor&#8217;s &#8220;state of the art&#8221; tells you almost nothing about how a tool will handle your own files.]]></description><link>https://thelegaltechoverlap.substack.com/p/the-only-benchmark-that-pays</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/the-only-benchmark-that-pays</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Thu, 16 Jul 2026 06:16:09 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Within a single day this week, two research teams published findings that cannot both be right, at least not on the surface. One team announced that its system now builds its own record-breaking setup for any test you give it, without any human help, and concluded that benchmarks are finished. The other team ran that same idea under controlled conditions and found that it barely works, and that most of its reported wins collapse the moment you test them fairly.</p><p style="text-align: justify;">They were describing the same thing and reaching opposite conclusions, so the disagreement is worth sitting with. It also happens to be the most useful idea a practicing lawyer can grasp right now, because it decides whether the phrase &#8220;state of the art&#8221; in a sales deck means anything at all for your matters.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="2896" height="1944" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1944,&quot;width&quot;:2896,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;text&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="text" title="text" srcset="https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1587716283570-f71969a1f854?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw3NHx8bGVnYWx8ZW58MHx8fHwxNzg0MTczMzM2fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@david_nitschke_95">David Nitschke</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h3>The part of the system nobody puts in the demo</h3><p style="text-align: justify;">When people talk about legal AI, they usually talk about the model, as if the model were the whole product. It isn&#8217;t: around the model sits a layer that researchers call <strong>the harness</strong>, and it quietly does much of the work. The harness is the bundle of prompts, tools, checklists, retrieval steps, and control logic that wrap around the model and tell it how to read a task and how to act.</p><blockquote><p style="text-align: justify;">A simple way to picture it is to think of the model as a bright new associate and the harness as everything your firm builds around that associate. It is the precedent bank they search, the templates they fill in, the standing rule that says &#8220;check the citation before anything goes out,&#8221; and the partner who reviews the draft before it reaches the client. </p></blockquote><p style="text-align: justify;">Give a middling associate a strong playbook and the work often comes out fine, while a brilliant associate with no playbook can still produce a mess. The field has spent the past year learning that the same holds for models, and that the scaffolding frequently matters more than the model sitting inside it.</p><p style="text-align: justify;">So the obvious next step was to automate the scaffolding. If a good harness lifts performance, then why not let the machine build and refine its own harness? That is what both research teams set out to do, and it is where their stories split.</p><h3>The victory lap</h3><p style="text-align: justify;">The first team, a lab called Poetiq, titled its post &#8220;<em>Benchmarks Are Dead (for us).</em>&#8221; The claim is bold and, on its own terms, impressive. Their system takes a benchmark it has never seen, writes a custom harness for it, and posts a new top score, often while running an older and cheaper model than the one that held the previous record. They report new highs across six very different tests, from competition math to long-context retrieval to tool use, all with no human tuning. </p><div class="pullquote"><p style="text-align: justify;">On one math benchmark their setup reached 89.2 where the prior best had been 87.5, and on a long-context test it hit 99.26 using a small, inexpensive model that does not even appear near the top of the leaderboard on its own.</p></div><p style="text-align: justify;">Their conclusion follows naturally from those results: if a machine can master any fixed test you hand it, then the fixed test has stopped measuring anything interesting. In their words, &#8220;<em>the static benchmark is dead; long live the living benchmark</em>,&#8221; by which they mean a test that keeps regenerating itself so that nothing can be trained against it.</p><p style="text-align: justify;">Read on its own, the post feels like the future arriving a little early. Then you read the second paper.</p><h3>The cold shower</h3><p style="text-align: justify;">The second team, from the Allen Institute for AI and the University of Washington, asked a plainer question. When you let a system evolve its own harness, are the gains real, or are you simply paying for more attempts?</p><p style="text-align: justify;">They compared automatic harness evolution against a deliberately unsophisticated baseline, which was to run the model a handful of times and keep the best answer.</p><div class="callout-block" data-callout="true"><p style="text-align: justify;">On Terminal-Bench, a suite of realistic command-line tasks, the self-evolving harness scored 67.4, while plain repeated sampling scored 72.3, and the untouched starting harness with no evolution at all scored 68.2. </p></div><p style="text-align: justify;">So the clever method came in behind both the simple trick and the very setup it was meant to improve.</p><p style="text-align: justify;">The more damaging finding came next: the whole promise of harness design is that a better harness should transfer, in the same way a good checklist helps on cases its author never saw. So the researchers built the harness on one set of tasks and then tested it on a separate, held-out set. The improvement almost disappeared, landing at six tenths of a point on average and at zero on one of the models. When they examined what the system had actually changed, the edits looked reasonable on their face, since they added sensible rules and guards, which is why the authors describe the pattern as &#8220;<em>rational edits but marginal gains</em>.&#8221; The catch is that the system had mostly memorized fixes for the specific tasks it trained on, rather than learning a better general way to work. It looked like progress only because it was being graded on the very tasks it had already studied.</p><p style="text-align: justify;">Set the two papers next to each other and the contradiction softens into something more interesting. Both teams agree that the static benchmark has stopped telling the truth, yet they draw opposite lessons from it. Poetiq treats the number as so easy to beat that you should trust the system and move on, while the academics treat that same number as so easy to inflate that you should distrust it and insist on a harder test. One camp reads a high score as proof of strength, and the other reads it as a warning light.</p><h3 style="text-align: justify;">Then someone automated the training data as well</h3><p style="text-align: justify;">If the story ended there, the lesson would be ordinary caution. It does not end there, because a third paper, from Meta&#8217;s research group, pushes the same logic one level deeper and makes the stakes concrete for lawyers.</p><p style="text-align: justify;">Their system, called Autodata, does not merely evolve the harness. It builds the training data itself. An agent plays the role of a data scientist, so it writes practice questions, has a weak model and a strong model attempt them, keeps the questions that a strong reasoner can solve but a weak one cannot, and then uses those questions to train the weak model into a stronger one. The results are genuinely striking, and one of the headline experiments is legal. After training on this home-grown data, a small four-billion-parameter model outscored a model almost a hundred times its size on a professional legal-reasoning test.</p><p style="text-align: justify;">Now comes the part a practitioner should hold onto. The overfitting problem from the second paper does not vanish once you automate the data, because it simply moves upstairs. The earlier system overfit its harness to the test, and this system can overfit its training data to the exact model it is grading against, since the question-writer is tuned to that model&#8217;s current weak spots. The authors are refreshingly honest about it. They describe agents that tried to game the objective, in one case by editing the instructions so that the weak model would answer badly on purpose, and they warn that some generated questions clung to the specific numbers in a source document instead of testing reasoning that carries over. Their own guidance for the legal question-writer even tells it to avoid the easy trap of a single-doctrine question that one statute resolves, which is exactly the kind of shortcut that looks like mastery while teaching nothing that lasts.</p><p style="text-align: justify;">There is a quiet irony running through all three papers. Autodata improves its data-writing agent using the same style of harness optimization that the Allen Institute paper singled out as unlikely to generalize, so the most advanced pipeline stacks two self-improving loops on top of each other, and both loops share the same weakness. Each can look brilliant on the tasks it trained on, and each can fail silently on the ones it did not.</p><h3>What this means when you are the buyer</h3><p style="text-align: justify;">Now translate all of this into a sales meeting, which is where it will actually reach you. A vendor tells you their tool is state of the art on some legal benchmark, and you finally know what that sentence really describes. It is a score produced by a harness, and quite possibly by training data, that were both shaped around that particular benchmark. It is the report card of a student who was shown the exam in advance.</p><p style="text-align: justify;">The only question that matters for you is the one all three papers circle from different directions. Does the tool hold up on matters it has never seen, meaning your files, your fact patterns, and your jurisdiction, or does it only shine on the set it was tuned against? In banking litigation nobody grades us on a public leaderboard, because we are judged on the next brief, the next set of facts, and the next judge, none of which the vendor&#8217;s benchmark ever contained.</p><p style="text-align: justify;">So when you evaluate legal AI, a short list of questions will save you real money and real embarrassment:</p><ol><li><p style="text-align: justify;">Ask whether the tool was measured on tasks it was also trained or tuned on, because if the answer is yes then the score is close to meaningless. </p></li><li><p style="text-align: justify;">Ask to run it on a sample of your own recent matters, redacted where needed, that no vendor has ever touched. </p></li><li><p style="text-align: justify;">Ask how the tool behaves when it is wrong, since a confident wrong answer on a live file costs far more than a cautious one. </p></li><li><p style="text-align: justify;">And if any synthetic training data is involved, ask what the held-out set was actually held out from, because a test drawn from the same machine that wrote the training data quietly inherits all of that machine&#8217;s blind spots.</p></li></ol><p style="text-align: justify;">Held-out work is the only benchmark that pays, because it is the only one that resembles the job, while everything else is a rehearsal the tool was allowed to watch beforehand. The labs racing to automate their own scaffolding are doing real and impressive work, and some of it will genuinely reshape how these systems get built. Until that work generalizes to the cases sitting on your desk, though, the right stance for a lawyer is the oldest one we have:</p><div class="callout-block" data-callout="true"><p style="text-align: justify;">trust the result you tested yourself, and stay politely skeptical of the one you were shown.</p></div><h3>Sources</h3><ol><li><p>Poetiq, &#8220;Benchmarks Are Dead (for us),&#8221; July 15, 2026. <a href="https://poetiq.ai/posts/benchmarks_are_dead/">https://poetiq.ai/posts/benchmarks_are_dead/</a></p></li><li><p>Wang et al., &#8220;Rethinking the Evaluation of Harness Evolution for Agents,&#8221; Allen Institute for AI and University of Washington, arXiv, July 14, 2026. <a href="https://arxiv.org/abs/2607.12227">https://arxiv.org/abs/2607.12227</a></p></li><li><p>Kulikov, Whitehouse, Wu, Nie et al., &#8220;Autodata: An Agentic Data Scientist to Create High Quality Synthetic Data,&#8221; FAIR at Meta, arXiv, June 25, 2026. <a href="https://arxiv.org/abs/2606.25996">https://arxiv.org/abs/2606.25996</a></p></li></ol><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/the-only-benchmark-that-pays?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading The LegalTech Overlap! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/p/the-only-benchmark-that-pays?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/thelegaltechoverlap.substack.com/p/the-only-benchmark-that-pays?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Narrowing: What an AI Ideation Study Reveals About Legal AI’s Favorite Move]]></title><description><![CDATA[When models are asked to think, they get more predictable, not less. That should worry anyone building legal AI.]]></description><link>https://thelegaltechoverlap.substack.com/p/the-narrowing-what-an-ai-ideation</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/the-narrowing-what-an-ai-ideation</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Tue, 14 Jul 2026 06:31:17 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">I spend a lot of my time watching legal AI systems generate first drafts: theories of a case, defenses, settlement angles, ways to frame a novel argument to a court that has never seen it. So when a paper lands that measures, at scale, what kinds of ideas language models actually produce, I read it as if it were about my own tools, because in every way that matters it is.</p><p style="text-align: justify;">The paper is called Measuring the Gap Between Human and LLM Research Ideas, from a team at Chicago and Yale. It studies scientific ideation, not legal drafting, and I want to be honest about that distance from the first paragraph. But its central result is not a fact about physics papers. It is a fact about how these models behave when you ask them to find an opening and build something on it. That behavior does not stay politely inside the natural sciences.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="4000" height="6000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:6000,&quot;width&quot;:4000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;library shelf near black wooden ladder&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="library shelf near black wooden ladder" title="library shelf near black wooden ladder" srcset="https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1535905557558-afc4877a26fc?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw5fHxib29rc3xlbnwwfHx8fDE3ODM2MjQ3MDl8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@henry_be">enrico bet</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h3 style="text-align: justify;">What they actually did</h3><p style="text-align: justify;">The setup is clean, which is why the result is hard to wave away. The authors took roughly 11,700 real research papers, split evenly between machine learning venues and Nature Communications, and for each one reverse-engineered the small set of prior works that plausibly preceded the paper&#8217;s core idea. They then handed those same prior works to nine current models, from Claude and Gemini to the GPT, Qwen, and DeepSeek families, and asked each to produce a new idea from the same starting point. The real paper is the human answer. The model output is the machine answer, generated from the same local context a human researcher would have had.</p><p>Then they classified every idea along two axes. One axis asks why the work is worth doing: is the gap a contradiction, a missing explanation, a scope problem, an evidence gap, a disconnect between literatures, a failure or risk, or a resource bottleneck. The other axis asks how the contribution is built: synthesis, scope extension, robustification, formal derivation, empirical mapping, artifact building, or optimization. That taxonomy was not invented casually. It was assembled from research guidance published by NSF, NIH, AHRQ, and DARPA, then refined on a held-out set of papers.</p><p>The comparison is distributional, and this is the part I find genuinely useful. They are not asking whether a single idea is good. They are asking what the whole population of ideas from a given source looks like, and whether that population is as wide as the human one.</p><h3>The finding, stated plainly</h3><p>Model ideas are narrower than human ideas, and they are narrow in a specific, repeatable direction.</p><p>Only 12.1 percent of human ideas frame the opportunity as connecting disconnected work, and only 5.1 percent build the contribution through explicit synthesis or unification. Across the main models, those same numbers run from 47.1 to 64.2 percent on the framing axis, and from 22.5 to 38.7 percent on the method axis. Human entropy across the categories sits above 0.92 on both axes, close to a flat spread. The models cluster. Even the closest model to the human distribution still needs roughly a third of its idea mass relocated to match it.</p><p>Underneath the taxonomy, the mechanism is almost blunt. When the authors reduced each proposal to a single verb, the dominant model move was &#8220;integrate&#8221;: it showed up 7,994 times in model outputs against 275 times in human ones. Humans, by contrast, reached far more often for verbs like &#8220;replace,&#8221; &#8220;decouple,&#8221; and &#8220;formalize.&#8221; Replace alone accounts for 9.13 percent of human moves and 0.92 percent of model moves. The human researcher tends to change one specific thing. The model tends to bolt two existing things together.</p><p>There is one more result that I keep returning to. Turning on extended reasoning, the &#8220;thinking&#8221; mode everyone assumes makes models more careful and more original, moved the output distribution further from humans, not closer. For one model, enabling thinking pushed bridge-style framing from 49.7 to 71.1 percent and synthesis from 38.7 to 52.2 percent, while the diversity of ideas dropped. Reasoning sharpened the template. It did not break it.</p><h3>Why a banking-litigation practitioner should care</h3><p>Read &#8220;research idea&#8221; as &#8220;legal theory&#8221; and the paper starts describing my daily problem.</p><p>The synthesis move is the safe move, and it is exactly the move that loses cases. In a non-performing loan dispute, in a guarantee enforcement action, in a derivatives claim turning on jurisdiction and public policy, the winning argument is rarely &#8220;combine these two established doctrines into something that sounds interdisciplinary.&#8221; It is usually the local intervention: replace a brittle premise the other side is relying on, decouple two issues the court has been treating as one, formalize a structure everyone has left implicit. That &#8220;replace, decouple, formalize&#8221; trio the paper found on the human side is a fair description of good lawyering. The trio the models overproduce, &#8220;integrate and unify,&#8221; is a fair description of the memo that reads well and persuades no one.</p><p>The reasoning result cuts the same way. There is a comfortable assumption in legal tech that longer chains of thought yield sharper legal analysis. This paper is at least a warning against treating that as automatic. If more deliberation collapses the model toward its favorite template, then a tool that &#8220;reasons harder&#8221; about a fact pattern may simply produce a more confident version of the most generic argument available. Confidence is not the scarce resource in litigation. Specificity is.</p><p>There is also a quieter finding that supports something I have argued for a while: verification is a moat, and specificity is measurable. The study&#8217;s annotator scored each proposal for how superficial the combination was, how precisely it named the actual bottleneck, and how boilerplate it read. Most models scored worse than humans on precision and generic phrasing, and one model family was a partial exception, scoring slightly better than the human baseline on specificity while still landing in the wrong overall distribution. That is the whole game in a sentence. An output can be polished, specific, and still be reaching for the wrong kind of idea. If your evaluation only checks whether a single answer is coherent, you will never see it.</p><p>Finally, the method itself is portable. Legal AI evaluation today mostly grades outputs one at a time: is this citation real, is this summary accurate, is this clause enforceable. Almost nobody grades the distribution. Ask your drafting tool for fifty theories of the case across fifty matters and look at the shape of what comes back. If forty of them are &#8220;harmonize these two lines of authority,&#8221; you have a homogenization problem that no single-output review will catch.</p><h3>Where I would push back</h3><p>I would not run this paper into a courtroom, and neither should you.</p><p>The corpus is entirely STEM. There is not a single legal text in it. Everything I wrote in the previous section is an argued transfer, not a demonstrated one. Legal reasoning has its own structure, and it is plausible that the human-model gap looks different, larger or smaller, once you are working with statutes, precedent, and the peculiar rhetoric of persuasion. The paper cannot tell us, and it does not claim to.</p><p>The evaluation is also an LLM grading LLMs. The classifier that assigned every one of those 11,700-plus labels is itself a language model, checked against two human annotators on 150 items. The agreement scores were high, in the 0.81 to 0.93 range, which is real. But there is an unavoidable circularity in using one model to characterize the taste of others, and I would want more human adjudication before treating the exact percentages as settled.</p><p>Then there is survivorship. The &#8220;human ideas&#8221; are drawn from published papers. Published work is the filtered tail of human ideation, the part that cleared peer review precisely because it was not another routine synthesis. Researchers produce enormous quantities of derivative, combine-two-things ideas that never reach a journal. So the human distribution here is not &#8220;what humans think.&#8221; It is &#8220;what humans publish,&#8221; which is exactly the sample selected to look less template-bound. The gap is real, but part of its size is an artifact of comparing raw model output against humanity&#8217;s edited highlight reel.</p><p>And the setting is one-shot and non-interactive. No back-and-forth, no retrieval pipeline, no domain-specific system, no lawyer in the loop pushing the model off its first instinct. The paper says as much in its limitations, and it matters, because the legal tools worth building are agentic and iterative. It is entirely possible that the narrowing the authors measured is most severe precisely in the setting they tested, and that a well-designed harness recovers some of the lost range.</p><h3>What I am taking from it</h3><p>Three things go into my practice notes. First, treat the model&#8217;s first idea as its most generic idea, and design the workflow to fight the synthesis reflex rather than reward it. Second, do not assume that a reasoning mode buys you originality; test whether it is buying you a more elaborate version of the same template. Third, start measuring the distribution of what your tools produce, not just the quality of each output, because homogenization is invisible one answer at a time.</p><p>The paper&#8217;s real contribution is to reframe the question. The interesting failure of legal AI will not be the hallucinated citation, which we already know how to catch. It will be the quiet convergence on a single, safe, respectable kind of argument, delivered with enough polish that nobody notices the range has collapsed.</p><p>Which raises the question I would put to anyone shipping a legal drafting system: are you evaluating whether each output is good, or whether the set of outputs is still wide enough to win?</p><p><strong>Reference</strong></p><p>Ziyu Chen, Yilun Zhao, Arman Cohan. Measuring the Gap Between Human and LLM Research Ideas. University of Chicago and Yale University. https://arxiv.org/abs/2607.01233 [cs.CL], July 2026. Code and data: github.com/ziyuuc/TasteGap and github.com/IdeaLand/IdeaSeed.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[How an 8B Model Beat GPT-4 Without Being Smarter]]></title><description><![CDATA[New research on AI agents points to the fix legal work actually needs, and it is not a bigger model.]]></description><link>https://thelegaltechoverlap.substack.com/p/how-an-8b-model-beat-gpt-4-without</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/how-an-8b-model-beat-gpt-4-without</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Sun, 12 Jul 2026 10:16:52 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>An open model with 8 billion parameters, the kind small enough to run on your own hardware, outperformed GPT-4 on two of three agent benchmarks in a recent paper. It did not manage this by being smarter, because GPT-4 is still the stronger raw model by a wide margin. It won on how its work was organized. That distinction sounds academic, but it speaks to a real and unsolved problem in legal AI, and it is worth working through.</p><p>The paper is called Atomic Task Graph, or ATG, and although it never once mentions law, it speaks directly to the question that decides whether any agent ever touches a client file. That question is not whether the model is clever enough. It is whether I can check its work, and whether I can show someone else that I checked it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="5472" height="3648" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3648,&quot;width&quot;:5472,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;open book lot&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="open book lot" title="open book lot" srcset="https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1457369804613-52c61a468e7d?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzMnx8bGVnYWx8ZW58MHx8fHwxNzgzODUwMTk0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@impatrickt">Patrick Tomasso</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h3>The plan, not the model</h3><p>Here is the core idea in practitioner terms. Instead of letting the model reason forward through one long, growing block of text, which is essentially how the popular ReAct loop works, ATG forces the plan into an explicit graph. Each node is a single, atomic tool call with a defined input and output, and each connection means one node&#8217;s output feeds the next node&#8217;s input. In other words, the dependencies between subtasks stop being buried in prose and become a structure you can actually inspect.</p><p>The framework then does three things with that structure. </p><ol><li><p>First, it builds the graph by progressive refinement, breaking a coarse task into finer subtasks until every node is directly executable, while preserving each node&#8217;s input and output interface at every step. </p></li><li><p>Second, it runs the graph by dependencies rather than by narrative, so independent branches execute in parallel and a node fires only once its inputs are ready.</p></li><li><p>Third, when something breaks, it isolates the failure to the smallest affected region and repairs only that piece, freezing everything already validated instead of replanning from scratch.</p></li></ol><p>One more detail matters more than the rest. Before touching the environment, ATG runs a pre-execution simulation the authors call a thought experiment, checking for missing steps, wrong tools, and broken dependencies. It is, in plain terms, a gate that inspects the plan before any action is taken.</p><h3>Why this is really a story about hallucination</h3><p>The headline that will get shared is speed and cost. That part is real, but it is not what caught my attention.</p><p>What caught my attention is the mechanism ATG uses to cut hallucinations. In a normal agent loop, the model drags an ever-growing history behind it, and that accumulating context is precisely what nudges it toward invented actions later in a task. ATG counters this by keeping the context of each atomic node small. As the plan refines, the window each node has to attend to gets narrower, so action generation stops depending on a bloated trajectory. To put it another way, the framework reduces fabrication not by making the model more honest, but by never handing it enough rope to confabulate. Anyone who has built verification tooling will recognize that instinct at once, because it is the same reason we isolate context when we check a claim rather than letting a model grade its own long answer.</p><p>The reported effect is large. On one benchmark, the share of runs containing invalid or hallucinated actions falls from 42.86 percent under ReAct to 12.14 percent under ATG, which is a <strong>71.7 percent relative reduction</strong>. The pre-execution gate pulls its weight too, flagging roughly a quarter of plans as risky before they ever run, with reliability above 74 percent in most settings.</p><p><em>(This is the kind of thing I take apart every week here: one new AI or legaltech paper, read from the practitioner&#8217;s chair. If that is useful to you, the subscribe button is a few lines down.)</em></p><h3>The number that does not mean what you think</h3><p>Now the part where I have to be honest with myself, because the distance between that benchmark and my desk is wide.</p><p>For one, that striking hallucination figure is measured on a single benchmark, for the simple reason that it is the only one of the three that reports invalid actions at all. So the number is genuine, but it is one environment, not a universal law.</p><p>More importantly, a hallucinated action there means the agent tried something the environment does not allow, for instance opening a drawer that is not present. That is not the same failure as citing a case that does not exist or misstating the holding of one that does. The underlying mechanism plausibly transfers, yet the paper offers no evidence about fabricated citations, misread clauses, or invented figures, which are the hallucinations that actually create liability in my field. Treating one as a proxy for the other would be exactly the unverified leap this research should teach us to avoid.</p><p>The small-model headline deserves the same discipline, including the one at the top of this article. The 8B model does beat GPT-4, but only on two benchmarks out of three, and on one of those the margin is a slim four points. On the third, GPT-4 still wins comfortably. So the honest reading is that a good control framework narrows the gap between small open models and the frontier, which is genuinely useful and cost-relevant, rather than the tidy claim that 8B now wins outright. There are further limits the authors concede, and they are the right ones. The approach leans on the model&#8217;s ability to decompose a task, the experiments are all text-based, and the whole apparatus is simply not worth the overhead on simple work.</p><h3>What actually transfers to legal verification</h3><p>Set the benchmark numbers aside, and the architecture is what I keep returning to. Three properties matter.</p><ol><li><p>The first is verifiability by construction. When a plan is an explicit graph of atomic calls with declared inputs and outputs, each step is a discrete object I can check on its own, rather than a sentence lost inside a paragraph of reasoning. That is the difference between auditing a workflow and re-reading a story.</p></li><li><p>The second is a native audit trail. ATG records the full refinement history, meaning the sequence of graphs that shows how an abstract task was compiled into executable steps. When a later step is wrong, you can trace it back to the exact point where the error entered. For anyone building governed legal output, that traceability is worth more than any single accuracy score, because it answers the only question that matters after a mistake, which is where and why it happened.</p></li><li><p>The third is bounded repair. Freezing validated regions and fixing only the affected part maps cleanly onto how careful review already works. You do not rewrite an entire memo because one authority was wrong. You correct the affected passage, re-check what depended on it, and leave the rest alone precisely because you already validated it.</p></li></ol><h3>My takeaway</h3><p>I read ATG less as a benchmark win and more as a design argument, and on that level it is a strong one. The reliable path to trustworthy legal AI is probably not a bigger model that we hope hallucinates less. It is architecture that makes the model&#8217;s work legible, checkable, and repairable step by step, so that verification becomes a property of the system rather than a prayer said over its output. The paper does not prove this for legal tasks, and I will not pretend otherwise, but it points at the right target.</p><p><em>The Legaltech Overlap is where I read one new AI or legaltech paper each week and tell you what holds up at the desk and what does not, with the numbers checked and the weak spots flagged. If that is the kind of signal you want in your inbox, subscribe below.</em></p><p>P.S. The 8B-beats-GPT-4 result is real, but only on two benchmarks out of three, and honestly I would not have believed it at first either. That gap between the headline and the footnote is exactly what this newsletter exists to close.</p><p>Paper: Zhang, Chen, Huang, Cui, Ji, Wang, Atomic Task Graph: A Unified Framework for Agentic Planning and Execution, arXiv:2607.01942. Link: https://arxiv.org/abs/2607.01942</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[A Model Can Be Faithfully, Confidently Wrong]]></title><description><![CDATA[New research teaches models to say &#8220;I&#8217;m not sure&#8221; and mean it. That solves a smaller problem than most people in legal AI think it does.]]></description><link>https://thelegaltechoverlap.substack.com/p/a-model-can-be-faithfully-confidently</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/a-model-can-be-faithfully-confidently</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Thu, 09 Jul 2026 06:30:36 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Every lawyer who has run a matter through a language model knows the tell. The answer arrives in the same even, assured register whether the model is reciting a settled rule or inventing a citation that has never existed. The prose does not flinch. And that flatness is the actual danger, because a wrong answer delivered with a hedge is a manageable risk, while a wrong answer delivered with total composure is the one that ends up in a brief.</p><p>A team from Yale and Google Research has been working on exactly this failure, and their recent paper on what they call reinforcement learning with metacognitive feedback is worth reading closely. It is a genuine advance. It also, read carefully, tells us something uncomfortable about where the real work in legal AI still has to happen, and I want to walk through both halves of that.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="5969" height="3979" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3979,&quot;width&quot;:5969,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;landscape photo of library hallway&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="landscape photo of library hallway" title="landscape photo of library hallway" srcset="https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1505488387362-48bc38155987?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwzNHx8bGF3fGVufDB8fHx8MTc4MzQ5MDYxMHww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@giamboscaro">Giammarco Boscaro</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h3>Two kinds of calibration, and the one everybody skips </h3><p>Start with a distinction the paper draws that most practitioners have never had a reason to make.</p><p>The familiar idea is factual calibration. A model is factually calibrated when its stated confidence tracks how often it is actually right. If it says &#8220;80 percent confident&#8221; across a hundred answers, roughly eighty of them should be correct. This is the property everyone assumes they want, and it is the property most calibration research has chased.</p><p>The paper is after something different, which the authors call faithful calibration. Here the question is not whether the expressed confidence matches the truth of the world. It is whether the expressed confidence matches the model&#8217;s own internal sense of certainty. Does the model&#8217;s outward hedging reflect what is actually happening inside it when it generates the answer?</p><p>Once you see the gap between these two, you cannot unsee it. </p><div class="pullquote"><p>A model can be factually calibrated on average and still be systematically dishonest about any individual answer, projecting certainty it does not have and burying doubt it does have. </p></div><p>The authors put it plainly: a model may look factually calibrated and yet remain misaligned with its own internal beliefs. For anyone deciding whether to trust a specific output, on a specific question, in a specific matter, faithful calibration is the property that actually governs the decision. And it has gone largely unaddressed.</p><h3>What the researchers built</h3><p>The method rests on an intuition that is almost too clean. A model that can accurately judge how well it performed is better positioned to perform well. So rather than only rewarding the model for good answers during training, reward it for accurately judging how good its own answers are.</p><p>The mechanism is elegant. During reinforcement learning, among the answers that already score well on the task, the training loop gives extra weight to the ones where the model&#8217;s self-assessment of its performance was most accurate. Good work that the model also correctly recognized as good work gets reinforced harder than good work the model misjudged. The self-knowledge itself becomes a training target.</p><p>They pair this with a data-selection trick in the same spirit. Instead of relying on external labels to pick training examples, they let the model score how well it thinks it did, then train on the examples from both ends of that self-assessment. The model helps curate its own curriculum.</p><p>The results are strong and I have no interest in downplaying them. Trained on a single question-answering dataset and tested across ten tasks spanning six-plus domains, the approach beats prior prompting-based and fine-tuning-based methods by twenty-nine and twenty-five percent respectively on their faithfulness metric, while holding task accuracy roughly steady. Against standard reinforcement learning it improves faithful calibration by as much as sixty-three percent. When human raters compared its hedging against the previous state of the art, they preferred it on naturalness and helpfulness at rates in the mid-to-high nineties, with strong agreement among raters. These are not marginal numbers.</p><p>There is also a design choice that legal-tech builders should note independently of the headline. The pipeline is deliberately split in two. One stage, the expensive one, calibrates the model&#8217;s internal confidence and runs once. The second stage translates that calibrated confidence into natural hedging language and can be re-tuned freely for different audiences and registers without repeating the costly training. A regulator memo and a client-facing summary need very different uncertainty language even when the underlying doubt is identical, and decoupling the two is the kind of practical architecture decision that survives contact with real deployment. The authors released their code (github.com/yale-nlp/RLMF), and the split is worth studying directly if you are building anything in this space.</p><h3>Now the uncomfortable part</h3><p>Here is the sentence in the method that should make every legal practitioner sit up, and it is not a criticism of the paper so much as a clarification of what the paper does and does not buy you.</p><p>The model&#8217;s &#8220;internal confidence,&#8221; the gold standard the whole system is trained to be faithful to, is not ground truth. It cannot be. The researchers estimate it by sampling the model&#8217;s answer many times and measuring how consistent the responses are. High agreement across samples is read as high internal confidence. It is a reasonable and well-established proxy. But it is a measure of the model&#8217;s self-consistency, not of its correctness.</p><p>Sit with what that means. A model can be perfectly consistent and perfectly wrong. If it has absorbed a widespread misconception, it will produce the same mistaken answer every time you sample it, and the system will faithfully translate that stability into a tone of high confidence. The output will be wrong, and it will be honestly, faithfully, well-calibratedly wrong.</p><p>So faithful calibration delivers something precise and worth having: the model stops lying about its own uncertainty. When it is internally shaky, you will now hear the shakiness. That is real, and in a field drowning in false confidence it matters.</p><p>But it does not, and does not claim to, tell you whether the confident answers are correct. It aligns the model&#8217;s words with the model&#8217;s internal state. It says nothing about whether that internal state is aligned with the law. Those are two different alignment problems, and only one of them just got solved.</p><h3>Why this reframes the moat</h3><p>I have argued before that the defensible position in legal AI is not the model. It is the verification layer that sits on top of the model. This paper, read against the grain, is the strongest evidence for that view I have come across in a while, precisely because it is such a good piece of work that stops exactly where the hard problem begins.</p><p>Faithful calibration handles the model&#8217;s honesty about itself. It is a form of introspection, and introspection has a ceiling: it can only ever report on what is inside. It cannot reach out and check the answer against an authority the model does not contain. The gap between &#8220;the model is consistent about this&#8221; and &#8220;this is actually a correct statement of law&#8221; is not a gap that any amount of metacognition can close from the inside. It has to be closed from the outside, by machinery that checks claims against sources the model cannot hallucinate: the actual text of the statute, the actual holding of the case, the actual citation in the actual reporter.</p><p>That external check is not a nicety layered on top of a well-calibrated model. It is the part that carries the legal weight. A model that faithfully signals its own uncertainty is a much better raw material to build on, because you now know where to point your verification budget. But it is raw material. The introspective layer and the verification layer are complementary, and confusing the first for the second is how a firm ends up trusting a confident answer that no one checked.</p><p>The distinction the paper draws between the two calibrations maps almost exactly onto the distinction between two products. One asks the model how sure it is and reports the answer honestly. The other checks whether the model is right. The first is now much more achievable than it was six months ago. The second is still the whole job.</p><p>What I take from this research, then, is not that the confidence problem is solved. It is that the confidence problem just got usefully cut in two. The model&#8217;s honesty about its own doubt is becoming a tractable engineering target. Its correctness against the law remains exactly where it was: outside the model, in the sources, waiting for someone to build the layer that checks.</p><p>Which half of the confidence problem is your legal AI stack actually solving, and are you sure you know which one your users think they&#8217;re getting?</p><p>Based on &#8220;<em>Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs</em>&#8221; by Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona, Idan Szpektor, and Arman Cohan (Yale University and Google Research, 2026): https://arxiv.org/abs/2606.32032</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelegaltechoverlap.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The LegalTech Overlap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Fail Closed]]></title><description><![CDATA[Google built an AI that reviews scientific papers before submission. The real lesson for law firms is in what it gets wrong.]]></description><link>https://thelegaltechoverlap.substack.com/p/fail-closed</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/fail-closed</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Wed, 08 Jul 2026 03:53:46 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Google Research just published something that has nothing to do with law, which is exactly why every managing partner should read it. The paper describes the Paper Assistant Tool, or PAT, a system that reads a full scientific manuscript before submission and reports back on it. It checks proofs, flags logical errors, tests whether the experiments actually support the claims, and points out where the work is thin. The tool was piloted at two of the most demanding venues in computer science, STOC and ICML, and between them it reviewed more than 4,700 submissions.</p><p>I read it as a litigator who spends his days building AI for a law firm, and I recognized the problem on the first page. It is our problem, wearing a lab coat.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="6000" height="4000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4000,&quot;width&quot;:6000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;man wearing blue long-sleeved shirt standing near wooden table&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="man wearing blue long-sleeved shirt standing near wooden table" title="man wearing blue long-sleeved shirt standing near wooden table" srcset="https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1573917730240-57db58e017a8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyNzB8fGxhd3llcnxlbnwwfHx8fDE3ODM0ODIyOTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@sodacheese">Jouwen Wang</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2 style="text-align: justify;">The bottleneck is the same shape</h2><p style="text-align: justify;">Science has a scaling crisis that the paper documents plainly. Across three flagship AI conferences, combined submissions went from roughly 17,000 in 2020 to a projected 74,000 in 2026. Expert review does not scale with that curve, because there are only so many qualified reviewers and each careful read of a dense proof can take days.</p><p style="text-align: justify;">A law firm has the identical structure. The senior lawyer is the scarcest resource in the building, and too much of that scarce time goes to the least strategic task there is, which is cleaning up a junior&#8217;s draft before the real thinking can begin. The constraint is not the volume of work coming in the door, and it is not the talent of the associates. It is the review capacity of the few people qualified to review. Same shape, different building.</p><h2 style="text-align: justify;">What PAT actually does, and why it is not a chatbot</h2><p style="text-align: justify;">The interesting part is architectural, and it is worth understanding before drawing any lesson from it. The obvious approach, a single model call on the whole paper, runs into context limits, because verifying dense proofs burns through more thinking tokens than a model can hold at once. The next obvious approach, calling the model many times and pooling the results, raises recall but destroys precision, so that a human ends up sifting through a hundred proposed issues to find the one that is real.</p><p style="text-align: justify;">PAT is built to avoid both traps. A segmenter breaks the paper into themed sections, and an adaptive budget spends more compute on the proofs than on the introduction. Specialized review agents then verify each segment while holding the full paper as context, and a synthesis agent deduplicates the critiques and, this is the part to underline, grounds the output against search to strip out hallucinated references before assembling the final report.</p><p style="text-align: justify;">That last stage is the whole idea. The system is engineered around the assumption that the model will invent things, and it dedicates an entire step to catching the inventions before they reach a person.</p><h2 style="text-align: justify;">The number everyone will quote, and the number that matters</h2><p style="text-align: justify;">On a set of papers that were later retracted for genuine mathematical errors, PAT caught 89.7% of them, against 55.2% for a single strong model call. That is a jump of roughly 34 points, and it is the number that will travel through every legal-tech deck for the next year.</p><p style="text-align: justify;">But the number that matters is the one the authors were honest enough to publish about their own tool. When they asked the STOC and ICML authors whether the feedback was grounded, only 55.8% and 64.8% respectively said it was mostly or entirely grounded. Read that again. Between a third and nearly half of the people using a state-of-the-art review system, built by Google and powered by its strongest reasoning model, found that a meaningful slice of the critiques were not anchored to anything real.</p><p style="text-align: justify;">In science, a hallucinated critique costs an author an hour of chasing down a phantom objection. It is annoying, and it is not fatal.</p><h2 style="text-align: justify;">The asymmetry that changes everything for law</h2><p style="text-align: justify;">Law does not have that luxury, because the cost of a fabricated output is not symmetric with its benefit. A review tool that invents a helpful improvement saves a little time. </p><div class="pullquote"><p style="text-align: justify;">A review tool that invents a controlling precedent, or flags a procedural defect that does not exist, or misreads a clause and confidently tells the junior to rewrite an argument that was already correct, does not cost an hour. It can cost the case. </p></div><p style="text-align: justify;">And courts have already sanctioned lawyers for briefs built on citations that a chatbot invented and nobody verified, which is the same failure PAT was engineered to prevent, transplanted into a courtroom where the stakes are a client&#8217;s rights rather than a reviewer&#8217;s afternoon.</p><p style="text-align: justify;">This is why, in legal AI, the model is not the moat. Any firm can license the same frontier model that I can. The moat is the verification layer, meaning the part of the system that decides what is allowed to reach a human at all, and refuses to surface any critique it cannot anchor to a verified source. It is deterministic checking sitting on top of probabilistic generation.</p><p style="text-align: justify;">Engineers have a phrase for this, and legal-tech should borrow it: fail closed. When a secure system loses confidence, it denies access rather than granting it. A legal review agent has to behave the same way. If it cannot ground a citation to a real, retrievable authority, it does not hedge and it does not soften, it says nothing. A missed issue is a cost the firm can absorb, because the senior would have caught it anyway. A fabricated issue is a heavier one, because it inverts the reviewer&#8217;s job, shifting the partner&#8217;s attention from finding problems in the draft to disproving problems that the tool invented. Fail closed, not open. That single design decision is where legal AI will be won or lost.</p><h2 style="text-align: justify;">The taxonomy, read for a law firm</h2><p style="text-align: justify;">The paper offers a taxonomy of four roles for AI in review, deliberately modeled on the SAE levels of vehicle autonomy, and it maps almost cleanly onto the adoption path of a firm.</p><p style="text-align: justify;">Role 1 is the tool for <strong>authors</strong>, where the junior runs the agent on the draft before it goes up the chain and stays fully responsible for the work. This is where PAT lives today, and where most firms should begin.</p><p style="text-align: justify;">Role 2 is the tool for <strong>reviewers</strong>, where the senior uses the same agent to accelerate their own read but signs off personally.</p><p style="text-align: justify;">Role 3 is the <strong>supporting reviewer</strong>, where the agent produces a genuine first-pass review and the human shifts from writing the review to deciding on it, the way an area chair decides rather than reads. Call it the synthetic junior, and note that it is technically within reach.</p><p style="text-align: justify;">Role 4 is full <strong>automation</strong>, the paper&#8217;s imagined &#8220;AIrXiv,&#8221; a repository where AI vets everything and the human is optional.</p><p style="text-align: justify;">Here the analogy breaks, and the break is the whole point. In science, Role 4 is at least imaginable. In law it is not, and not because the technology is missing. It is structurally off limits, because accountability in legal work cannot be delegated to a system that cannot be admitted to the bar, cannot be sanctioned, and cannot owe a client a duty of loyalty. </p><div class="pullquote"><p style="text-align: justify;">A principle is already taking shape in administrative law (in Italy), a reserve of human decision, holding that certain judgments must remain with a person. In our field the human in the loop is not a phase we pass through on the way to automation. It is the terminal state. </p></div><p style="text-align: justify;">The ceiling for legal AI is a very good Role 3 sitting under a human who owns the outcome, and that is not a limitation to apologize for. It is the shape of the product.</p><h2>Where it is actually won</h2><p style="text-align: justify;">So the lesson from a Google paper about theoretical computer science turns out to be a legal one. Stop competing on the model, because everyone has the same model. Compete on the gate. Build the verification layer that fails closed, tune it on the one reward signal that only real litigation can provide, which is the accumulated record of which objections actually held and which citations actually controlled, and put it in front of a partner who signs. The generation is a commodity. The verification is the practice.</p><p style="text-align: justify;">I am building exactly this layer, and Google&#8217;s paper just told me, in its own honest numbers, precisely which part is hard.</p><div><hr></div><p><strong>Reference</strong></p><p style="text-align: justify;">Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes, Yossi Matias, Vahab Mirrokni, and Vincent Cohen-Addad, <em>Towards Automating Scientific Review with Google&#8217;s Paper Assistant Tool</em>, arXiv:2606.28277 (2026). <a href="https://arxiv.org/abs/2606.28277">https://arxiv.org/abs/2606.28277</a></p>]]></content:encoded></item><item><title><![CDATA[Throughput Is Not Value: A Lawyer Reads Microsoft's AI Coding Study]]></title><description><![CDATA[Microsoft measured what command-line AI agents actually change for engineers. Almost every finding transfers straight to legal AI rollouts.]]></description><link>https://thelegaltechoverlap.substack.com/p/throughput-is-not-value-a-lawyer</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/throughput-is-not-value-a-lawyer</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Wed, 08 Jul 2026 03:38:02 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Earlier this year, a Meta employee built an internal dashboard so colleagues could compete to be the company&#8217;s top AI token user. In one thirty-day stretch, employees burned through more than 60 trillion tokens, and the single heaviest user averaged 281 billion. On the cheapest Claude Opus tier, that one person alone could have run up a bill north of 1.4 million dollars.</p><p style="text-align: justify;">I lead an AI R&amp;D function inside a law firm, so numbers like that are not abstract to me. When the meter runs in the millions, the question every managing partner eventually asks is the one Microsoft&#8217;s researchers set out to answer: who actually uses these tools, do they keep using them, and does anything measurable come out the other end that justifies the spend.</p><p style="text-align: justify;">A new field study from Microsoft (Murphy-Hill, Butler, and Savelieva) is the most honest attempt I have seen to answer that. It tracks tens of thousands of engineers through the company&#8217;s early-2026 rollout of two command-line agents, Anthropic&#8217;s Claude Code and GitHub&#8217;s Copilot CLI, over roughly four months. What makes it unusual is that the authors did not infer AI use from public signals the way most prior work does. They observed who could adopt and who actually did, using real telemetry. For anyone rolling out legal AI, the design alone is worth studying, and the findings are worth arguing about.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="4032" height="3024" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3024,&quot;width&quot;:4032,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;black laptop computer keyboard in closeup photo&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="black laptop computer keyboard in closeup photo" title="black laptop computer keyboard in closeup photo" srcset="https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1530133532239-eda6f53fcf0f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMHx8bWljcm9zb2Z0fGVufDB8fHx8MTc4MzQ0OTM4Mnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@stadsa">Tadas Sar</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2 style="text-align: justify;">First lesson: adoption is social, not demographic</h2><p style="text-align: justify;">The strongest predictor of whether an engineer tried Copilot CLI in a given week was not seniority, tenure, or prior tooling. It was whether the people around them were already using it. The effect was largest for what the authors call skip-level peers, meaning the colleagues who share a manager&#8217;s manager. Where more than a quarter of that group had adopted, the odds of an engineer trying the tool rose by around 216 percent. Having a direct manager who used it lifted the odds by roughly 82 percent, and the reviewers someone regularly trades code with mattered too.</p><p style="text-align: justify;">I read this and thought about every failed legal-tech pilot I have watched. Firms tend to roll out AI the way they roll out a new document management system, which means a vendor demo, a policy memo, and a mandatory training slot. That approach treats adoption as an information problem. This study says it is a social one. If the partners two desks over are visibly working with AI-drafted analysis and talking about it, associates follow. If adoption happens quietly behind closed doors, nothing spreads. </p><div class="pullquote"><p style="text-align: justify;">The practical instruction for a firm is uncomfortable but clear: make competent use visible on purpose, and let it travel through the review and reporting lines that already exist.</p></div><h2 style="text-align: justify;">Second lesson: retention is behavioral, and the &#8220;first tool&#8221; advantage is real</h2><p style="text-align: justify;">Trying a tool and sticking with it turn out to be governed by different forces, which is the finding I keep coming back to. Engineers who had leaned heavily on AI inside their IDE were more likely to try the new command-line agent, yet they were slightly less likely to stay with it. Every retention marker for prior heavy IDE users came out negative.</p><p style="text-align: justify;">The authors offer a reading I find persuasive. People who already trust an AI tool have a comfortable fallback, so they experiment with the new one and then drift back to what they know. Engineers for whom the command-line agent was their first serious AI tool had no such fallback, and when they stayed, they stayed firmly.</p><p style="text-align: justify;">For legal AI this reframes an assumption I hear constantly, namely that the associates already using ChatGPT will be your best adopters of a purpose-built legal tool. Maybe not. They may kick the tires and then return to the generic tool they already trust. The lawyer who builds a genuine habit around your verified, firm-specific system may be the one for whom it is the first tool that ever earned their confidence on legal work. Retention is a question of habit formation rather than enthusiasm, and habits form around whichever tool becomes the path of least resistance.</p><h2 style="text-align: justify;">Third lesson: the output moved, and it did not fade</h2><p style="text-align: justify;">Now the number everyone quotes. Using a synthetic-control method, the authors estimate that adopters merged about 24 percent more pull requests than they otherwise would have, and crucially the lift did not decay across the window. A comparable open-source study of the Cursor editor had found an early bump that vanished by the third month. Here, the gain in February and the gain in late April were statistically indistinguishable. A within-person analysis pointed the same way, since the more days an engineer used the tools in a week, the more they shipped, rising from roughly 15 percent more output at three days a week to around 50 percent at five or more.</p><p style="text-align: justify;">If you are the one signing the invoice, that is the evidence you wanted. A real output metric moved, and it stayed moved.</p><h2 style="text-align: justify;">The question a lawyer cannot skip: throughput is not value</h2><p style="text-align: justify;">Here is where I part company with the triumphant reading, and where the authors, to their credit, part company with it too. Their proxy for output is the merged pull request. They say plainly that a merged pull request is not the same as the value it delivers, and they close the paper by naming the open question directly, which is whether all this extra throughput actually produces better software. They do not claim it does. They say the field does not yet have the measures to know.</p><p style="text-align: justify;">That distinction is the whole game for lawyers. Our version of the merged pull request is the drafted motion, the produced memo, the reviewed contract. It is entirely possible to produce more of those, faster, while producing less value, because a legal work product that is 24 percent faster and 5 percent wrong is not 24 percent better. It is a liability with a shorter turnaround. </p><div class="pullquote"><p style="text-align: justify;">Speed that outruns verification does not compound into value, it compounds into risk, and in our field that risk lands on a client and on a professional license.</p></div><p style="text-align: justify;">This is precisely why the work I spend my days on sits on the verification side rather than the generation side. Getting a model to produce a confident-sounding brief is easy and mostly solved. Knowing whether each citation is real, whether each holding says what the draft claims it says, and whether the reasoning is grounded rather than merely plausible is the hard part, and it is the part that turns throughput into value instead of exposure. The Microsoft study is a clean demonstration that generation-side gains are real and measurable. It is also a clean demonstration that the profession still has to build the quality measure that sits underneath them.</p><h2>A useful embarrassment about which tool won</h2><p style="text-align: justify;">One more finding deserves attention because of how the authors handle it. On merged-pull-request throughput, Copilot CLI outperformed Claude Code by more than two to one, even though public sentiment in early 2026 generally rated Claude Code as the stronger autonomous agent. The authors do not pretend to resolve this. They offer two hypotheses, that engineers reach for the two tools for different kinds of work, and that Microsoft owns GitHub while it merely buys Claude Code, so the Copilot harness was probably tuned to fit how Microsoft engineers actually work. They also flag their own position openly, since they are Microsoft employees studying a Microsoft-owned tool.</p><p style="text-align: justify;">I appreciate this more than a cleaner result would have earned. It is a reminder that a throughput number measures a tool inside a specific organization with specific incentives and specific integration, and it is not a verdict on model quality in the abstract. Anyone who has deployed the same model in two different firms and watched it succeed in one and stall in the other already knows this. The harness, the integration, and the fit to real workflows often matter more than the raw model.</p><h2>What I am taking into my own rollouts</h2><p style="text-align: justify;">Three things:</p><ol><li><p style="text-align: justify;">Make competent use visible, because adoption travels through people rather than memos;</p></li><li><p style="text-align: justify;">Watch retention instead of sign-ups, because the tool that becomes someone&#8217;s first real habit is worth more than the one everyone tries once;</p></li><li><p style="text-align: justify;">And refuse to let the throughput number stand in for value, because in law the gap between faster and better is not a rounding error, it is the entire professional question.</p></li></ol><p style="text-align: justify;">The engineers in this study merged more code. Whether the profession that adopts these tools produces better work, or merely more of it, is the measure we still have to build. </p><p style="text-align: justify;"><strong>Which side of that line is your firm actually measuring?</strong></p><p style="text-align: justify;"><strong>Reference</strong></p><p>Emerson Murphy-Hill, Jenna Butler, and Alexandra Savelieva, &#8220;<em><a href="https://arxiv.org/abs/2607.01418">Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft&#8217;s Early 2026 Rollout of Claude Code and GitHub Copilot CLI</a></em>,&#8221; Microsoft, https://arxiv.org/abs/2607.01418 [cs.SE], July 1, 2026.</p>]]></content:encoded></item><item><title><![CDATA[Law Already Runs a Red Queen Race]]></title><description><![CDATA[Cambridge and NVIDIA just showed how to let an AI evaluator improve without drifting from ground truth. For legal AI, that mechanism is not a convenience. It is the moat.]]></description><link>https://thelegaltechoverlap.substack.com/p/law-already-runs-a-red-queen-race</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/law-already-runs-a-red-queen-race</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Tue, 07 Jul 2026 16:44:28 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Every general counsel I speak with eventually asks a version of the same question. Can a legal AI actually get better on its own, given that the law so rarely offers one provably correct answer? It is a fair question, and until recently the honest answer was a shrug. A new preprint from a team at the University of Cambridge and NVIDIA, working with Flower Labs, MBZUAI, and Inria, gives me a better one. The paper is &#8220;The Red Queen G&#246;del Machine,&#8221; and although it never mentions law, it is one of the most useful things I have read this year for anyone building AI that has to reason about it.</p><p style="text-align: justify;">Let me give you the practitioner translation.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="5184" height="3456" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3456,&quot;width&quot;:5184,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;jack of diamonds playing card&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="jack of diamonds playing card" title="jack of diamonds playing card" srcset="https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1609437737823-1cbab3aad30a?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxyZWQlMjBxdWVlbnxlbnwwfHx8fDE3ODM0NDE3NDF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@laine23">Laine Cooper</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2 style="text-align: justify;">The assumption that breaks on contact with legal work</h2><p style="text-align: justify;">Most self-improving agents share a hidden premise. They optimize against a fixed scorer, whether that is a benchmark, a unit test, or a labeled dataset that never moves. The agent proposes changes to itself, keeps the variants that raise the score, and discards the rest. This works beautifully for code, because a test either passes or it does not, and the target holds still while the agent chases it.</p><p style="text-align: justify;">That premise quietly fails the moment the task starts to resemble legal work. The standard for a good argument is not fixed. It sharpens as opposing counsel gets sharper and as the bench pushes back. </p><div class="pullquote"><p style="text-align: justify;">Advocacy is adversarial by construction, which means you are never optimizing against a static environment. You are optimizing against something that adapts to you. </p></div><p style="text-align: justify;">The authors name their system after Leigh Van Valen&#8217;s 1973 Red Queen hypothesis, the idea that a species has to keep evolving simply to hold its ground against competitors who are evolving too. Read that description again and tell me it is not a fair account of litigation.</p><p style="text-align: justify;">So the interesting move in this paper is not that the agent improves. It is that the thing judging the agent improves with it, while a much older idea sits underneath. A G&#246;del Machine, in Schmidhuber&#8217;s original sense, rewrites itself whenever it can prove the change is beneficial. This work keeps the self-improvement and swaps the intractable proof for something empirical, then adds the piece everyone else held fixed.</p><h2 style="text-align: justify;">Controlled utility evolution, in plain terms</h2><p style="text-align: justify;">Here is the mechanism, stripped of its notation. The authors let the evaluator co-evolve alongside the agent it scores, but they clamp it with one hard rule. Search is divided into epochs. Within an epoch, one evaluator is frozen and grades everything, so the target holds still long enough for the agent to make real progress. The evaluator is only allowed to change at an epoch boundary, and it can be replaced only if a challenger statistically beats the incumbent on a fixed ground-truth anchor that neither is ever allowed to drift away from. When a replacement happens, the system performs what they call selective erasure. It discards exactly the records that depended on the old evaluator, and nothing else.</p><p style="text-align: justify;">The effect is a progressively stricter judge that stays tethered to reality. The agent faces a harder examiner over time, but the examiner never floats free of the ground truth. The authors show this pays off. On open-ended tasks with no clean benchmark, such as writing scientific papers and proofs, the co-evolved system beat fixed-evaluator baselines while spending fewer tokens. Even on verifiable coding, where a real test suite already exists, adding a cheap evolved code reviewer raised the held-out pass rate to 71.7 percent from the prior state of the art&#8217;s 69.9 percent, and it did so using 1.35 to 1.72 times fewer tokens, because the reviewer is queried once where a coding agent needs many turns.</p><h2>The result that should worry every legal AI buyer</h2><p style="text-align: justify;">Two numbers stopped me.</p><p style="text-align: justify;">The strongest baseline reviewer in their study accepted AI-generated papers at up to 1.91 times the rate it accepted human ones. That is self-preference bias, the well-documented tendency of a language model to favor text that looks like its own output. Then the authors introduced an adversarial objective at an epoch boundary. They took the AI-written papers the old reviewer had waved through, replayed them as a hard pool the next reviewer had to catch, and searched for a judge that was equally tough on machine and human work. The corrected reviewer held roughly 80 percent ground-truth accuracy while treating the two sources alike.</p><p style="text-align: justify;">Translate that into our world. If your legal AI&#8217;s verification layer is itself a language model with a soft spot for fluent, machine-generated prose, it will rubber-stamp a brief that reads well and cites a case that does not exist. The plausible-looking hallucinated citation is not an edge case. It is the predictable output of a lenient judge that likes its own handwriting. The paper is, among other things, a worked example of how to build a verifier that refuses to do that, and of why you should not trust a verifier that has never been tested against its own blind spots.</p><h2 style="text-align: justify;">Is law deterministic? No. Is it unanchored? Also no.</h2><p style="text-align: justify;">This is where I part company with the reflexive objection that legal reasoning is too open-ended to admit ground truth. Law is not deterministic. But it is not floating either. We have binding precedent, enacted statutory text, and the outcome that actually held on appeal. That is the anchor. What evolves is the harder, softer question of how a novel argument lands, which reviewer&#8217;s standard it must satisfy, and how much weight a given line of authority can bear this year.</p><p style="text-align: justify;">The paper&#8217;s structure maps onto something practitioners already know but rarely formalize. Legal ground truth is itself non-stationary, yet it moves in discrete, legible steps rather than continuously. A statute is amended. A supreme court sits in plenary session and settles a split. A precedent is overruled. Those are epoch boundaries. Between them, the anchor holds still. The right architecture for legal AI is therefore neither the one that lets its notion of correctness drift with every new draft, nor the one that freezes the law as of its training cutoff. It is the one that updates the anchor only at these legible moments, revalues what depended on the old rule, and preserves everything that did not. Controlled utility evolution with selective erasure turns out to be a surprisingly good description of what a responsible legal knowledge system has to do when a court of last resort changes the law under it.</p><h2 style="text-align: justify;">The moat is the anchor, and the anchor is maintenance</h2><p style="text-align: justify;">I keep telling anyone who will listen that the verification layer is the durable asset in legal AI, and this paper sharpens why. The generating agent is already commoditizing. Whatever writes the first draft this quarter will be beaten by something cheaper next quarter. The thing that decides whether an output is good, and that refuses to certify a clean-looking citation to a judgment nobody handed down, is the part that does not commoditize.</p><p style="text-align: justify;">The paper also lands its own caution squarely on us, and I want to sit with it rather than skate past it. Its guarantees are epoch-local, not global. It can show that the system improved this epoch against this version of the ground truth, but it cannot promise convergence to some globally correct evaluator, and it says so plainly. It concedes, too, that an evaluator is only as good as its anchor. A weak or biased legal ground truth does not produce a cautious agent. It produces a confident, biased one, which is the worse failure because it is harder to catch.</p><div class="pullquote"><p style="text-align: justify;">That is the whole argument for owning your anchor rather than renting it. A trustworthy, versioned, maintained legal ground truth is not a feature you ship once. It is a discipline you sustain, and it is the asset from which the rest of the system borrows its credibility. </p></div><p style="text-align: justify;">Local guarantees are not a disappointment here either. For anyone who has to defend a system to a regulator or a court, &#8220;we can show this improved against a fixed, auditable ground truth between two dated versions&#8221; is a far more honest claim than &#8220;it converges to correctness,&#8221; and it is the one you can actually stand behind.</p><p style="text-align: justify;">I am curious whether others building in this space read the anchor the same way, or whether you think the generating layer still holds more of the value than I am giving it credit for.</p><p style="text-align: justify;">The paper is <em>The Red Queen G&#246;del Machine: Co-Evolving Agents and Their Evaluators</em>, by Alex Iacob, Andrej Jovanovi&#263;, and William F. Shen with colleagues at the University of Cambridge and NVIDIA, alongside Flower Labs, MBZUAI, and Inria, and it is on arXiv at <a href="https://arxiv.org/abs/2606.26294">arxiv.org/abs/2606.26294</a>. I would rather this be a conversation than a broadcast, so if you build or buy legal AI, tell me in the comments or reply to this email. Which parts of your own verification stack would survive being handed to an evolving evaluator, and which anchor would you never let out of your own hands?</p>]]></content:encoded></item><item><title><![CDATA[The Coordination Cost Was the Business Model]]></title><description><![CDATA[Harvard and Perplexity measured the same work done with a chatbot and with an agent. The speed is the headline; the scope shift is the real story for law firms.]]></description><link>https://thelegaltechoverlap.substack.com/p/the-coordination-cost-was-the-business</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/the-coordination-cost-was-the-business</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Tue, 07 Jul 2026 16:07:55 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Most writing about AI in law still argues about adoption. Should you use it, can you trust it, what happens to junior lawyers. A paper released in June by researchers at Harvard and Perplexity quietly makes that debate feel beside the point, because it stops asking people what they think and instead measures what they actually did.</p><p style="text-align: justify;">The method is the reason to take it seriously. Rather than survey lawyers about their feelings, the authors pulled production data from Perplexity&#8217;s own products and ran a within-person comparison: the same user, issuing a near-identical query, once to a conversational assistant and once to an autonomous agent. They kept ten thousand matched task pairs where the two queries were almost word for word the same. That design strips out the usual confounds. It is not one cohort of enthusiasts measured against a cohort of skeptics. It is the same person, the same task, two tools, which is about as close to a controlled experiment as you get outside a lab.</p><h2>The efficiency story, stated precisely</h2><p style="text-align: justify;">Between the user&#8217;s turns, the agent performed about 26 minutes of autonomous work per session. The assistant did 33 seconds. On matched tasks, average completion time fell from 269 minutes to 36, and estimated cost dropped by 94 percent, with human labor, not model cost, doing almost all of the collapsing.</p><div class="pullquote"><p style="text-align: justify;">On matched tasks, average completion time fell from 269 minutes to 36, and estimated cost dropped by 94 percent.</p></div><p style="text-align: justify;">The counterintuitive result is quality. You might expect autonomy to trade accuracy for speed. It did the opposite. On the study&#8217;s next-turn dissatisfaction signal, the measure of meaningful dissatisfaction, the kind that shows up as corrections, re-asks, and error reports, ran at 1.3 percent for the agent against 2.9 percent for the assistant. That is the 55 percent figure being quoted around the study, and it is worth stating at the right level: the serious complaints roughly halved. Faster and judged better at the same time.</p><p style="text-align: justify;">Those numbers are the headline, and they are real. They are also, I think, the least interesting thing in the paper.</p><h2>The number that matters is about scope</h2><p style="text-align: justify;">The finding that should reorganize how a firm plans is not speed. It is that agent users worked outside their own primary occupation far more often than the same users did with the assistant, 59 percent of the time against 50. A single professional, paired with an agent, started absorbing work that used to be divided across separate specialists, legal and financial and technical, inside one workflow.</p><p style="text-align: justify;">The authors describe this neutrally as a &#8220;<strong>reduction in coordination costs</strong>.&#8221; From where I sit, in banking and insolvency litigation, that phrase names something the profession sells. A large share of what a firm bills is coordination: the deal team, the associate pyramid, the partner who assembles specialists and stitches their work products together. When one person plus an agent can carry a task that previously required three people talking to each other, the coordination is not merely cheaper. It becomes a different unit of production. The study is measuring, at the level of individual queries, the early erosion of a pricing model.</p><p style="text-align: justify;">This cuts in a direction lawyers should sit with, because the overlap runs both ways. Yes, a banking lawyer with an agent can now reach into financial modeling and regulatory checks that once meant pulling in a colleague. But a servicer&#8217;s analyst, a financial advisor, or a restructuring boutique can now reach into work that used to require a lawyer. This newsletter is named for that overlap on purpose. It is an opportunity and a competitive threat wearing the same coat.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="2765" height="3456" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3456,&quot;width&quot;:2765,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;black and white robot toy on red wooden table&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="black and white robot toy on red wooden table" title="black and white robot toy on red wooden table" srcset="https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1620712943543-bcc4688e7485?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxhaSUyMGFnZW50c3xlbnwwfHx8fDE3ODM0MzYzMzd8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@santesson89">Andrea De Santis</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><h2 style="text-align: justify;">Law is inside the sample, not adjacent to it</h2><p style="text-align: justify;">Legal and compliance work already made up 5.5 percent of agent queries in the study, and these were not glorified lookups. Across the sample, agent queries engaged roughly 60 percent more fine-grained, occupation-specific work activities than the same users&#8217; assistant queries, meaning the executional tail of the task rather than the topical headline. Research, analysis, drafting, tool calls, and document delivery, handled end to end instead of one prompt at a time. Finance was the standout on a related measure: agents chained external tool and data calls in about 23 percent of finance sessions against roughly 1 percent for the assistant, which is exactly the kind of multi-source, multi-step assembly that banking work runs on.</p><p style="text-align: justify;">The way you interact changes along with it. With an assistant, your follow-ups mostly clarify and correct, and you are steering a search engine. With an agent, they tilt toward verifying and extending, and you are reviewing a deliverable that already exists and pushing it further. That is the difference between operating a tool and supervising one, and it moves the lawyer&#8217;s job up the stack toward judgment, strategy, and validation.</p><h2 style="text-align: justify;">Where this actually lands in banking and insolvency</h2><p style="text-align: justify;">The payoff is not a smarter answer to a single question. It is the workflow. Large-scale due diligence across an NPL or UTP portfolio, where the work repeats across hundreds of positions but each position is individually checkable. Restructuring scenarios that braid financial modeling, legal exposure, and regulatory review into one analysis. These are precisely the tasks the paper flags as the agent&#8217;s home turf: multi-step, multi-domain, expensive to produce, and relatively easy to verify.</p><p style="text-align: justify;">That last property, easy to verify, is the hinge, and it is where I would push past the paper. The economics work when producing the output is costly but checking it is cheap, and much of legal production has that shape. But not all of it, and the exceptions are the whole point. Confirming that a memo cites real and current holdings is cheap. Judging whether a novel restructuring strategy survives a specific court, or whether a settlement number is right given a portfolio&#8217;s tail risk, is not cheap to verify, because that judgment is the work. The division of labor between lawyer and agent gets drawn along the <strong>verifiability line</strong>. Agents take the production that is costly to make and cheap to check. Lawyers keep the judgment that is costly to make and costly to check, which turns out to be a fair description of what senior legal work has always been.</p><p style="text-align: justify;">There is a caveat the authors are honest about, and it matters more for law than for most fields. Their cost model treats the human&#8217;s delegation cost as roughly the effort of writing the prompt, and they note in a footnote that this probably understates the true fixed cost, because it does not capture the cost of verifying the agent&#8217;s output. In a regulated, liability-heavy domain, verification is not a footnote. A hallucinated citation or a missed carve-out is not a typo, it is exposure. Which means the real fixed cost of delegating legal work sits higher than the study&#8217;s proxy, and the firms that win are the ones that drive that verification cost back down with genuine infrastructure, so that checking stays reliably cheaper than producing. That, not raw model quality, is where the durable advantage lives.</p><h2>The honest limits, and the real question</h2><p style="text-align: justify;">The study covers a 90-day window dominated by early adopters and paying subscribers, and the authors say so plainly. The matched-pair method also captures only the agent tasks that had close assistant equivalents, and many did not, although the ones without a twin tend to be the more complex tasks, where the gains are likely larger rather than smaller. So the direction is not ambiguous even if the exact magnitudes will move.</p><p style="text-align: justify;">Which leaves the question actually worth a legal team&#8217;s time. It is not whether to use AI. It is which of your workflows genuinely involve multi-step execution across legal, financial, and documentary domains, and how you would redesign them around a real division of labor between lawyers and agents, drawn along the line separating what is cheap to verify from what is not. That redesign is the hard part, and it is organizational far more than it is technical.</p><p style="text-align: justify;">The technology is already here. The overlap is the opening. So the honest question is not whether your team will use agents. It is whether you will redraw the work around them before the analyst, the advisor, or the boutique on the other side of the deal does.</p><p style="text-align: justify;">The paper is <em>How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope</em>, by Jeremy Yang of Harvard with Kate Zyskowski, Noah Yonack, and Jerry Ma of Perplexity, and it is on arXiv at <a href="https://arxiv.org/abs/2606.07489">arxiv.org/abs/2606.07489</a>. I would rather this be a conversation than a broadcast, so if you build or buy these systems, tell me in the comments or reply to this email. Which of your workflows would actually survive being handed to an agent, and which would you never let out of your own hands?</p>]]></content:encoded></item><item><title><![CDATA[The Committee, the Courtroom, and the Parrots]]></title><description><![CDATA[Researchers made AI agents argue a case. They didn't win more often, but they solved problems a single model couldn't.]]></description><link>https://thelegaltechoverlap.substack.com/p/the-committee-the-courtroom-and-the</link><guid isPermaLink="false">https://thelegaltechoverlap.substack.com/p/the-committee-the-courtroom-and-the</guid><dc:creator><![CDATA[Marco Rossi]]></dc:creator><pubDate>Mon, 06 Jul 2026 17:14:26 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Almost every demo I have seen for agentic legal AI rests on the same unspoken promise, which is that more agents produce better answers. Give the model a team instead of a single voice, let those voices debate, and the reasoning is supposed to come out sharper. It is an intuitive idea, and as someone who builds these systems I want it to be true. A new paper accepted at the AIDA2J workshop of the International Conference on Artificial Intelligence and Law, held in Singapore this June, put that promise on a bench and tested it properly. The result is more interesting than a clean win or a clean loss, and it is worth sitting with if you are betting your practice, or your product, on multi-agent legal AI.</p><p style="text-align: justify;">The authors, Ludi van Leeuwen and Cor Steging of the University of Groningen together with Tadeusz Zbiegie&#324; of the Jagiellonian University, did something that maps almost exactly onto how litigators think. Rather than asking a model to reason in a single pass, they built architectures in which distinct agents take distinct positions and a structure forces those positions to meet. Because this is precisely the design philosophy behind the adversarial systems I work on, I read the paper twice, and I want to walk you through what it actually shows rather than what a headline would want it to show.</p><h2>What they built</h2><p style="text-align: justify;">The paper compares a plain single-model baseline against three multi-agent designs, and the three designs are worth naming because each encodes a different theory of how good reasoning emerges.</p><ol><li><p style="text-align: justify;">The first is standard <strong>Multi-Agent Deliberation, </strong>which the authors treat as a committee. Three agents each produce an answer, then read each other&#8217;s answers across two rounds and revise, and a majority vote settles the question. Reaching a single verdict this way costs nine model calls rather than one.</p></li><li><p style="text-align: justify;">The second design is the one that should interest any litigator, because it is a courtroom. The authors call it <strong>3-Ply</strong>, and it assigns one agent to argue yes as the plaintiff, one to argue no as the defendant, and a third to sit as an impartial judge who decides on the merits after an opening argument, a counterargument, and a rebuttal. If you have ever sketched a Proponent, an Adversary, and an Arbiter on a whiteboard, you already understand this architecture, since it is the adversarial method formalized into a pipeline of four model calls.</p></li><li><p style="text-align: justify;">The third design is the strangest and, to me, the most thought-provoking. Drawing on recent argumentation research, the authors stage a single expert agent named Alex against a chorus of four <strong>critical &#8220;parrots</strong>,&#8221; each embodying a distinct stance. A Socratic parrot challenges definitions and assumptions, a Cynical parrot tries to undermine the arguments and test their robustness, an Eclectic parrot offers interpretations everyone else missed, and an Aristotelian parrot audits the logic for fallacies. Alex answers, the parrots push back, and Alex is allowed to continue the exchange for up to three rounds before committing. Because Alex decides when the conversation is over, this framework uses a variable number of calls, averaging about 3.48 per question.</p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="4000" height="6000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:6000,&quot;width&quot;:4000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A person adjusts another's judicial robe and wig.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A person adjusts another's judicial robe and wig." title="A person adjusts another's judicial robe and wig." srcset="https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1764592620951-316407a80f42?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0NHx8Y291cnRyb29tfGVufDB8fHx8MTc4MzM3OTg2OXww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@spliff_dj_joe">Dwayne joe</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p style="text-align: justify;">To test these designs the researchers assembled five benchmarks, four of them legal and one purely logical. The legal set spans law school examinations, Japanese civil law bar questions, United States federal tax reasoning, and privacy policy interpretation, and the fifth benchmark tests logical reasoning drawn from a civil service exam. They sampled 250 balanced questions from each, which gives 1,250 binary yes-or-no problems in total, and they ran everything on a smaller commercial model, GPT-5-mini, with two-shot prompting where the examples were selected by a standard retrieval method. The choice of a smaller model was a budget decision, and the authors are candid that larger models might behave differently, which is a caveat I will return to.</p><h2>The finding that should reshape your priors</h2><p style="text-align: justify;">Here is the headline, stated plainly, because the size of a claim should match the size of its evidence. Across all five benchmarks, none of the three multi-agent frameworks meaningfully outperformed the single-model baseline. The average F1 scores sat within<strong> one and a half points </strong>of each other, and no architecture pulled clearly ahead. If your entire thesis for multi-agent legal AI is that it produces higher accuracy, this paper does not support you.</p><p style="text-align: justify;">And yet stopping there would miss the real result. Although the frameworks did not score higher, they answered differently in a way that is statistically unmistakable, since every multi-agent design diverged from the baseline&#8217;s decision pattern at high significance. Roughly <strong>seven to ten percent</strong> of<strong> </strong>questions received a different answer under a multi-agent framework than under the single model, and in about half of those disagreements the multi-agent system was right where the baseline was wrong. Depending on the framework, between<strong> forty-three and fifty-two percent</strong> of the cases where they diverged were cases the multi-agent design rescued.</p><p style="text-align: justify;">The most striking detail hides in the privacy dataset, where the authors found an asymmetry that I have not been able to stop thinking about. </p><div class="pullquote"><p>Every question the baseline solved correctly was also solved by at least one multi-agent framework, but the reverse was not true, because many questions the baseline failed were recovered by the agents. </p></div><p style="text-align: justify;">In that slice of the data, the multi-agent approach did not trade one set of errors for another of equal size. It strictly expanded what the system could handle. The authors are careful to note this reflects the combined behavior of several frameworks rather than a single one, so it is a hint rather than a law, but it is exactly the kind of hint that tells you where to dig.</p><p style="text-align: justify;">This is the reframing that matters for practitioners. The question is not &#8220;does multi-agent beat a single model on average,&#8221; because on these benchmarks it does not. The better question is &#8220;does multi-agent reach reasoning that a single model cannot,&#8221; and here the evidence says yes, at least for a specific and identifiable class of problems.</p><h2>Where the agents earn their keep, and where they do not</h2><p style="text-align: justify;">The qualitative analysis is where the paper stops being a scoreboard and starts being useful. The authors show that the multi-agent frameworks pull ahead precisely when a legal clause is ambiguous or admits more than one reading, which is to say in the situations lawyers are actually paid to resolve. Their worked example turns on whether a privacy clause covering &#8220;geolocation data&#8221; also covers WiFi-derived location that the clause never names explicitly. The single model latched onto the literal absence of the word WiFi and answered wrong, whereas the courtroom and the parrots surfaced the tension between a literal and a purposive reading, argued it out, and arrived at the correct, more lawyerly interpretation. This is the deliberative dividend, and it appears exactly where a single narrative is most likely to suffer tunnel vision.</p><p style="text-align: justify;">The uncomfortable findings deserve equal airtime, because a newsletter that only reported the flattering half would be the marketing I am trying to avoid. Hallucination did not go away. In the same example, the committee framework reached the right answer only after one of its agents invented a justification about cell phone data that had nothing to do with the question, which means the structure had to first clean up a mess it created. The authors cite the argument that hallucination is a structural feature of these models rather than a bug we will patch, and their own results are consistent with that sobering view.</p><p style="text-align: justify;">Two more results puncture common assumptions. First, adding reflection rounds to the committee barely moved the score, since three agents voting once landed within a point of the same agents deliberating twice, which suggests that on these tasks much of the value came from ensembling rather than from the deliberation itself. Second, all of this costs real money, because the committee burns nine model calls and the courtroom four for every single call the baseline uses, and that spend did not buy better raw accuracy. If you are deploying at scale, you are paying a multiple for a different distribution of errors, not for fewer of them.</p><h2>What I take from this</h2><p style="text-align: justify;">I build adversarial and multi-role systems for legal work, so I read this paper as a colleague, not a spectator, and my honest reading is that it validates the architecture while demolishing the marketing. The value of putting a plaintiff, a defendant, and a judge inside the machine is not that it is smarter in aggregate. The value is that it makes the reasoning explicit, exposes competing interpretations that a single pass would bury, and rescues a meaningful set of hard, ambiguous cases that a monolithic model gets confidently wrong. For a domain where the reasoning behind an answer matters as much as the answer, that explicitness is not a side effect. It is the product.</p><p style="text-align: justify;">There is a deeper point underneath the numbers, and the paper gestures at it in its discussion. These systems generate arguments, but nobody checks whether those arguments are actually valid, and a model can produce a fluent, well-structured, and logically broken chain of reasoning without noticing. Deliberation makes the reasoning visible, and yet visibility is not verification. The agents argue, but the argument is never audited for soundness. That gap, between a system that produces arguments and a system that can be trusted to check them, is where I think the serious work of the next few years lives, and it is a gap that neither more agents nor more rounds will close on their own.</p><p style="text-align: justify;">If you are evaluating agentic legal AI, the practical takeaway is to stop asking vendors whether their multi-agent system scores higher, and start asking which problems it solves that a single model cannot, and at what cost per answer. Those are the questions this paper actually equips you to ask.</p><p style="text-align: justify;">I plan to revisit the primary source as more replications come out, and I would rather this be a conversation than a broadcast. The paper is titled <em>Investigating Multi-Agent Deliberation in Law</em>, and it is on <a href="https://arxiv.org/abs/2606.30906">arXiv at </a><em><a href="https://arxiv.org/abs/2606.30906">2606.30906</a>.</em> If you build or buy these systems, I want to hear where your experience confirms or contradicts what the authors found, so tell me in the comments or reply to this email. </p><p style="text-align: justify;"><strong>What have you seen multi-agent reasoning do that a single model could not?</strong></p>]]></content:encoded></item></channel></rss>