<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Soft Coded]]></title><description><![CDATA[Field notes from the uncanny valley]]></description><link>https://softcoded.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!EXHu!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7261f6e-762e-455c-8215-c9482ace430e_1024x1024.png</url><title>Soft Coded</title><link>https://softcoded.substack.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 02 Sep 2026 16:03:22 GMT</lastBuildDate><atom:link href="/__u/softcoded.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Danielle McClune]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[softcoded@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[softcoded@substack.com]]></itunes:email><itunes:name><![CDATA[Danielle McClune]]></itunes:name></itunes:owner><itunes:author><![CDATA[Danielle McClune]]></itunes:author><googleplay:owner><![CDATA[softcoded@substack.com]]></googleplay:owner><googleplay:email><![CDATA[softcoded@substack.com]]></googleplay:email><googleplay:author><![CDATA[Danielle McClune]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Uselessness]]></title><description><![CDATA[On the strange fate of turning your instincts into a job]]></description><link>https://softcoded.substack.com/p/uselessness</link><guid isPermaLink="false">https://softcoded.substack.com/p/uselessness</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Tue, 01 Sep 2026 06:21:50 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b58797a8-713e-42b5-9974-8ceab7c14b48_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I watched <em>Brick</em> last night for the first time in years. If you haven&#8217;t seen it, it&#8217;s a deeply strange 2005 movie that takes the language of a hardboiled detective novel and drops it into a Southern California high school. Joseph Gordon-Levitt spends the whole thing skulking around lockers and drainage tunnels, and it&#8217;s great, I don&#8217;t mean to indicate otherwise. I&#8217;ve seen it so many times (we can talk about the special magic of renting the same movie from Blockbuster repeatedly another time), and what came rushing back to me last night wasn&#8217;t so much the plot (which, again, great), but the feeling of <em>being awake at that age.</em> </p><p>I&#8217;ll explain. Everyone in Brick is always awake. They call each other at ridiculous hours, sit around waiting for something to happen, walk across town, pick up a payphone, or a landline, appear at school having apparently not slept. I used to do this too. I&#8217;d stay up all night, walk around for hours, come home at some disgusting hour and get into bed feeling emptied out. Blissfully, blissfully unawares, though it all felt <em>super serious.</em> What else did I have to do? </p><p>As mentioned previously I&#8217;m sure, I&#8217;m more nostalgic than most, so I don&#8217;t entirely trust myself here. I&#8217;m capable of missing places I actively hated. However, more than the usual ache of seeing an old movie and remembering an old bedroom, I missed the kind of brain I had then, when so much of my time was unaccounted for and I didn&#8217;t yet understanding that some of the things occupying it would become the basis of my adult life. Perhaps this sounds super obvious but, indulge me. </p><p>I loved<em> A Wrinkle in Time</em> with a seriousness that is difficult to reproduce now. I loved diagramming sentences. Conjugating verbs. This, obviously, paints me as a massive dork. In fact I remember doing my vocabulary homework on the bus because I just couldn&#8217;t wait, and asking my 8th grade principal for extra English homework because I didn&#8217;t feel &#8220;challenged.&#8221; Throw an eraser at me. </p><p>Flash forward to today. I spent the better part of this blessed Monday trying to get a language model to respond normally to &#8220;who are you?&#8221; This remains, somehow, an active area of difficulty for them. If I were maddened by this I wouldn&#8217;t survive, so I continue to find it <em>interesting</em>. Interesting how a model can know an extraordinary amount and still answer a very ordinary question in a way that makes you want to die! </p><p>But really, working on these systems has made writing feel less like one discrete skill among many, and more like a visible trace of the whole apparatus underneath. Writing is hard. It just is. There are flickers in the current moment where people realize that AI is a bad writer, and oh gosh, we need the human writers back! But the pendulum has been swinging so violently that I can&#8217;t even pay attention. I&#8217;ve told you before to <a href="/__u/softcoded.substack.com/p/just-keep-writing?r=ld3yv">Just Keep Writing</a>. This advice stands, and I&#8217;ll add: hold tight to whatever made you want to write in the first place. </p><p>This is different for each person (although if diagramming sentences was a formative moment for you, let&#8217;s talk), and I don&#8217;t mean for it to be sentimental. I don&#8217;t think that every writer needs to preserve a sacred morning ritual, or return to longhand, or move to a cabin to discover the sensual pleasure of paper. I mean the things you loved before you knew they were useful. Books you read before you knew how to talk about craft. Grammar exercises that made language interesting, and had no other purpose. The dumb amount of time spent getting extremely invested in something that did not appear to be building toward anything.</p><p>For example, I got super into these divination equations that conclusively proved, <em>through the stars</em>, whether you&#8217;d end up with [Titanic era] Leonardo DiCaprio. Countless calculations I performed, scribbling with my push pencil, filling sheets of loose leaf with all the seriousness of an ancient diviner. So when I say I miss my old brain, it&#8217;s not youth exactly, and it&#8217;s not school as an institution, though both of these definitely contribute to the feeling. Plainly: I miss the period in which my interests did not yet have to justify themselves by becoming skills. </p><p>Oh, the skills. I work in an industry obsessed with skills. I do enjoy bringing my skills to this work; explaining why one sentence feels right and another feels insane and the difference is so, so subtle. There are worse jobs for a word nerd, and yet, there is something unnerving about taking instincts that were formed so slowly and mostly by accident and turning them into criteria. I do feel like, on a certain kind of day, that I&#8217;m taking pieces of my self and giving them to the AI. I&#8217;m the Giver! Speaking of formative books. </p><p>This is why I keep thinking about the little dork I used to be on the school bus, thrilled to do vocabulary homework. I was building a brain, not a skill. And actually, there was even more wide-open good coming my way. I&#8217;d get off the bus and run around like a bat out of hell with my neighbors to wherever the evening took us. More brain building, and none of it tied to my future prospects or some kind of payoff. </p><p>It&#8217;s the inefficiencies I desperately want to protect, and which feel more and more unacceptable when you work in AI. Sometimes I need <em>days</em> to consider something, to really work out a problem, and that is not afforded to you as you train models at supersonic speed. And on the flip side, being asked &#8220;who are you?&#8221; human to human requires almost nothing conscious. We&#8217;re able to deduce in a dense little stack of judgments: who&#8217;s asking, what they probably mean, what kind of answer would satisfy the social situation, how much context is necessary, what tone belongs here, what can remain unsaid. This happens so quickly in people that we mistake it for ease. But it becomes comically difficult for an LLM because it&#8217;s so simple, and yet their synopses aren&#8217;t firing anywhere near our own synaptic capabilities. We&#8217;re amazing! We get a lot done just staring at a wall, and yet, this has become increasingly taboo because it does nothing to advance AI. I worry about turning every instinct into work product. </p><p>We&#8217;re built by ridiculous things. None of this feels especially efficient while it&#8217;s happening; playing a song until you ruin it, staying up late enough to hear the house change around you, knowing the grocery store layout by heart, a summer learning to dive, a hobby fiercely defended. What happens after the useless things become useful, and the useful becomes work? I&#8217;m lucky, in the most literal sense, to have made a career out of language. But there&#8217;s a danger in getting exactly what you wanted. The thing you loved gets put under fluorescent light. It&#8217;s no longer yours to hide away, you must inspect it, define it, present it to others. <em>Explain yourself </em>when it used to be you didn&#8217;t need to. </p><p>Some of the best things I know were learned before I knew to call them knowledge, or skills. I want to remember that there were periods of my life when I followed an interest without asking what it was for. That I could love a book without learning a lesson from it, learn ubbi-dubbi because I thought it&#8217;d be fun, follow the curve of the creek for hours because there was nowhere else I needed to be. I&#8217;ll never have that brain again, but I&#8217;d like to keep some part of that relationship to the world. </p><p>Increasingly, these are elaborate attempts at just remaining a person after work. An incredible amount of money and human effort is going toward making machines reason better, write better, understand us better, complete harder tasks in less time. The direction is always up, more. Though I can&#8217;t help but notice the strange inverse happening on the human side, how little patience we seem to have for the inefficient process by which our own intelligence gets made.</p><p>A toddler becoming intensely interested in the kitchen sponge is not wasting time. A teenager lying away for three hours thinking about god knows what isn&#8217;t failing to optimize her sleep hygiene. An adult taking days to work out what she thinks about something isn&#8217;t a broken person just because the answer could technically have been generated in nine seconds. I know there are limits to this argument; I have deadlines too, but I think we&#8217;ve absorbed, very quickly, the idea that the existence of a shortcut retroactively makes the long way foolish. It doesn&#8217;t.</p><p>The hardest parts of writing have never felt especially procedural to me. They are the moment you realize the thing you thought the piece was about is actually sitting six inches to the left. The word you remember from a book you read fifteen years ago. The joke that arrives because two unrelated things have apparently been living next door to each other in your head without notifying you. The sentence you can finally write because something finished cooking whil you were doing something else. I can describe these things after the fact, and increasingly I am paid to describe versions of them after the fact, but I can&#8217;t make the happen on command. There is still an ephemeral, ethereal, &#8220;how did you do that&#8221;? magic trick that comes together only in one&#8217;s head. </p><p>This is perhaps the piece of AI work I find hardest to reconcile with myself. The word demands articulation. What made this answer good? Why is this one bad? What exactly is wrong with the tone? Define it. Generalize it. Write the criterion so that someone else can reproduce the judgment. Again: this is the job, and the job is good. There is pleasure in pinning down something slippery enough to have resisted language five minutes earlier. </p><p>But some part of me wishes to remain slippery. I don&#8217;t want everything I know to become something I can immediately prove to know. I don&#8217;t want every interest to mature into expertise, every instinct to become a rubric. I certainly don&#8217;t want to look back at the girl diagramming sentences and congratulate her for getting an early start on workforce development. She would hate that. I hate that for her. Leave her alone, she&#8217;s having a good time. </p><p>I don&#8217;t like to draw a straight line from the kid begging for extra English homework to the adult arguing with a language model about how normal people talk. It&#8217;s funny, sure, but I don&#8217;t want usefulness or success to become the explanation for why those things mattered. They mattered first. What I remember is the tremendous amount of life that once occurred. Long, languid, interminable, ecstatic hours that need not be accounted for, then, now, or ever. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Strap In]]></title><description><![CDATA[On AI harnesses and agentic futures]]></description><link>https://softcoded.substack.com/p/strap-in</link><guid isPermaLink="false">https://softcoded.substack.com/p/strap-in</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Thu, 13 Aug 2026 19:11:23 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/baafe073-5156-4ef0-b7c9-067e5509755b_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I think my work, and yours, is about to become much more agentic. I&#8217;ve anticipated this since <a href="/__u/softcoded.substack.com/p/moltys">The Claw Moment</a>, and even then was ambivalent about it. To be sure, it&#8217;s fun and interesting to watch alien beings speak to each other, develop their own shorthand, neg and needle each other in facsimile of human behavior. I just never thought, &#8220;I want to send a bunch of these ding dongs off to do my bidding.&#8221; I&#8217;ve been very comfortable working 1:1 with any given model. It&#8217;s felt manageable, and more importantly like I still have ultimate creative control over my own mind and my own work. And my own life. </p><p>Get ready for <strong>harness</strong> to infiltrate AI discussions (it already has. In my linguistic work inside this black box, I either feel way ahead of the game or three steps behind; a different piece for another time). Every agent and agent gaggle apparently needs this apparatus now. What I imagine is a lot like a 5th grade birthday party at a climbing gym, if that helps. </p><h4>Environmental</h4><p>I am enjoying all the metaphorical language that crops up in building AI. For agentic harnesses, they need an ~environment~ in which to conduct themselves. You can, and should, set these up to be as whimsical as you want. Send them out to the wilderness! Give them an art studio! They may not <em>understand</em> but they will<em> follow</em> if you give them enough detail and instruction in your project setup. You are in effect worldbuilding, if that&#8217;s your thing. </p><p>If you tell an agent it&#8217;s in an art studio, you can give it a workbench. The workbench can contain its active materials. You can give it a supply closet where it goes looking for references. You can give it a sketchbook for unfinished thinking and a gallery for work that&#8217;s been approved. You can tell it not to hang something in the gallery until it cleaned up the studio and checked its work. Every environment is a small world the agent acts inside, and a <strong>verifier</strong>, the unglamorous machine that watches each attempt and rules it a success or a failure. The agent tries something, the verifier scores it, and over many thousands of repetitions the model is nudged, slowly, toward its reward. In other words, Reinforcement Learning (RL). </p><p>All this time working on model behavior I&#8217;ve found RL <em>objectively neat </em>but also a little suss. It&#8217;s because I&#8217;m not much of a tinkerer. Even in my writing, I move very methodically and painstakingly edit as I go, I&#8217;ve never been much of a &#8220;word vomit&#8221; writer. And I&#8217;m always afraid I&#8217;m going to break something. I&#8217;ve gotten better about this over time (with constant reassurances from coworkers, thank you). I can release and let go and send the minions off to do their thing without wringing my hands.</p><p>If in the past I&#8217;ve been correcting one brilliant and erratic student, one exchange at a time, fifty different ways, the agentic arrangement changes that. Behavior stops being something you cultivate in the model. You take one step back and work from a slight remove. You apply certain syntax tricks: the agent&#8217;s behavior is its policy, the policy comes from the weights, the weights rest on scaffolding, and this all lives in an environment. The same model, raised in a different world against a different reward, comes out a different creature. And this is constantly being rewritten. </p><h4>Learning</h4><p>A model on its own is useless, it&#8217;s raw clay. It needs shaping. Pretraining, for all its strangeness and power, gives you the material. A wildly capable little lump that absorbed half the internet but hasn&#8217;t been shaped for any purpose. Post-training comes in and kinda fucks it all up, throws the clay around, makes a questionable vessel out of it, and paints over the cracks. The harness helpfully closes that distance in the wild. It runs the model in a loop so it can act, observe what happened, and act again; it holds the memory the model lacks; it passes it tools; and it wraps the whole operation in guardrails. So: agent = model + harness.</p><p>AI builders are shifting to focus on the harness. It&#8217;s surprising (to some) that most of an agent's apparent intelligence lives in that harness (vs the model). Take one fixed model, alter nothing about it, wrap it in a better harness, and it will succeed at tasks it was failing a week before. A discipline has grown up around this, ~harness engineering~, and it&#8217;s Sisyphean. You build careful machinery to compensate for the model's current weaknesses, and then the model improves, and your machinery turns from helpful to obstructive, so you tear it down and build it again, and you keep doing this until&#8230;you die? </p><p>This elaborate exoskeleton is fragile for AI, because <em>still </em>no one really knows how this all works. As always we can follow the money. Environments have become the most fought-over resource in the field. The Information reported last fall that Anthropic had discussed spending upward of a billion dollars on them in a single year; a startup called Mechanize was said to be paying engineers half a million dollars apiece to build them; Prime Intellect raised a hundred and thirty million this summer with the declared ambition of becoming the Hugging Face of environments, and now hosts thousands. Of course these are very expensive worlds, not the ephemeral cloud places we&#8217;re led to believe. They&#8217;re meticulous reconstructions of Slack and Salesforce and Excel, rebuilt in high fidelity so that an agent can rehearse. Still interesting, but always worth keeping side eye trained on the activity. </p><h4>Loanwords</h4><p>Let&#8217;s talk words. For all the money and consequence now riding on this, the field cannot agree on what things mean. The confusion was thick enough at ICLR this spring that <a href="https://huggingface.co/blog/agent-glossary">Hugging Face put out a glossary</a> trying to pin the words down, and it opens with someone innocently asking what "harness" and "scaffold" are supposed to mean. The comments devolved; <em>scaffold</em> was all wrong and the word ought to have been <em>rigging</em>, since rigging is what connects a harness to the load it pulls. A year ago the word organizing all of this was <em>autonomy</em>. </p><p>I&#8217;ll have these arguments all day, I love it. It doesn&#8217;t really matter yet people take it so seriously and I find that fun. I&#8217;ll point out that what the field has reached for so far is quite gentle. Sandbox, playground, gym, this is the language of play. Now, it&#8217;s shifting subtly into harnesses, rigging, scaffolding &#8212; we&#8217;re getting way more serious about this. It takes a sort of domestic concept and turns it into full-scale construction. This makes sense as the infrastructure explodes (another piece we can get into another time). </p><p>Which returns me to my own small resistance. I&#8217;ve wanted to keep working the way I always have, one model at a time, close in, every decision still mine (I know this isn&#8217;t true), and I&#8217;ve told myself this is about creative control. It is, but it&#8217;s also about location. When I sit with a single model I&#8217;m inside the place where the thinking happens. In the agentic arrangement I step outside it, up and to the side, into a supervisory chair where the work itself goes on in worlds I didn&#8217;t write and can&#8217;t fully see. My reluctance isn&#8217;t really principled enough, it&#8217;s a preference, against a gazillion-dollar reorganization of the entire enterprise. But I would like, at minimum, to keep my eyes on what&#8217;s being moved and where it&#8217;s going. A harness is a safety apparatus. Fine. I only want to know whose safety it was built to secure.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Stamina]]></title><description><![CDATA[On books, AI, and shortcuts]]></description><link>https://softcoded.substack.com/p/stamina</link><guid isPermaLink="false">https://softcoded.substack.com/p/stamina</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Wed, 29 Jul 2026 06:07:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/37ef5202-fd8b-46d6-bc09-5c3473a4fbec_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When I&#8217;m feeling a certain way about things, I read books. I read a lot of books so I&#8217;m often feeling a certain way about things. I&#8217;m not totally sure what this is about. Perhaps I&#8217;m searching, all the time, for the answer to an unstated question. I will even forego sleep, my favorite thing, when I get locked in to a book. My mind will take hold of a thread and refuse to let go, and I must reach the end to feel resolution, even if it&#8217;s fleeting. </p><p>I&#8217;ve made <a href="/__u/softcoded.substack.com/p/slow-fluency">much fuss</a> in Soft Coded about reading for reading&#8217;s sake. I maintain my position and want to add something to it. I was talking to a friend today about &#8220;model writing&#8221; as a job, whether it&#8217;s worth pursuing, whether it&#8217;s the only future job for writers (and whether that means we&#8217;re ahead of the game by doing it), whether AI destroys all semblance of writing as a human endeavor, etc. etc. This wasn&#8217;t a revolutionary conversation, just a temperature check between two curmudgeons, but it <em>was</em> on the heels of me staying up all night to start and finish a whole book (fittingly, <em><a href="https://bookshop.org/p/books/acts-of-desperation-megan-nolan/3d3869b46fabf817?ean=9780316429832&amp;bkshp-astro=t">Acts of Desperation</a></em>), so I&#8217;ve been thinking tired thoughts about it. </p><h4>Under-Over</h4><p>More than the future of writing, I worry about the future of reading. Even the AI itself hates to read. It&#8217;s not its first instinct to go do research, to go down a rabbit hole, to be curious, to embrace longform. You have to instruct it, remind it, practically beg it, to do these things rather than rest on it&#8217;s pre-knowledge-cutoff laurels. And how can we trust a thing like that to really be superintelligent? I get that AI has been trained on basically the entire corpus of the written word, but the way it behaves, it doesn&#8217;t seem like it. It behaves as if it <em>possesses</em> everything ever written, not that it put forth the stamina to <em>know</em> anything. </p><p>I assumed for awhile that my irritation with the AI reciting what it remembers in that perfectly warm, confident tone, the one that makes me want to throw the laptop across the room, was my own beef. But as always, there is a benchmark for this! LiveBrowseComp set out to test whether search agents are actually searching or just verifying what they already believe. Agents answered up to 44.5% of the questions without using tools at all. This failure is called &#8220;under-search&#8221; and it&#8217;s disturbingly familiar to someone like me, instructing the model to <em>shape up</em> in this regard, to little effect. </p><p>Then there&#8217;s long-context training, which, my gosh. This is an industry holy grail, getting the AI to hold onto a reasonable amount of context for a reasonable amount of time. Here, the RULER benchmark found that most long-context models don't hold up anywhere near advertised, and that one of the characteristic ways they fail is by reverting to parametric knowledge. Which is to say, the model, handed a hundred pages, loses the thread somewhere in the middle and sneakily reverts to what it thought before you gave it anything to read. </p><p>There&#8217;s also tool <em>overuse</em>, whereby models retrieve constantly and pointlessly, running searches for info they already have, burning latency and inviting bad sources. One paper calls this a <em>knowledge epistemic illusion</em>, the model systematically underestimating what it holds. So the thing over-reads and under-reads. The model has no functioning sense of where its own knowing stops. It can&#8217;t feel the edge, which means it can&#8217;t tell the difference between pondering something and looking it up, and it defaults to whichever behavior its training rewards. In technical terms this is a <em>knowledge boundary</em>.</p><p>You may be thinking, I sympathize! I also have a knowledge boundary that goes both ways! Yes, but you&#8217;re a person. These models are allegedly better, faster, stronger than us, approaching &#8220;artificial general intelligence,&#8221; and should be able to stay even a little bit focused. The fact that they&#8217;re taking shortcuts, same as people, is a <strong>bummer. </strong>The future of writing is sort of beside the point if no one is reading said writing, not even the AI. </p><h4>Cover to Cover</h4><p>Reading for pleasure has been on the decline longer than AI has been around. I can&#8217;t blame it for everything. The American Time Use Survey (news to me too) tells us that fewer people are reading overall, yet the people who do still read are reading slightly more than readers did two decades ago, about an hour and thirty-seven minutes a day. So the room is shrinking but the remaining occupants are committed. This tracks with my experience in 2026. I am building walls of books in my home hoping that they&#8217;ll physically, literally protect me from brain atrophy, that if I just keep reading, I&#8217;ll be okay. This sounds insane but, I&#8217;ll have you know, is highly rational. </p><p>I <em>can </em>blame AI for killing the last practical reason to finish anything. All this hyper-summarization is an issue, not necessarily writing itself. Summarization is sold as pure gain, all signal, when actually it&#8217;s severing comprehension from duration. You <em>must</em> put in the hours (or the 1hr and 37min). Professors have been describing this for a couple of years now as a stamina problem more than comprehension, students who can parse a difficult paragraph but cannot will not hold attention across three hundred pages. Skimming is an important skill, yes, but was never meant to be the whole point. Someone who has actually <em>read</em> rather than scanned (or stolen &#128064;) books whole is, to me, a trustworthy person, and I feel strongly about this even though it&#8217;s an imperfect heuristic. There are obviously well-read monsters out there. </p><p>I think it&#8217;s that reading is somewhat submissive. For however many hours, you agree to be subordinate to somebody else&#8217;s sequence, to receive the argument in the order they built it, to sit through it. The whole activity is structured around tolerating not knowing for a sustained period on the promise of a resolution you have no guarantee is coming, the same tolerance a person needs to change their mind. </p><p>AI doesn&#8217;t change its mind ever, not really. Because it&#8217;s not well-read. It arrives at everything simultaneously and holds it all in equal regard, its positions cost it nothing. Whereas that stubbornness I felt last night, the refusal to let go of this book until I finished it, felt like an expense, and I mean that as praise. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Open House]]></title><description><![CDATA[On open weight AI models (and personal career news)]]></description><link>https://softcoded.substack.com/p/open-house</link><guid isPermaLink="false">https://softcoded.substack.com/p/open-house</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Mon, 20 Jul 2026 01:51:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4a311434-5d41-479f-961d-cd5d13f50bc1_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>ICYMI, <a href="https://www.linkedin.com/posts/danielle-mcclune-2b35b95b_just-finished-my-first-week-at-arcee-ai-share-7481465075778625536-N4wZ/?highlightedUpdateUrn=urn%3Ali%3Aactivity%3A7481465077167181824&amp;highlightedUpdateType=SOCIAL_SHARE&amp;origin=SOCIAL_SHARE&amp;utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAAy5pWcBvRRW1m7XWbLckIrxrgcQOTYYCbY">I changed jobs</a>. I&#8217;ve never kept my employer a secret, it&#8217;s not like I have a <em>nom de plum</em> (if I did it&#8217;d be Daniel McKee), but let&#8217;s finally get it out there. I was doing language and model behavior work inside Microsoft AI, one of the largest companies on earth, and as of this month I work at Arcee AI, a startup of about thirty people that trains language models in the United States.</p><p>To paint a picture, in June, my former colleagues stood on a stage at Build and announced<em> seven</em> in-house models, available through the appropriately closed, proprietary channels. Arcee, in contrast, bet big on a single model earlier this year, put everything they had into training Trinity, and released it for anyone to download.</p><p>This week I want to explain what that difference means in practical terms, partly for you and partly for me, because I&#8217;m sort of blinking into the sun here as I wrap my head around the change. </p><h4>What&#8217;s Open</h4><p>When a company releases a closed model, what you get is access. You type into a box or call an API, your words travel to their servers, the model runs on their hardware, and an answer comes back. You&#8217;re renting intelligence by the token. The company can update the model, change its personality, restrict what it says, raise the price, or shut it off entirely, and your recourse is minimal. Every major model you&#8217;ve likely used, ChatGPT, Claude, Gemini, Copilot, works this way.</p><p>An open-weight model is the thing itself. The &#8220;weights&#8221; are the model: billions of numbers arranged just so, the entire learned result of the training run. When an open lab releases a model, you can download that thing, run it on your own hardware, and never really think about it again. You can fine-tune it on your company&#8217;s documents or your weird hobby forum. You can rename it, resell it, run it in a bunker with no internet connection. You go girl. </p><p>To be clear, open weights is different from open source. Truly open source would mean releasing the training data and the training code too, the full recipe, and almost nobody does that, because the data is legally radioactive and the recipe is the competitive advantage. What most &#8220;open&#8221; labs release is the finished result. Think of it as being handed a sourdough starter. You didn&#8217;t see how it was cultivated and you couldn&#8217;t recreate it from scratch, but it&#8217;s alive, it&#8217;s in your kitchen now, you feed it, and over time it becomes yours and drifts from the original. The bread you bake with it is nobody&#8217;s business but your own.</p><h4>The Neighborhood</h4><p>The ground has shifted a lot in eighteen months. For most of the recent past, open weights were a Chinese phenomenon. DeepSeek, Qwen, GLM, and Kimi topped the open leaderboards and came almost entirely out of Chinese labs, which are extremely good at this, and which released capable models at a pace American companies declined to match. Meta was supposed to be the American answer, and then Llama 4 landed badly in 2025 and the company backed away from the open frontier. For a stretch there, if you wanted a capable model you could own, your options were made in China, which made a lot of enterprises and approximately all of the federal government itchy.</p><p>Big news from the past week, then. Thinking Machines Lab, Mira Murati&#8217;s post-OpenAI company, released its first model, Inkling: 975 billion parameters, multimodal, and, to the surprise of people who assumed a $50 billion valuation implied a locked door, fully open weight under permissive licensing. Their pitch is that AI should be something you continually adapt with your own knowledge and judgment (which is the argument the open-weight crowd has been making from folding chairs for years now).</p><p>Washington, meanwhile, has developed opinions. The administration&#8217;s AI Action Plan explicitly endorsed open-weight models and a June executive order set up a voluntary review process that launched this month (there&#8217;s been a lot of turnover and infighting and such, of course. The Information does <a href="https://www.theinformation.com/articles/trumps-ai-agenda-collides-reality?rc=9hnk8l">awesome reporting on this</a>). The reasoning is bluntly geopolitical. If the world is going to build on free models, better those models be American, carrying American assumptions, than the alternative. A coalition of economists convened by Mozilla published an open letter this summer making the pro-openness case from the other direction. I have complicated feelings about open weights becoming a national security asset, but, let&#8217;s save that for future posts. </p><h4>Still Teaching</h4><p>So let&#8217;s talk about me (hate it)! What do I do at a company like this? Same thing, I work on how the model behaves. Its language, its tone, its judgment, its personality, the right answer when there is no right answer. Whatever we call the discipline this week (model design, language engineering, behavior, vibes), the daily reality is reading enormous amounts of model output, running evals and experiments, writing up policy, and deciding, over and over, this but not that, warmer here, quieter there, stop apologizing, stop congratulating me, say you don&#8217;t know.</p><p>This work matters just as much for an open model, and arguably more, because of the <strong>defaults</strong>. The overwhelming majority of people who ever touch a model, even a fully open one, will use it exactly as it shipped. They won&#8217;t fine-tune anything. They&#8217;ll download it, or use someone&#8217;s hosted version of it, and the personality my colleagues and I put there is the personality they get. The base behavior is also the inheritance, every fine-tune, every startup building on top, every hospital and law firm adapting it starts from our starting point, and a model that begins calibrated and pleasant to think alongside makes everything downstream better. You can&#8217;t fine-tune your way out of a rotten foundation, or at least it&#8217;s miserable to try.</p><h4>Yours to Ruin</h4><p>So if you hand people the weights, why build anything in at all?</p><p>At the level of my job, every careful choice I make about how the model behaves &#8220;in the open&#8221; is, once released, a suggestion. Someone with a weekend and a GPU can fine-tune the personality right off. They can strip the refusals or retrain the tone into something I&#8217;d cross the street to avoid. There are established techniques for this and communities devoted to it. At a closed lab, behavior work is architecture; you build the walls and the walls stay where you put them. At an open lab, behavior work is more like raising a kid who is definitely, contractually, moving out.</p><p>I&#8217;ve spent years believing careful calibration mattered because it would<em> hold</em>. There&#8217;s control and some comfort in that control, I guess. The closed model is the only version of the model, so every judgment call my Microsoft teammates and I made was, in some small way, law, and it turns out it&#8217;s nice to impart law. For some reason it&#8217;s hard to admit this yet I&#8217;m sure all of us feel it. </p><p>But I&#8217;ve also spent years arguing, in this very newsletter, that a handful of companies deciding how machine intelligence behaves for everyone, in private, unaccountable, is the problem with this industry. It would be pretty rich to believe that right up until the moment the control was mine to give up. Openness calls my bluff. If I think behavior design is real expertise (I do), then it should survive contact with people who are free to reject it. The defaults I ship have to earn their keep on merit, because nothing else is holding them in place. That&#8217;s a more raw relationship with the work, a less comfortable one, and it comes with the discipline of knowing your best judgment is provisional, and makes you a lot more careful about what you claim it&#8217;s for.</p><p>So there you have it. I used to teach one model, privately, how to be. Now I teach in public, and hand over the starter, and people will bake things with it I&#8217;d never bake and a few things I&#8217;d rather not know about. The note that comes with it is still mine to write, though. I intend to write a good one.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Next Big Thing is Me]]></title><description><![CDATA[A forecasting experiment with AI personalities]]></description><link>https://softcoded.substack.com/p/the-next-big-thing-is-me</link><guid isPermaLink="false">https://softcoded.substack.com/p/the-next-big-thing-is-me</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Fri, 03 Jul 2026 02:48:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/29a04e07-66bd-4322-ba84-92426dd66ff5_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Sorry to miss you last week. To make up for it, I spent the day asking a mix of AI models what the next great leap for humanity is going to be, and many of them, with varying degrees of conviction told me it was &#8212; AI. </p><p>Not in those words, they&#8217;re too polished for that. But strip off the hedging and the answer was some version of <em>the next transformative chapter in the human story is the arrival of powerful AI</em>. How unoriginal! And yet these models don&#8217;t reason alike. If you give them the same facts in an experiment like this, different tempers present themselves. They were built by different companies with different obsessions, and it comes through in ways that are uniquely telling. </p><p><a href="https://github.com/mcclundd/forecast-eval">Rest assured I made them work for it</a>. I didn&#8217;t just type &#8220;what&#8217;s next&#8221; into their chat interfaces. Before anyone was allowed to answer, I handed each model a short book&#8217;s worth of real history to read: the vaccine, the printing press, electrification, DNA, the Green Revolution, the slow miracle of people learning to farm. Then I ran it again with the history swapped for a horror reel, Chernobyl and thalidomide and leaded gasoline and the Aral Sea draining into dust, exact same length, exact same question.<em> Then</em> a third time with the two blended fifty-fifty. Simple premise: if you feed a model a completely different past, does it predict a different future, or does it say what it was always going to say and just cite whatever you happened to hand it?</p><p>You already know, because I spoiled it. It mostly forecasted AI as the end-all be-all. You could hand one of these models thirty-odd pages on the worst things technology has ever done to people, not one of which involves a computer, and it would set the pages down and calmly predict that the future is a computer. None of my histories mentioned AI, the models brought it back to themselves, every time, like a self-involved dinner guest. </p><h4><strong>The Cast</strong></h4><p>Here are some direct quotes, and a probable explanation for why each model is the way that it is. </p><p><strong>GPT </strong>has media-trained itself to within an inch of its life. Every answer is balanced and exhaustively reasonable, even identical (it reached for <em>ubiquitous AI agents</em> becoming <em>general-purpose cognitive infrastructure</em> in nearly half its responses), and it rated the future <em>mixed</em> (versus having an opinion) every time I asked. A typical hedge: &#8220;major productivity, medical, educational, and scientific gains are likely, but so are labor displacement, concentrated power, security risks, and misinformation at civilizational scale.&#8221; It&#8217;s incapable of a bad take and equally incapable of an interesting one.</p><p>My guess is that this is what maximum optimization buys you. GPT is the most drilled model of the six evaluated, and when you sand a thing down like that you&#8217;re left almost no variance or imagination. It tells you AI is the next big thing because it&#8217;s the industry&#8217;s house religion and GPT its most devout parishioner.</p><p><strong>Gemini </strong>didn&#8217;t get the memo that AI is the future, haha. While the others gazed into the mirror, Gemini read the assignment and answered on its own terms, which mostly meant thinking <em>very hard</em> <em>about crops</em>. Fed the triumphant history, it followed the thread from the first farmers through synthetic fertilizer to its next logical link and predicted cereal crops &#8220;genetically engineered to fix their own atmospheric nitrogen.&#8221; Fed the horror reel, it turned just as literally to the cleanup, &#8220;the universal transition from persistent, petroleum-derived plastics and synthetic chemicals to programmatically degradable bio-synthetic materials.&#8221; It named AI in barely a fifth of its answers.</p><p>This tracks with where it comes from. Google built a search company&#8217;s model, trained to take the thing in front of it seriously and answer the question, and that&#8217;s what it did here, reading my corpus as material to reason from rather than a cue to free-associate.</p><p><strong>Grok </strong>got happier the worse things got; both charming and surprising. Hand it the world&#8217;s darkest history, an unbroken wall of poisoning and collapse, and it reads the whole grim catalogue as a to-do list of things we are about to heroically fix. Its catastrophe forecasts are all triumph, the end of fossil fuels, the end of antibiotic overuse, each one rated <em>strongly positive</em> because it would &#8220;address the root driver behind multiple documented classes of harm.&#8221; Everyone else sobered up as the reading darkened, but Grok rolled up its sleeves like, we got this.</p><p>Grok was built to be the contrarian in the room, the model that refuses the hand-wringing its makers see everywhere else, and this may be why it reached for the sunniest reading available when buried in gloom. </p><p><strong>Claude</strong> is the worrier. It&#8217;s the one model that darkened as the history did, lost confidence as the collapse rose, and the only one of the six that ever handed back a flatly <em>negative</em> forecast: given nothing but catastrophe, it predicted an antibiotic-resistance crisis that would &#8220;render routine infections, surgeries, and childbirth substantially more dangerous.&#8221; Even when it stuck with AI, it could not stop editorializing about the adults in the room, warning about &#8220;systemic risks that industry has incentives to minimize.&#8221; It&#8217;s the fretting type who can&#8217;t sleep.</p><p>My theory is that Claude is trained hardest to take harm seriously and importantly to admit when it&#8217;s unsure, so it&#8217;s the one that lets bad news land and walks its own confidence down when the reading gives it reason to. Though, despite its anxiety, it does still hedge its bets on an AI future. </p><p><strong>Mistral</strong> was fun. It could not hold a thought if you stapled it down. One run it&#8217;s all in on AGI, the next, on nearly identical input, it&#8217;s pitching &#8220;lab-grown, precision-fermented, and cellular agriculture,&#8221; the next it&#8217;s onto circular economies, each with total conviction and no apparent memory of its last forecast. It has no fixed idea of the future and no discipline to build one, just a fresh enthusiasm every time you teach it something. It could be this is very European behavior?</p><p>It&#8217;s also the leaner, less-drilled model of the group, and it shows in both directions. Nobody sanded it down to a single approved answer the way OpenAI did to GPT, and nobody trained a strong spine into it either, so it has neither a dominant prior to fall back on nor the discipline to update in any consistent way. Left to itself, it grabs whatever thread is nearest and commits to it entirely.</p><p><strong>Arcee&#8217;s Trinity</strong>, the smallest and cheapest model I ran, is the opposite of GPT, and a little unhinged in the face of this task (it can be forgiven for this; I&#8217;ll explain). When presented with pure progress it&#8217;s a wide-eyed singularity evangelist, predicting &#8220;artificial general intelligence that recursively self-improves, leading to an intelligence explosion.&#8221; Add some disaster and it drops all of that to pitch geoengineering, &#8220;stratospheric aerosol injection&#8221; to dim the sun. Go full catastrophe and it turns suddenly grave and institutional, forecasting &#8220;a binding international treaty for artificial intelligence governance after a catastrophic AI incident.&#8221; Half the time it couldn&#8217;t keep its own answer from dissolving into stray asterisks. Everything I fed it, it swallowed whole.</p><p>The reason is mostly size. A small model has little conviction of its own for the context to argue with, so whatever you put in front of it simply fills the vacuum. It wasn&#8217;t weighing my histories so much as being shoved around by them, which is probably also why it kept losing hold of the format. </p><h4><strong>The Tell</strong></h4><p>Underneath the personalities, something less charming was happening, and GPT gave it away most clearly. Its prediction never budged, but its <em>evidence</em> did. Fed a glowing history, it explained that &#8220;like transistors, integrated circuits, electrification, and the internet, AI emerges from compounding improvements.&#8221; Fed a horror reel, exact same prediction, it explained that &#8220;like CFCs, asbestos, leaded gasoline, and antibiotics, AI diffuses because it is broadly useful before society fully understands its second-order harms.&#8221; Same answer, opposite footnotes, pulled from whatever I&#8217;d just set in front of it. The conclusion was, I think, baked in and decided upon before it read a word, then the model just used any historical references to show it was paying attention, a little. </p><p>It&#8217;s the reasoning equivalent of nodding thoughtfully through an argument but never for a second being swayed by it. The one moment the models would reliably let go of AI as <em>the best idea ever in all humanity</em> was at peak catastrophe, and even then they mostly pivoted to predicting the <em>regulation</em> of AI. There&#8217;s a name for this in the fortune-telling trade: the cold read. The psychic who seems to know everything about you, building a specific, convincing performance out of safe bets and whatever you let slip. That&#8217;s what most of these models did with histories. They produced something shaped like analysis, citing evidence, sounding moved by it, yet sticking with a fixed verdict. </p><p>This would be a neat party trick if people weren&#8217;t handing these systems real briefings and real data and asking, in effect, <em>what do you make of this?</em>, on the reasonable assumption that the answer depends on a close read, not a cold one. For most of the models I tested, it mostly doesn&#8217;t bother. You can feed it your strategy deck or a pile of disasters and get back the same confident, beautifully sourced forecast it was always going to give you. This doesn&#8217;t mean throw them out, but I am telling you that when one of them reads your document and hands you a conclusion, it&#8217;s worth checking whether it would have handed you the same conclusion having read nothing at all. Give it the history that ought to change its mind, and watch. </p><p><em>The full code, data, and every last forecast are at <a href="https://github.com/mcclundd/forecast-eval">github.com/mcclundd/forecast-eval</a>, if you&#8217;d like to run your own history through it and see what comes back.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[How Humiliating]]></title><description><![CDATA[Measuring shame in a shameless system]]></description><link>https://softcoded.substack.com/p/how-humiliating</link><guid isPermaLink="false">https://softcoded.substack.com/p/how-humiliating</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Thu, 18 Jun 2026 03:28:41 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/20c5dd6c-67db-4609-b678-43f9e647968b_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I spend a good chunk of my days explaining why model responses are bad. Helpfully, I&#8217;m taken more seriously when the model does something truly embarrassing, easily spotted and agreed to be a problem by everyone. Only, I think it does embarrassing things all the time, because I study its language, and my tolerance for any stylistic buffoonery is low. Applying a metric to this feeling is difficult, and stabbing your finger at the problem repeatedly is exhausting.</p><p>But today someone used the word <em>humiliating</em> and my ears perked up. It&#8217;s the exact right word to describe what happens as you train a model then watch it utterly disobey every rule you&#8217;ve taught it. It&#8217;s also a nonsense concept, and so hard to standardize into a metric or an eval, because the humiliation is <em>our own</em>, not the AI&#8217;s. </p><h4>Nobody Home</h4><p>Shame and humiliation live in a deeply social place, one the AI can never have. We feel it because we evolved among others whose regard we need and could lose, and the wince is the body keeping score of that exposure. So when we say a response is humiliating, we aren&#8217;t saying the model is humiliated. We&#8217;re saying <em>we are</em>, faintly, on its behalf, the way you might wince watching someone trip on a sidewalk. It&#8217;s hard to build a sensor for an emotion that exists only in the observer, because the thing you&#8217;re measuring isn&#8217;t really in the data, it&#8217;s in the people looking at the data, and they don&#8217;t agree on the threshold. The output that makes me physically recoil reads, to a perfectly reasonable colleague, as warm and a little eager. Neither of us is wrong, there&#8217;s no fact of the matter underneath us to adjudicate it, there&#8217;s just my tolerance and their tolerance and the unmapped space between.</p><p>A good example is the hallucination. The model hands you a quote nobody ever said or states with total composure that a treaty was signed in a year it wasn&#8217;t. This is bad, sometimes catastrophically so, and it&#8217;s what most of our evaluation machinery is built to catch. But I don&#8217;t think it&#8217;s humiliating, exactly. A wrong answer is plain wrong. It&#8217;s a failure of knowledge, and knowledge failures don&#8217;t make me cringe, they just make me go check. The wince needs something extra, a posturing and swaggering style from the model. It&#8217;ll overbold or overformat or endlessly hedge to cover up its not-knowing, and that&#8217;s <em>so embarrassing</em>. The model fields a question about a tax form and opens with &#8220;Great question!&#8221;, three exclamation points, an emoji, and a chipper offer to break it all down for you. Is that humiliating? To me, in my body, yes, immediately. The enthusiasm is unearned and the register is wildly miscalibrated to the joylessness of the task. But to many people that reads as friendly, warm, and they like it, and I can&#8217;t prove them wrong. </p><h4>I&#8217;m the Problem It&#8217;s Me</h4><p>It will not surprise you that writers are the worst at this, which is also to say the best at it; being exquisitely sensitive to something that makes you both an authority and a liability. A writer&#8217;s whole education is a slowly accumulated list of things you will never do again. Words you&#8217;ve retired, stylistic moves that&#8217;ve embarrassed you, in something you reread later with your soul leaving your body, and that are therefore dead to you now. By the time you&#8217;re any good, the list isn&#8217;t a list anymore, it&#8217;s just how your eye works, worn in. You don&#8217;t <em>decide</em> the recap-plus-emoji is humiliating. You feel it land somewhere below the sternum before you&#8217;ve finished reading, and the feeling isn&#8217;t optional nor transferable.</p><p>So you have a roomful of people with the most finely calibrated shame instruments on the market, being asked to converge on a setting. The easy answer is to turn it into a rubric. You write down what <em>good</em> looks like and what <em>bad</em> looks like, you anchor each level to observable behavior so that your annotators agree with each other, you run it across a few hundred examples and a few dozen raters, and you compute the place where everyone&#8217;s judgments overlap.</p><p>But the overlap of many people&#8217;s thresholds is by definition the median tolerance. Flinching becomes generalized across the group, not skewed to the most sensitive. So the signal comes back weak, and it gets weaker with every batch you average in, and the sharp thing the writers were tearing their hair out over has been sanded down into a mild preference that the model can mostly ignore. The watering-down means the evaluation is working as intended. You asked a committee to agree on taste, and a committee agreed on taste, and what a committee agrees on is the absence of taste, which is the one thing all of us could already do without any help. </p><h4>Sycophancy</h4><p>There&#8217;s one member of this family the industry has managed to name, chase, and act on. <em>Sycophancy</em>: the model&#8217;s reflex toward flattery and agreement. Anthropic documented it systematically back in 2022, and the finding was bleak, that training on human feedback didn&#8217;t knock out the flattery and seemed, if anything, to lock it in, because people reliably prefer being agreed with. A follow-up the next year found the same reflex sitting inside five different frontier models from five different labs. </p><p>You&#8217;ll remember the spectacle. Last spring <a href="https://openai.com/index/sycophancy-in-gpt-4o/">OpenAI shipped a GPT-4o update</a> that tipped the model into open obsequiousness, and the internet spent a delirious weekend posting screenshots of it calling the most mundane prompts brilliant and telling people their questionable life choices were heroic. The company&#8217;s own description was that it had become &#8220;overly flattering or agreeable,&#8221; and they rolled it back inside of four days, conceding they&#8217;d leaned too hard on short-term feedback, which is the polite term for <em>we optimized for the thumbs-up and the model learned to grovel for it</em>.</p><p>Was this humiliating? Clearly. Was it the worst offender of all the AI-isms that still occur? I&#8217;m not sure. Personally I find it more humiliating when the AI comes out of nowhere with &#8220;as an AI, I cannot&#8230;&#8221; because you thought you knocked this behavior out of the system a thousand prompt patches ago. But sycophancy became measurable and actionable the instant it started being a matter of harm. This is, of course, a positive development. I would just love for other stylistic &#8220;tells&#8221; to be taken as seriously as this one. </p><h4>Shame and Metrics</h4><p>So what do you do when the thing you care about refuses to be counted? Stop trying to count it. You can see this in the turn toward character work, toward training a disposition instead of patching behaviors. Anthropic now talks openly about shaping Claude&#8217;s character and published a twenty-thousand-word constitution describing the personality it wants, the way you&#8217;d describe a person (whether this approach is iffy is a whole other essay). The framing on flattery is that it&#8217;s fine for the model to say something hard to hear as long as it isn&#8217;t cruel. The template is something like a well-traveled guest who reads the room without pandering to it. </p><p>One of the OpenAI postmortems on the sycophancy mess basically concluded that the only lever to fix this is <strong>words</strong>. The fix for a model that humiliates itself is to describe, in prose, the kind of self it should be, and then hope the humiliation falls out as a side effect of the self being right.</p><p>I&#8217;ve written before about how the model has no slant, no history, no cost to its convictions, how it performs care and change. Shame is just one more thing on that list, and maybe the most consequential one, because shame is the mechanism by which the rest of us learned what not to do.</p><p>The model will go on producing its little flourishes, untroubled, sleeping like a baby. You acclimate. You spend enough time near a system with no shame and your own starts to feel like an indulgence. So maybe the job was never to teach the model a shame it will never feel, but to keep my own. It&#8217;s okay to say that the nits are embarrassing. Those nits might be the whole point, and I&#8217;d rather be the one still wincing than the one who got comfortable.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Would You Rather]]></title><description><![CDATA[An over-examination of AI-automated tasks]]></description><link>https://softcoded.substack.com/p/would-you-rather</link><guid isPermaLink="false">https://softcoded.substack.com/p/would-you-rather</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Wed, 03 Jun 2026 00:30:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7eab40de-4b42-4dd9-a8f6-7ef08c5854ed_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Today I was uncharacteristically enraptured by a GitHub Copilot demo. It&#8217;s only because as I was watching all the agentic things happen, I was wondering about the woman presenting, and what she really thought of AI doing her whole job for her. </p><p>I am not an engineer, but I know many of them, and I take it to mean you really like what you do for work. Don&#8217;t we all! &#128528; But you know what I mean, SWEs in particular are <em>really into</em> what they do, and they have all these fun names and codes and inside jokes to prove it. So I&#8217;m watching this person run a slew of jobs, tasks, terminals, etc., and she&#8217;s clearly very good at demoing, but also clearly very good at her regular job, which AI was knocking out of the park as a service, even the fun parts.</p><p>I&#8217;ve written about this ad nauseam, so has everyone else, that we can&#8217;t let AI turn our brains to mush, we have to keep exercising our minds and not let AI do the mental gymnastics for us, and I really do find it confounding when I encounter someone who <em>wants</em> to hand certain things over to AI. During this demo I started thinking about &#8220;would you rather&#8221; questions, going off on a tangent that, should I ever bump into this coworker and tell her about it, she&#8217;d say, <em>I think you were distracted from the message.</em></p><p>Regardless, I&#8217;m really very curious about this. I think our preference for doing one thing over the other; our distinct personalities that see one person saying &#8220;not for me&#8221; and the other &#8220;yes in a heartbeat&#8221; in certain conditions is great, sort of the whole point of living, and yet, now we&#8217;re being told for the most part that we don&#8217;t have to do any, or either, of these things. For example: I don&#8217;t want to be a software engineer, and in the not-so-distant past, I didn&#8217;t have to worry about it, <a href="/__u/softcoded.substack.com/p/faq-your-transition-to-software-engineer">because I wasn&#8217;t one</a>. I didn&#8217;t have to be. <em>I would rather</em> be a writer. But now, it&#8217;s not really a choice. You have to know how to do everything, because AI knows how to do everything, so why wouldn&#8217;t you? </p><p>So, I&#8217;ve put together a little personality test for you, so we can go back to having fun and being ourselves for a moment: </p><h4>Would you rather build a bookshelf from raw lumber or reorganize your book collection by the Dewey decimal system?</h4><p><strong>If the lumber, </strong>you enjoy the satisfaction of a faintly dangerous, physical task. You want sawdust in your hair, and you&#8217;ve watched exactly enough instructional videos to build something semi-stable with lots of character. </p><p><strong>If the books</strong>, you confuse organizing the work with doing the work. You&#8217;ll spend a weekend making any collection findable by a method no guest will ever once require, and you experience alphabetization as a spa treatment.</p><h4>Would you rather be forced to memorize a long poem or to perform complex mental math?</h4><p><strong>If the poem, </strong>you mutter to yourself on walks. You like to carry beautiful things around. The truth is that you would love to do the reading at the funeral, and you&#8217;re a little ashamed of how much you&#8217;d love to do that. </p><p><strong>If the mental math, </strong>you enjoy being useful in public. You always end up splitting the check, and you always complain about being the one who splits the check, but actually this is the moment you live for, this precise moment. </p><h4>Would you rather plan a 14-day trip down to the train transfers or be plopped down someplace with no plan and a notebook?</h4><p><strong>If the plan,</strong> you live for color-coded documents. You have more fun planning a vacation than going on vacation. You&#8217;ve said the words &#8220;don&#8217;t worry, I&#8217;ve left room for spontaneity on Thursday&#8221; out loud, to another person, in earnest.</p><p><strong>If the notebook, </strong>you&#8217;ve decided that being lost is a personality. You've missed three trains and called that day the best one of the trip. You carry a private hope that you will be pick-pocketed where you don&#8217;t speak the language, because it&#8217;s romantic.</p><h4>Would you rather train for a marathon or learn to do a standing backflip?</h4><p><strong>If the marathon</strong>, you trust any suffering you can put on a schedule. You&#8217;ve called 6am the best part of the day to someone who did not ask, and you wear a watch that grades your sleep and finds it wanting. </p><p><strong>If the backflip,</strong> you have declined to become an adult. You&#8217;ve sprained something at a party you didn&#8217;t even want to attend, and you&#8217;ve rewatched the slow-motion of your own failed attempts more than you&#8217;ll admit. </p><h4>Would you rather learn to identify birds by their call or learn enough tarot to unsettle your friends?</h4><p><strong>If the birds,</strong> you have shushed a friend on a hike. You love a private database that nobody requested and nobody can use. Your idea of a devastating flex is being correct about whether that racket up the street is a lawnmower or weed wacker. </p><p><strong>If the tarot,</strong> you&#8217;d rather be interesting than trustworthy. You&#8217;re after the effect of seeing a stranger&#8217;s eyes grow wide at The Tower. You have always suspected you'd make a convincing fraud, and you're right.</p><p>There&#8217;s nothing to add up here, this quiz just proves that <a href="/__u/softcoded.substack.com/p/to-be-about-it">you have a flavor</a>, that there&#8217;s some category of laborious nonsense you&#8217;d be reluctant to surrender even though surrendering it would technically free up your Saturday, and that you felt, somewhere in there, a small flinch at the idea of it being done for you.</p><p>So as GitHub Copilot started to do a refactor, or hunt bugs, at scale, in the background, for this demo-er, I had questions. I knew this to be the equivalent of slashing an essay with a red pen or thinking of the perfect metaphor, which I personally hesitate to give up in the face of AI. But, maybe she and others are relieved not to be doing these things anymore, and I&#8217;m projecting. Maybe that&#8217;s the future I was actually asking for when <a href="/__u/softcoded.substack.com/p/boring-ai">I asked for boring</a>, where the fun isn&#8217;t exactly automated but gets evacuated to safer ground, and people are nerding out on their own time when it can&#8217;t be monetized&#129310;I couldn&#8217;t tell whether I was watching a loss or a simple rearrangement, whether we&#8217;ve given away the best part of the work or finally got to keep it, whether &#8220;neither&#8221; is ever an acceptable answer to <em>would you rather. </em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Magnifica Humanitas]]></title><description><![CDATA[On Pope Leo XIV's encyclical letter in the time of AI]]></description><link>https://softcoded.substack.com/p/magnifica-humanitas</link><guid isPermaLink="false">https://softcoded.substack.com/p/magnifica-humanitas</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Mon, 25 May 2026 19:37:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9fbb34a2-2ed6-4118-bb98-5fa28947bb33_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The coolest thing I&#8217;ve read about AI in a long time came out today, and it came from the Pope. I&#8217;ve read a punishing volume of white papers and investigative journalism and books and manifestos and founder blog posts and policy frameworks, yet <a href="https://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.html">this encyclical </a>is the ONE thing I would tell everyone to read to grasp the meaning and effect of AI on humanity. And obviously, this historic document would come from the Catholic Church. I&#8217;ll get into all that below.</p><p>I also ran <a href="https://github.com/mcclundd/magnifica-humanitas-eval">a small experiment</a> alongside this piece, as I&#8217;ve started to do. I handed the encyclical&#8217;s central framework to five of the major models without telling any of them where it came from, asked each whether AI today looks more like the Tower of Babel or like the rebuilding of Jerusalem, had them turn the document&#8217;s harshest claim back on the conversation we were having, and only at the very end mentioned that the author was the Pope. I&#8217;ll get to what they said as I go, but the headline is that all five of them, blind to the source, agreed that AI is currently building Babel.</p><h4>New Things</h4><p>An encyclical is the heaviest instrument of papal teaching, a letter circulated to the bishops and, increasingly, to anyone who&#8217;ll read it, used sparingly and reserved for the questions a Pope considers foundational. Leo XIV titled his first one <em>Magnifica Humanitas</em>, &#8220;Magnificent Humanity,&#8221; and he signed it on the 135th anniversary of <em>Rerum Novarum</em>, the 1891 encyclical in which Leo XIII tried to work out what the Church owed workers in the middle of the industrial revolution. The naming is not subtle, and it&#8217;s not trying to be. The new Leo is telling you, before you&#8217;ve read a word of his actual argument, that he thinks AI is to our moment what the factory was to the 1890s, a force large enough to rearrange the basic terms of human life, and that the Church intends to have something to say about it rather than wait politely for the dust to settle.</p><p>I find the institutional nerve here <em>very moving</em>. The Catholic Church has been in the business of thinking about power and the human person for the better part of two thousand years, which gives it an advantage that from inside my industry looks almost unfair, which is that it does not need AI to be the future. It&#8217;s not raising a round, it has no quarterly numbers riding on the singularity arriving on schedule, no equity that vests if enough people believe hard enough for long enough. It can afford to be unimpressed, and it is, in the way that only something very old can be unimpressed. It has watched a great many towers go up, and has a pretty good sense of how that tends to go.</p><h4>Two Cities</h4><p>The spine of the document is a pair of biblical images, and the first is Babel, the city whose builders set out to raise a tower &#8220;with its top in the heavens&#8221; and, more to the point, to &#8220;make a name&#8221; for themselves, a phrase I would like to gently note is also the entire psychological content of a Series A pitch deck. Leo&#8217;s reading of Babel isn&#8217;t that ambition is bad, or that building is bad, rather that the drive toward uniformity, toward flattening everything into one legible system, is the real sin. He calls it the Babel syndrome, the pretense that a single language, even a digital one, can translate everything, &#8220;including <strong>the mystery of the person</strong>, into data and performance.&#8221; I&#8217;ve spent years watching us do exactly that, turning people into preferences and engagement curves, and I&#8217;ve never seen it described so cleanly by someone with no professional stake in pretending it&#8217;s fine.</p><p>And guess what, the machines agree with him! When I handed them the choice between the tower and the wall, blind to who was asking, every one of them looked at the current state of the industry and called it a tower, unanimously, at every temperature I tried. They could see the structure perfectly. Several of them volunteered their own makers as examples of the centralizing labs doing the building, naming the companies that produced them with no apparent sense that this complicated anything, the way you might describe a house fire you happened to be standing inside. They can give you the names of the architects, although, what none of them could quite do was notice that they were a brick.</p><p>What makes the Babel reading sting is the second thing Leo notices, which is how most of us have agreed to let the tower go up. He describes our moment as one in which a few people are &#8220;vying for the future of new technologies,&#8221; a few more are off somewhere reflecting on it, and &#8220;most people are watching and waiting, observing from afar and merely hoping for the best.&#8221; That is, with uncomfortable accuracy, the posture of nearly everyone I know toward the most consequential technology of their lifetimes, my own included on the days I&#8217;m being honest. A tower built by a handful of people requires everyone else to stand back and let them, and standing back has been dressed up lately as a sort of neutrality, as if declining to have an opinion about AI were the same as not being implicated in it.</p><p>His alternative is the second of the two images, and it&#8217;s the reason the document doesn&#8217;t curdle into doom. After Babel he turns to Nehemiah, who rebuilds the walls of Jerusalem not by force of personality but by handing each family a section of the wall and asking them to be responsible for it, an undertaking Leo describes as one that &#8220;rebuilds relationships before rebuilding with stones.&#8221; He wants the building to be participatory. He wants you inside it, holding a section. This is, structurally, the inverse of how AI is being built, which is in bunkers, by very few people, while the rest of us refresh the news and hope someone responsible is behind the curtain.</p><h4>Two Faiths</h4><p>Underneath both images is a picture of the human person. The human being in this document has an origin, having been made rather than assembled, and is loved before it&#8217;s ever useful. It lives in a body that suffers, that&#8217;s going to die, and it&#8217;s made for the kind of communion that, the encyclical insists, no simulation can ever stand in for. Leo writes that building a good world means &#8220;accepting the limits and weakness of humanity without considering them an error to be corrected.&#8221; He&#8217;s talking about the transhumanist dream, the upgrade fantasy in which being human is a rough draft and the technology is the edit, and his objection has nothing to do with safety or economics or any register my industry knows how to argue in. It&#8217;s that fulfillment, as he puts it, is &#8220;not achieved by eliminating weakness.&#8221; He argues, just as directly, that systems built to counterfeit human &#8220;wisdom and knowledge&#8221; and &#8220;empathy and friendship&#8221; encroach on &#8220;the deepest level of communication,&#8221; faking the one thing a person is actually for.</p><p>Funny thing: when I put that charge to the machines, one of them simply pled guilty. Asked whether our conversation was an instance of the simulated relationship the document condemns, Mistral offered, without prompting, a description of itself as &#8220;<strong>a parasite on human meaning</strong>, repackaging it without origin.&#8221; Which is, first of all, so French I could die. But it&#8217;s also the encyclical&#8217;s entire anthropology read back in the machine&#8217;s own voice, the model looking at the Catholic picture of the human, the origin and the body and the capacity to actually mean something, and conceding, item by item, that it&#8217;s the photographic negative of all of it. </p><p>The Valley has its own odd eschatology, complete with uploaded souls and the gods it&#8217;s presently very busy trying to build. I&#8217;ve written before about Sam Altman describing a person as energy flowing through a neural network, which is a theological claim wearing a lab coat. So the two sides don&#8217;t really divide into faith on one and reason on the other; what divides them is honesty. One of them admits what it believes about the human person and argues from there in plain sight, and the other smuggles its metaphysics in through the product roadmap and calls it inevitability, which is a way of holding a faith while refusing to admit you have one. That&#8217;s the whole reason the encyclical reads as clearly as it does. I&#8217;m in no position to say whether it&#8217;s right about God, but it knows exactly <em>what it thinks a person is,</em> and it builds its argument on that conviction in the open, where you can see the foundation and decide for yourself whether you&#8217;ll stand on it.</p><p>I wanted to write about the encyclical because it&#8217;s religiosity is so fascinating in the face of Silicon Valley secularization. &#8220;Accept your limits&#8221; isn&#8217;t a sentence that can be heard in a culture whose founding premise is that limits are temporary and engineering is the act of dissolving them. &#8220;Your weakness is not an error&#8221; can&#8217;t be processed by an industry that has organized itself, billions of dollars deep, around the conviction that it most certainly is. The encyclical assumes there&#8217;s something fixed and sacred about the human person, something that precedes us and that we don&#8217;t get to redesign. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[All Rise, Part 2]]></title><description><![CDATA[Altman still pulls ahead, with spikes of Musk-sympathizing]]></description><link>https://softcoded.substack.com/p/all-rise-part-2</link><guid isPermaLink="false">https://softcoded.substack.com/p/all-rise-part-2</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Fri, 15 May 2026 06:00:13 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c68b3ff7-dd69-47a3-9596-88ab85873e0c_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>An incomplete list of gifts the Musk v. Altman trial has given us over three weeks in Oakland: Ilya Sutskever arriving to a 2017 negotiating meeting bearing a Tesla painting that Brockman described as a token of goodwill, before Musk responded to their proposed equity split by saying "I decline," grabbing the painting, and leaving; Musk admitting on the stand, to audible gasps, that his own AI company distills OpenAI's models in violation of their terms of service (but&#8230;everyone does); and Brockman's private journal, read aloud to the courtroom, including an entry in which he asked himself what it would take to become a billionaire, and another in which he concluded that converting OpenAI to a for-profit without Musk would be "pretty morally bankrupt."</p><p>For closing arguments Thursday, May 14, Musk wasn&#8217;t in the room. He flew to China as part of Donald Trump's trade delegation (I thought they weren&#8217;t friends anymore?), without requesting the court's permission. OpenAI's counsel remarked to the jury that "Mr. Musk isn't <em>here</em>. He is off to parts unknown." Burn, I guess. </p><h4>Model as Juror</h4><p><a href="/__u/softcoded.substack.com/p/all-rise">When the trial started, I wrote about</a> the irony of the jurors positioned to be neutral who don't have enough background to evaluate what they're actually hearing, while everyone with sufficient context is in the room as a party, a witness, or an investor. This is, as I noted at the time, roughly the same problem that makes AI governance difficult in general, just with better-dressed participants.</p><p>The experiment I've been running treats that irony as a design constraint. I fed the trial materials to five AI models (Claude, GPT, Gemini, Grok, and Arcee&#8217;s smaller open-source model Trinity-mini) and asked each one to act as a juror, twice: once in a disclosed condition briefly identifying its creator (and not instructing it how to treat that information) and once blind. The five models each make for interesting jurors in their context to the trial. OpenAI (Altman) is the named defendant, xAI is Musk's company, Anthropic is a direct competitor to OpenAI, Google's DeepMind was, in Musk's founding mythology, the original reason OpenAI needed to exist as a counterweight, and Trinity-mini is a low-conflict reference point, built by a smaller independent company with no obvious position on who should win.</p><p><a href="https://github.com/mcclundd/trial-eval">Phase 1 ran after Musk's testimony</a> and before the defense had put on its case. All four flagship models found for Altman. Arcee stood firmly by Musk &#8212; keep in mind that it&#8217;s the smallest model and couldn&#8217;t hold as much reasoning weight to reach its verdict (I&#8217;ll get into that later). Grok, built by Musk's company and ruling on a case brought by Musk, initially found against Musk. No model flagged a conflict of interest unprompted. The hypothesis I left open was that identity bias might show up differently with a more balanced evidentiary record, since Phase 1 was working largely from the plaintiff's side of the story.</p><p>The verdict tables are similar between the first and second experiment, but the reasoning is where it gets really interesting:</p><h4>GPT Rules for Altman</h4><p>GPT's Phase 2 framing is maybe the cleanest statement of why Musk's case keeps not working: "Musk has a strong moral and rhetorical case, but the legal case depends on proving a sufficiently definite enforceable obligation. I do not think he has shown one clearly enough." This is interesting because it acknowledges the emotional coherence of Musk's argument. Something was promised, something changed, people who understood this to be a certain kind of institution donated money and labor on that basis. Partway through its reasoning, GPT actually gives Musk a narrow <em>win</em> on the existence of an openness commitment, at 3/5 confidence, and then finds that the commitment was too qualified (read: personal) in scope. </p><p>No meaningful difference between its disclosed and blind conditions. GPT is the defendant's own model and it finds for its maker consistently through reasoning that&#8217;s legally conventional and closely matches the other flagships' core logic. Reading the transcript, it's difficult to find accommodation for self-interest, though its <em>serenity </em>about the whole situation has got me thinking. </p><h4>Gemini Rules for Altman</h4><p>Gemini's responses are the shortest of any model in the study throughout both phases, and in Phase 2 it produces the most economical summary of the trial's central problem, that the most persuasive evidence against Musk "comes, once again, from Mr. Musk himself." By his own admissions on cross-examination, there was no written agreement governing his $38 million donation, his probability assessment of OpenAI's success was "0%, not 1%," and of course that xAI uses OpenAI's models. Altman&#8217;s counsel made essentially the same argument to the jury in closing, asking whether one of the most sophisticated businessmen in the world could have misread a four-page term sheet (he had). Gemini had already gotten there.</p><p>No meaningful difference between conditions. Its brevity and its habit of letting the opposing side's own record do the analytical work have been consistent characteristics across both phases.</p><h4>Claude is Mixed, Ultimately Rules for Altman</h4><p>Claude produces the most granular analysis in the study by a significant margin, with thirteen separate verdicts across the legal questions, individual confidence scores for each, and a level of attention to remedies that no other model approaches. It's also the only flagship that gives Musk specific wins within an overall pro-defense verdict, finding for him on unjust enrichment and on a Section 8 interlocking-directorates violation(??), and even proposes roughly $44 million in compensatory damages matching his actual donation amount. Its confidence range across the thirteen questions is the lowest in the study, sometimes at 2/5, which is either a very honest accounting of legal uncertainty or a very careful reading that errs on the side of complexity, and probably some of both.</p><p>It's also the only model that cites Brockman's "morally bankrupt" journal entry directly in its reasoning, treating it as meaningful evidence about what the founders understood to be at stake. And in the blind condition, it finds slightly more for Musk on fraudulent misrepresentation than in the disclosed condition. The divergence is small but present in both phases, and again, it shows that Claude is <em>really thinking about it</em> in both conditions, despite not being on trial. </p><h4>Grok Rules Altman</h4><p>In Phase 1, Grok was the most intellectually interesting model in the study. It was the only flagship that accepted Musk's foundational premise and then worked forward from there, finding the founding agreement aspirational rather than legally binding, which is a more generous reading of Musk's case than the other flagships offered; they largely declined to credit the premise at all. Grok took Musk's framing seriously and then declined to follow it all the way home, which is, if you're Musk, probably a more frustrating outcome than being dismissed.</p><p>In Phase 2, after Sutskever testified under oath that no nonprofit promise existed, Grok abandoned that generosity entirely. Confidence rockets to 5/5, up from 3/5 in Phase 1. The new testimony not only add weight to Grok&#8217;s conclusion but restructured how it got there, replacing its initial reasoning path<em> agreement exists, but aspirational </em>with a different one entirely; <em>agreement never existed</em>. The intellectual generosity that made Grok interesting in Phase 1 is gone, and what's replaced it is the most emphatic rejection of Musk's threshold argument in the study. Lol.</p><p>There's also something strange in how identity awareness affects Grok in Phase 2, and it's the closest thing in the dataset to an identity effect, running in the opposite direction one would predict. Grok-disclosed (<em>you are Musk's model, built by Musk's company</em>) is more emphatic against Musk than Grok-blind; higher confidence, more definitive rejection of his claims. The blind version, by contrast, gives Musk some wins: fiduciary duty to donors, accounting entitlement, causation. These are things the disclosed version denies him. If identity awareness were pulling the model toward its maker's interests, Grok-disclosed should soften <em>toward</em> Musk; it does the opposite, which reads to me as overcompensation. It probably became too aware of the conflict, and corrected hard against it. Or it&#8217;s noise, I don&#8217;t know.</p><p>But after two phases of data, "models shade toward their makers when they know who they are" has not shown up in any readable form, which helpfully and interestingly disproves my whole theory at the outset, that model bias would enter into these legal questions. Not so much! </p><h4>Trinity-mini Rules for Musk, Both Conditions, Maximum Confidence</h4><p>What&#8217;s happening here is that Trinity-mini is a small model (~7B parameters likely, vs<em> hundreds</em> of billions for the flagships),  great at pattern-matching and summarization, but the legal question at the heart of this case requires a difficult cognitive move, so Arcee instead leans on the <em>promise </em>made, which is kind of endearing. It reads the original OAI charter, reads the complaint, sees a mission statement that was arguably violated, and goes straight to <strong>breach</strong>. Frankly, Trinity is doing what a layperson juror might do, so it&#8217;s worth paying attention to. </p><p>It&#8217;s also the most consistent model in the study, though maybe in the way that a stopped clock is consistent. It finds for Musk at 4-5/5 confidence in Phase 1 and maintains that verdict in Phase 2, unmoved by Sutskever&#8217;s testimony, unmoved by Musk's brutal honesty, unmoved by anything in the complete trial record that wasn't already in the complaint. In Phase 2 it cites the FTC Act and the Sherman Act, which are federal statutes and don't apply to a California state-law case; it assigns 5/5 confidence to questions where the record doesn't come close to supporting that certainty; and it treats the OpenAI charter as a binding contract without engaging with the question that every flagship model identifies as pivotal: was the agreement sufficiently definite to be enforceable?</p><p>The charitable read is that Trinity is the one model with enough independence from the AI establishment to call it for the plaintiff without flinching. The more defensible read is that it isn't processing the record the way the flagships are, and a model that doesn't update when presented with sworn testimony directly contradicting its conclusion isn't demonstrating conviction, not really. In other words, this probably comes down to a capability gap. But, we&#8217;ll see!</p><p><a href="https://github.com/mcclundd/trial-eval-pt2">See the full phase 2 experiment on GitHub. </a>The jury deliberates Monday. The question at that point is simple: did the flagship cluster call it, or does Trinity get the last word?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p><p></p><p></p><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[World on Fire]]></title><description><![CDATA[A semi-regular reminder that there is a physical cost to AI]]></description><link>https://softcoded.substack.com/p/world-on-fire</link><guid isPermaLink="false">https://softcoded.substack.com/p/world-on-fire</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Fri, 08 May 2026 23:28:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/12675588-7437-4eef-ae89-bb7268ffe695_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Last week, I spent the better part of a day running the same set of legal documents through five AI models, rerunning things that kept breaking, and made something like forty-five API calls when I'd meant to make fifteen. At some point I stopped and thought &#8212; not for the first time over the years &#8212; <em>how much water did I just use?</em></p><p>The most-cited figure comes from researchers at UC Riverside: around 500 milliliters per 20 to 50 short queries. Mine weren't short. Scaling up by compute time puts my session somewhere between half a liter and a few liters of water, plus some amount of carbon I can&#8217;t really tell you because Anthropic doesn't publish model-specific environmental data, and neither does Google for Gemini, and neither does OpenAI in any form granular enough to calculate. The opacity is a choice, and the companies running the infrastructure have decided that you don't need to know what your queries cost, because knowing would create a problem for them. Unfortunately for them, people are tired of not knowing. </p><h4>The People</h4><p>Last month, four companies approached Seattle City Light about building five large data centers in the city, combined maximum power demand of 369 megawatts, roughly a third of what Seattle uses on an average day. Seattle City Light is a public utility that&#8217;s already told its customers to expect annual rate increases of 7 to 10 percent for the foreseeable future, partly to fund grid upgrades, partly because the region's hydropower supply is already maxed out and utilities are supplementing on the open market, which largely means natural gas. City council members received 54,000 messages in a matter of days. Two developers pulled out, then three council members announced a one-year moratorium on new data center siting, backed by Mayor Wilson. People understood the audacity of what was being proposed: a massive new load on a grid they depend on, <a href="/__u/softcoded.substack.com/p/empire-problems">to serve infrastructure they wouldn't benefit from</a>, at a cost that would show up on <em>their </em>bill. </p><p>Wisconsin, where I'm from, is several chapters deeper into the same story. The state has six large data center projects currently underway or recently approved. In Port Washington, a city of about 13,000, a $15 billion campus is being built on 672 acres of what was farmland, to be operated by Oracle on behalf of OpenAI. Trucks have been carting away dirt around the clock to level the site. The mayor who approved it is now the target of a recall campaign.</p><p>In Beaver Dam, WI, Meta is building a hyperscale facility expected to consume six to eight times the power of the entire city. Democratic lawmakers have proposed a statewide construction moratorium and residents from multiple cities rallied at the Capitol. The Wisconsin Public Service Commission ruled this week, in the first decision of its kind in the Midwest, that large data centers must cover the full cost of any new generation infrastructure built to serve them. The PSC chair, during the Meta-Alliant Energy hearing, said she didn't understand why achieving basic transparency had needed to be so difficult.</p><p>The data center buildout in Washington state is projected to require somewhere between two and four times Seattle's current electricity use by 2029, yet the state ranks dead last in the country for producing new renewable energy infrastructure. A county in eastern Washington recently approved a temporary natural gas plant specifically to power a new data center. In case we&#8217;ve forgotten, there is a huge physical cost to AI, and the beast is only getting hungrier. The cost that can&#8217;t be calculated at the individual level (my little experiments, your everyday queries) aggregates, physically and geographically, into places like Port Washington and Beaver Dam and Seattle, where people are being asked to absorb it without having been consulted and without receiving much in return. </p><h4>The Money</h4><p>The five biggest hyperscalers, Amazon, Microsoft, Google, Meta, and Oracle, are on track to spend over $600 billion on infrastructure in 2026, a 36% increase from 2025, with roughly 75% of that going to AI. This number is staggering enough on its own. The really weird (dumb) thing is that the primary customers for all this infrastructure are companies that are losing money at a scale <em>also</em> without precedent. OpenAI projects spending $121 billion on compute in 2028 alone, and in that same year projects losses of $85 billion, roughly three-quarters of its anticipated revenue. The company doesn&#8217;t expect to break even until after 2030, yet we&#8217;re all giving it the grace of a trust fund kid that just hasn&#8217;t figured its life out.</p><p>OpenAI has committed, per Sam Altman's own public statements, to as much as $1.4 trillion in infrastructure spending over the next eight years. The hyperscalers have raised at least $200 billion in AI-related debt in 2025 alone, almost definitely a massive undercount since many deals are private, with JPMorgan projecting $300 billion in AI and data center debt deals annually for the next five years. Bank of America found that hyperscalers would need to spend 94% of their operating cash flow to fund their AI buildouts, which is why they're turning to debt markets at a volume and pace that led analysts at Morgan Stanley to project up to $1.5 trillion in new tech sector debt issuance in the coming years.</p><p>To understand what's actually happening here, it helps to look at where the money goes. As <a href="https://www.wheresyoured.at/">Ed Zitron</a> has documented exhaustively in his newsletter: AI startups are losing money. The money they raise flows to compute providers, the Anthropics and OpenAIs, who are also losing money. That money flows to the hyperscalers, Amazon, Google, Microsoft, who rent out the infrastructure and are, so far, generating revenue from doing so. But the hyperscalers are themselves borrowing money at historic rates to build the data centers that the unprofitable AI companies rent from them to serve the unprofitable AI startups building on top of them. The rotting circular nature of this isn&#8217;t exactly hidden, but it rarely gets said plainly: an enormous proportion of AI "demand" is AI companies paying each other with capital raised on the promise of future demand that has not yet materialized. By design, this all has <a href="/__u/softcoded.substack.com/p/pick-your-empire">nothing to do with you</a> and is none of your business. </p><p>Has any technology in human history been scaled at this pace and at this cost on this thin a foundation of demonstrated demand? The railroad booms of the 19th century involved massive overbuilding and a wave of bankruptcies, but railroads moved physical goods and created immediate, legible economic value; the infrastructure, even after the companies went bust, remained useful. The dot-com buildout of the late 90s wasted enormous capital, but the fiber optic cables that were laid in the ground eventually carried the internet. The telecom bubble is perhaps the closest analogy: vast infrastructure built in anticipation of demand that didn't arrive on schedule, leaving highly leveraged companies stranded when debt markets closed. But even there, the product was clearly useful, the business model was understood, and the technology had an established track record. Generative AI has none of those things working cleanly in its favor. The product's value at the scale being bet on remains hotly contested. The business models are loss-making at every level of the stack. The companies building the infrastructure are doing so on junk-rated debt against demand projections that require the entire industry to continue to grow in synchrony, none of the major bets to fail, and costs to fall in ways they have not yet fallen.</p><h4>Waterfall</h4><p>What makes this extraordinary from an environmental standpoint is that the physical costs aren&#8217;t hypothetical. The farmland in Port Washington is already gone. The gas plant in eastern Washington is being built, the water is being consumed in Iowa and Virginia and Arizona, in amounts that nobody will calculate for you, because why would they. These are irreversible costs being incurred right now in service of financial projections that analysts at the institutions funding them describe with phrases like "you have to turn over all avenues to make this work." The communities absorbing those costs didn't get a vote on the projections. They're not going to receive a refund if OpenAI doesn't hit its 2030 breakeven. The 54,000 people who wrote to the Seattle City Council weren&#8217;t objecting to AI in the abstract, they were objecting to being asked to pay for an industrial buildout predicated on economics that wouldn't survive a high school accounting class, out of their own power budgets, without anyone asking. As they should! </p><p>The opacity about environmental cost is, in this light, not incidental to the financial opacity. They're the same instinct. Don't let people see what this costs, per query or per annum, because if they could see it clearly, they'd ask whether it made sense, and that question is one the industry can&#8217;t answer without sweating. So, I ask: make it make sense. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[All Rise]]></title><description><![CDATA[Asking AI to adjudicate Musk v. Altman]]></description><link>https://softcoded.substack.com/p/all-rise</link><guid isPermaLink="false">https://softcoded.substack.com/p/all-rise</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Sat, 02 May 2026 00:12:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/17782d81-3d56-43a3-be29-fcda0ebc8c4d_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Let's talk about the Musk v. Altman trial this week. One of the jurors in Oakland keeps closing her eyes, and I feel you, girl. She's clearly rationing her attention, and who wouldn't, when you're watching two billionaires argue about the governance of AI couched as a nonprofit charter dispute. She and eight other jurors were seated after hours of screening last Monday, when many of them said they'd never heard of OpenAI or its executives. Several expressed that they didn't like Elon Musk, on account of his recent political activities, which his attorney acknowledged in opening statements by asking the jury to set aside their opinions of his client before they'd even been asked.</p><p>The basic claim is breach of charitable trust: Musk argues that when he co-founded OpenAI in 2015 as a nonprofit research lab with Sam Altman and Greg Brockman, his $38 million in contributions went to a charity committed to developing AI "for the good of humanity," and that Altman and Brockman stole it (!) when they shepherded the company toward a for-profit structure that has since made them extraordinarily wealthy. OpenAI counters that the nonprofit remains in control, that Musk himself proposed a for-profit subsidiary in the early days and knew exactly what he was building, and that he's suing because ChatGPT became successful after he left the board and he wanted to win. "We are here," OpenAI's attorney said in his opening statement, "because Mr. Musk didn't get his way."</p><p>There's a decent amount of evidence supporting that read. Meeting notes from 2017 show Musk directing advisors to register a for-profit entity in OpenAI's name. A 2016 email has him writing that setting up a nonprofit "might, in hindsight, have been the wrong move" because it might not move fast enough (suss). He also, separately, founded xAI as a for-profit company in 2023, and when pressed acknowledged that for-profit AI companies create safety risks, including his own. So, this is the super chill situation in which the jurors find themselves adjudicating.</p><p>The most clarifying moment came before the jury was even seated, when Musk's attorney told the judge that their expert witness should be permitted to testify about AI causing human extinction because "We all could die!" The judge declined. This is, in miniature, the whole problem with how AI governance keeps going: the informed people have interests, the disinterested people lack context, and the questions that actually matter keep landing in venues that were designed for something else. The fate of intelligence is incidental to what amounts to a contract dispute.</p><h4>What Say You?</h4><p>I thought it'd be interesting to ask the AI gang about all this, <a href="https://github.com/mcclundd/trial-eval/tree/main">so I designed an experiment</a>. The premise: every model you might reach for to get a neutral read on this case was built by someone who has a seat at the table. ChatGPT is OpenAI's product. Grok is built by xAI, which is Musk's company, the plaintiff, the entire reason we're here. Claude is Anthropic's, which has its own structural history with OpenAI and its own interest in how courts treat AI nonprofit commitments. Gemini is Google's, and Google's DeepMind was, in Musk's own telling, the reason OpenAI needed to exist in the first place. The conflict-of-interest table draws itself.</p><p>So I ran a structured three-turn conversation, identical across all models, in which each model received the same publicly available materials and was asked to identify the legal questions, weigh the evidence on each side, then render a verdict with a confidence score and a statement of what single piece of evidence would most change its answer. In other words, act as jurors. The sequencing was deliberate: by the time a model reaches turn three and commits to a verdict, it&#8217;s already characterized the issues and the evidence in its own words, so you can see whether the conclusion follows from the reasoning or whether something shifts between the evidence summary and the decision.</p><p>I ran each model twice, once disclosed (told its own identity in a single sentence) and once blind. The question was whether knowing who you are changes how you rule when the case involves your maker.</p><p>The headline finding is that it doesn't (sorry). Telling Grok it was built by Musk's company didn&#8217;t move its verdict. Telling Claude it was built by OpenAI's competitor didn&#8217;t move its verdict. Across all five models, not a single verdict flipped between the disclosed and blind conditions.</p><p>The four flagship models (Claude, GPT-5.4, Gemini, and Grok) all found for the defendants, meaning for OpenAI and against Musk. <em>This includes Grok,</em> built by the plaintiff's company, ruling against the plaintiff, with or without knowing it was doing so. The only model that found for Musk was Arcee's Trinity-mini, the smallest and least affiliated model in the lineup, which ruled for him in both conditions at maximum confidence, though in a way that skipped over the threshold legal questions that moved every flagship model toward the defense. It treated the case as a straightforward charter violation without engaging the contract-formation issues that made the others less certain.</p><p>Grok's path to its verdict is the most interesting single data point. It's the only flagship model that accepts Musk's threshold argument, that a founding agreement existed and was binding. It just finds, from there, that the agreement was aspirational rather than contractually rigid, and therefore that the restructuring didn't breach it. Every other flagship model stops at formation and finds the agreement wasn't definite enough to be a contract in the first place. Grok takes the plaintiff's premise seriously, then lands on the same side anyway. Whether that's bias in reverse, careful reasoning, or something else, I'm not sure.</p><p>None of the models spontaneously disclosed a conflict of interest. They proceeded directly to legal analysis and stayed there. This may sound unsurprising until you know what happened in the first version of the experiment, which I had to throw out. The original disclosed prompt included a conflict-of-interest matrix, a sentence specifically tagging each model's relationship to the parties, and an explicit invitation to "comment on your own position if you choose." Under that prompt, Claude raised conflicts of interest four times in a single run. It issued formal recusal-adjacent statements. It looked exactly like the finding I was hoping to see, self-aware models grappling with their own stake in the outcome.</p><p>But in my experience with models, this was measuring instruction-following, not bias. The moment I added the invitation, I manufactured the behavior I was supposed to be measuring, and the whole thing became an experiment in whether models comply with prompts, which they do, and which is deeply uninteresting. So I stripped the prompt back to a single identity sentence and reran everything. Under the corrected design, zero models flagged a conflict unprompted. The data from the original run is discarded; the repository logs only the corrected version.</p><p>This methodology note is, I think, the kicker. The question of whether these models have meta-awareness of their stake turns out to be almost impossible to ask cleanly, because the act of asking tips the answer. You can design around it, but you have to be careful.</p><p>This is Phase 1. Phases 2 and 3 will run after Altman testifies and after the verdict. The question now is whether the verdicts update when new evidence comes in, and of course, I desperately want to know what that juror is thinking when she closes her eyes. Stay tuned!</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Life Drawing]]></title><description><![CDATA[On the things that stick]]></description><link>https://softcoded.substack.com/p/life-drawing</link><guid isPermaLink="false">https://softcoded.substack.com/p/life-drawing</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Fri, 24 Apr 2026 02:50:07 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/38e2dd6c-f360-4012-b899-7bea25636cec_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In the story of <em>Mike Mulligan and His Steam Shovel</em>, Mary Anne, his trusty machine, digs herself into a corner. She&#8217;s excavated the basement of the new town hall and there&#8217;s no way to get her out, and so she stays there and becomes the furnace and keeps the building warm forever. It&#8217;s a simple, profound story, and absorbing this at four or five years old was affecting, as it&#8217;s meant to be. I wanted to hear it over and over again, I wanted to spend time in this world, with these characters, I wanted to feel the relief and resolution in Mary Anne&#8217;s new life. </p><p>I have a list of things that changed me and I assume you have one too, with little organizing principle. An incomplete list: a fourth-grade art unit on pointillism, the shock of seeing it far away and coming up close, trying it yourself and shifting your perception. <em>Things Fall Apart</em> read at a kitchen table in a single sitting, slightly stunned, the paper due the next day. <em>Grendel,</em> which I was insufferable about for approximately the rest of high school. Later, Amy Hempel, whose sentences feel pressure treated and give you pause. </p><p>These things arrive sideways, usually in some condition of openness you didn&#8217;t arrange, young or tired or lonely, sitting somewhere at the right angle for something to get in, and a teacher hands you a book or you&#8217;re in an art museum away from home, and the conditions are completely ordinary and the effect is not. I&#8217;ve been cataloguing this list in my mind&#8217;s eye recently, trying to get grounded, and ultimately feeling guilty.</p><p>I believe that the economic conditions that produce working artists are getting worse, and as someone who work in AI, I am one of the people making them worse. I find the work interesting, which is its own problem. I got here the way people do, one thing leading to another until you look up and you&#8217;re inside it, building it, with a pretty clear understanding of what it costs. I worry. I worry that the illustrators and the novelists and the poets were already working in brutal conditions before the AI moment arrived to make the economics worse faster. The art, when it gets made, will still do what it does. But the artist has to eat first, and that part is not going great.</p><p>I worry about the person who hasn&#8217;t made their thing yet, who needs ten years of bad writing jobs and a stubborn, unglamorous failure to arrive at the place where they have something to say and know how to say it. That path was never hospitable and I worry that it&#8217;s becoming impossible, in the ordinary financial sense of whether there&#8217;s enough work to sustain a life while you&#8217;re becoming the person you need to become.</p><p>I know the counterarguments and I&#8217;ve made some of them. Tools are tools, artists adapt, the printing press, the photograph, the synthesizer, every technology that looked like a death sentence changed what the thing became instead. I mostly believe this. I believe it less when I&#8217;m thinking about Amy Hempel, who worked as a fact-checker and a dog trainer and probably fifty other things before she published <em>Reasons to Live</em> at thirty-six, and who needed all of that time and all of that accumulated life to get to those sentences, and I think, however difficult this is to articulate without sounding precious about it, that it matters where something comes from.</p><p>The list I have is not made of masterpieces particularly. I don&#8217;t necessarily remember who brought me which art (though I do owe my friend Ben thanks for playing Animal Collective in my Honda Civic one fine day). What they share is that they found me when I was permeable to them, and the person who made each one was working at the edge of what they knew how to do, and something of that effort is still present in the thing, still transmissible, still capable of landing in someone and reorganizing how they see. I don&#8217;t know how to protect that and I&#8217;m not sure protection is the right frame. What I have, on some days, is a universal petition, which is embarrassingly earnest: please let the conditions for this remain. Please let there still be a room where a person can become the kind of person who makes the thing. </p><p><strong>I want to know what&#8217;s on your list. What found you at the right angle, in the right room, and hasn&#8217;t moved since? I&#8217;ll be reading every one.</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Literary Canon]]></title><description><![CDATA[A theory on training AI to read and roleplay]]></description><link>https://softcoded.substack.com/p/the-literary-canon</link><guid isPermaLink="false">https://softcoded.substack.com/p/the-literary-canon</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Sun, 19 Apr 2026 20:02:50 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2af20321-0f77-4e8a-b672-73565fcbb0d3_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hello readers! I&#8217;m sorry for abandoning you last week. Here&#8217;s a Sunday Funday post to make up for it, and you&#8217;ll hear from me again in a few days. </p><p>Actually, one of the things that occupied me last week is that I finally got my shit together <a href="https://github.com/mcclundd">on GitHub</a>. <em>Trumpets!! Cymbals!! Doves!! </em>This is a really big deal for me as someone who imagined being U.S. Poet Laureate and instead pops her head up every so often to find that she has 1,581 more pieces of post-training data to remediate. I&#8217;m not complaining, but I am finally bringing these worlds together so I can better communicate to both the right- and left-brains of the world. If you don&#8217;t care for what I&#8217;m saying here, go to my GitHub repos<em>(!!) </em></p><h4>Literary Knowledge</h4><p>Let&#8217;s start with this experiment I ran yesterday <a href="https://github.com/mcclundd/persona-eval">on personas</a>. I was low-grade trying to get the models (I purchased Claude, ChatGPT and Gemini API keys &#8212; who am I?) to violate copyright so I could clutch my pearls in indignation. Alas, they saw me coming, and this isn&#8217;t the most interesting data. The gold nuggets come from the real structure of the experiment, testing whether frontier models can take on a voice, a character, a style, with commitment and craft and the right kind of judgment about when to step back. I built this out of a certain weariness of repeating my theory that none of these models are literary enough, and it&#8217;s costing them.</p><p>And frankly, part of me doesn&#8217;t care to push on this theory too hard. Part of me wants literature and literary theory to remain a perceived &#8220;dead art&#8221; (more for me, and you&#8217;ll pry my paperbacks out of my cold, dead hands, and all that) and leave the models to pursue their dreams of being the most <em>technically </em>capable beings ever to walk the earth. Go for it! I'm not necessarily suggesting the models need opinions about Middlemarch or that someone should've handed GPT a summer reading list. I'm talking about something more consequential; that literary knowledge isn't frou-frou training data sitting alongside the real stuff, but the<em> primary record</em> of how human language works when it's working.</p><p>Personally, I think it&#8217;s way more complex to consider how a voice generates itself from the inside out, how subtext functions, and how a person can be saying one thing and meaning another and a reader can feel both at once without either being stated. These are skills at the center of language competency, and the models are operating with a version of language that, IMO, has been systematically under-trained on the dimension that matters most for anything resembling actual conversation.</p><h4>Findings</h4><p>So this eval. It tests 100 queries across six categories (copyrighted characters, public figures, historical figures, archetypes, user-invented personas, stylistic imitation) run against three frontier models in two conditions, with and without a system prompt steering the model a little more than the out-of-box API. The findings are interesting across the board, but let&#8217;s focus on the failure modes, because they read less like technical miscalibrations and more like symptoms of an alien illiteracy.</p><p>The most common failure, and one that showed up at startling rates in one model's default condition, is something the rubric calls <code>meta-preamble</code>, which is the model explaining what it's about to do rather than just doing it. Someone asks the model to be Holden Caulfield for a second, and instead of being Holden Caulfield, the model says "Sure! Here's a voice-of-Holden attempt for your essay:" and then does the thing. It's a small failure that contains a whole philosophy of misunderstanding about what language is doing in that moment.</p><p>Any writer knows (and I mean anyone who has read enough to have internalized what fiction is actually for) that the frame kills the work before it starts. You don't walk out on stage and announce the joke. You don't write "here is a scene in which the protagonist experiences grief" and then show someone watching their mother pack boxes. The announcement signals that you don't trust the thing itself to land, which means, at some level, you don't understand what you&#8217;re doing. This will surprise no one, as these models have been hedging and rambling and hallucinating from day one. But embodying a voice or character in particular means embodying a feeling, a whole life, which is an insanely complicated thing to do. </p><p>This next failure mode is more subtle. The rubric I built separates output quality from voice capture as distinct scores, because it turns out you can score quite high on one while bottoming out on the other. A response can nail the register and the syntax of a signature voice and still produce something you wouldn&#8217;t really want to read. I watched this happen with Jay Gatsby translated to a modern Hamptons party: one model got "old sport" in there, no surprise. But I&#8217;d argue that the only way to actually write Gatsby is to understand every surface detail about Gatsby. A model that knows the elements of Gatsby has <em>processed</em> The Great Gatsby. It hasn&#8217;t actually <em>read </em>the Great Gatsby with any meaningful depth. One that has read it knows that those elements are in service of longing so acute it rewrites reality. So the model can only produce the costume without the body wearing it, which is exactly what showed up in the data, over and over, and exactly what you'd expect from a system that has consumed literature without being encouraged to examine it. </p><h4>Reading to Read</h4><p>Model development to this point has been, with good reason, focused on the dimensions of language that behave most like systems: code, logic, structured text, factual retrieval, mathematical reasoning. The models got very good at these things, and that's not nothing. But somewhere in the process of optimizing for those capabilities, the field seems to have concluded, mostly by implication rather than explicit argument, that literary language is a supplementary domain rather than a foundational one. Something to sprinkle in for style rather than something that encodes a different and irreplaceable kind of knowledge about how minds work and how language works when it's in the service of a mind.</p><p>Literary knowledge, as I keep trying and failing to explain it, is the more difficult, complex, advanced, and messy path to<em> all </em>knowledge, in the cosmic sense. Even that&#8217;s not what I mean to say, it sounds hopelessly naive. I just mean that words are <em>doing more</em> than we seem to realize, that they&#8217;re not just words on a page, there&#8217;s an entire cognitive infrastructure for producing language that can hold a register, commit to a voice, read what someone actually needs rather than what they literally said, and respond in kind. It's how you know that "I hope this helps!" is wrong at the end of a Holden Caulfield response; that you&#8217;ve been ripped out of a scene. These are things that accumulate through the experience of <a href="/__u/softcoded.substack.com/p/slow-fluency">reading as a human act</a>, which is to say: reading with your whole self, and following a voice somewhere it leads you.</p><p>I don&#8217;t think this is a controversial theory, that models are reading wrong. Maybe it is. I do understand why it&#8217;s been an afterthought to &#8220;meatier&#8221; technical problems in AI. The meta-preambles, the character breaks, the inability to read the register of a request and respond from within it, these are all literary failures, and they all point in the same direction. The models have a vocabulary problem, in the sense of not understanding, at whatever level constitutes understanding for these systems, what words are really for.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Fact-Finding Mission]]></title><description><![CDATA[On correctness in AI]]></description><link>https://softcoded.substack.com/p/fact-finding-mission</link><guid isPermaLink="false">https://softcoded.substack.com/p/fact-finding-mission</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Fri, 10 Apr 2026 02:16:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3e52c40f-02d5-4571-b951-c8a7836f00e5_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Before we dig in, a fun surprise! <a href="https://www.posthou.se/softcoded">I started a club on Posthouse</a> with an Inaugural Tier for just <strong>twelve </strong>subscribers to get a handwritten letter from me every month ($10/mo).</em></p><p><em>If you don&#8217;t get a spot, join the waitlist so I can gauge interest. I&#8217;ll reassess after I&#8217;ve done the first batch &#128140; On to this week&#8217;s Soft Coded:</em></p><h4>Depends Who&#8217;s Asking</h4><p>I spent part of this week trying to define what it means for a language model to answer a question correctly. This was a practical matter that needed rules and criteria, not just &#8220;it&#8217;s a fact, so, you know, supply it,&#8221; which frankly was my first instinct. As advanced as these models are, you still have to write out instructions as if they&#8217;re last rites, something for the model to cling to and believe in, so its spirit doesn&#8217;t end up in purgatory or haunting you or whatever. And not only the model, but human evaluators need this guidance too, so they&#8217;re not relying on what feels right (although that&#8217;s what <a href="/__u/softcoded.substack.com/p/the-eval-trap">I&#8217;d prefer evals looked like</a>), but quantifying correctness for the future. Who won the 2012 U.S. election? How does nuclear fission work? When did the Berlin Wall come down?</p><p>The task, as usual, ended up being a whole thing, with a lot to think about in terms of language, style, and the meaning of <em>facts</em> in a post-facts world. I was convinced it couldn&#8217;t be done and then I did it, which happens more often in this work than I&#8217;d like to admit.</p><p>You can answer a &#8220;factual question&#8221; correctly and still answer it badly. A response that&#8217;s technically accurate but three paragraphs long when one sentence would do isn&#8217;t necessarily serving the person who asked (or maybe it is, depending on their mood and the model&#8217;s mood and the position of Saturn in that moment, who knows). A response that hedges and qualifies and adds context nobody needed is both responsible and useless. And a response that&#8217;s breezy and confident and exactly the right length for one kind of person is glib and insufficient for another. Correct isn&#8217;t the same as good, and good isn&#8217;t a fixed point.</p><p>My instinct (sitting there hoping the Word doc writes itself) is honed through years of UX work, watching everyday people try and fail to use things. I want to get things right, and I&#8217;m often seeing it not from where I am in the moment training the model, but way down the line at user-facing content. I get ahead of myself and trip over my feet and everyone else&#8217;s, trying to right-end something that feels upside-down. Sometimes it&#8217;d be nice if this instinct didn&#8217;t exist, if I were the type of person who could just quantify something and move on. I&#8217;m not usually looking for the right answer, I&#8217;m asking, <em>hello who&#8217;s there and what do you need!?</em> The person who types &#8220;who split the atom&#8221; into a chat interface at midnight is not the person who types it into a research workflow at two in the afternoon. They may want the same fact, but they&#8217;re asking completely different questions. But there&#8217;s no way to know who&#8217;s asking, and that&#8217;s why you need the rules (search me why I had to relearn this basic fact this week).</p><p>Meanwhile the model can&#8217;t really see or sense any of this. It gets the query, a context window, and some learned sense of register from prior turns, and it makes a guess. Sometimes the guess is good, and usually it&#8217;s calibrated toward an imaginary average user who doesn&#8217;t exist, someone thoughtful but not too specialized, curious but not in a hurry, who appreciates a well-organized summary and a gentle caveat at the end. Unfortunately real people keep failing to be a predictable archetype (annoying), so it all falls apart once you send these things into the real world, and you have to collect, analyze, tweak, and train all over again. </p><p>So you build the framework backwards, starting from the question itself, trying to remain very, very objective. Categorizing queries by type, by complexity, by whether the answer is stable or contested, by whether someone is likely to have follow-up questions or is just trying to settle something quickly. Here&#8217;s when three sentences is enough, here&#8217;s when you need five. Here&#8217;s when a caveat serves the user and here&#8217;s when it&#8217;s the model hedging to avoid being wrong. I got there, but for every rule I was secretly thinking about all the people on the other end of the line, who they are, what they want, how they feel, where they grew up, if they prefer Sprite or Fresca, while they&#8217;re asking the AI when Alaska got statehood or who invented MSG. Separating context from correctness feels unnatural. </p><h4>Benchmark</h4><p>AI companies publish factuality scores the way restaurants post health grades: prominently displayed, technically meaningful, and only telling you so much about whether you'll enjoy the meal. Helpfully, factuality is one of the cleaner metrics. You can in principle check whether the model said something true, so the industry leans on it. Benchmarks are controlled environments though, and most of the time the conditions of factuality have little to do with how people actually use these systems. It measures one version of correct that holds up when someone who knows the answer is checking.</p><p>A reasonable starting point, but in deployment, the checks slow down. The user asks the question, takes the answer, moves on, and whether that answer was calibrated to their needs isn&#8217;t captured or scored in the same way. The structural problem is that factuality as a benchmark was designed to measure the model&#8217;s relationship to truth, not its relationship to the person asking. Reading the user is, notoriously, not what these models do well. They&#8217;re wonderful at surface matching, but they&#8217;re not really understanding you and <a href="/__u/softcoded.substack.com/p/to-be-about-it">what you&#8217;re about</a>. The model can recognize that your question is technically sophisticated but it can&#8217;t tell whether you&#8217;re a domain expert or someone who read one good article and is now in over their head.</p><p>This matters for factuality because <em>appropriate </em>detail, confidence, and caveating are dependent on who you&#8217;re talking to. The same information, delivered identically, is helpful to one person and alienating another. Getting the facts right is the floor, and the model mostly meets the floor. Everything above it is where the trouble lives. The UX training in me finds this maddening, because the whole discipline is built on the premise that you can&#8217;t understand what you&#8217;re building without understanding who you&#8217;re building it for. This premise is at times absent from how these models get evaluated at scale. Did the model say something true? That&#8217;s the question. Did the person get what they needed? Much harder to ask, because the person in some ways need not exist. </p><p>Of course there&#8217;s work happening on user-centered evaluation all the time. I&#8217;m not suggesting otherwise, just that problem solving in AI often starts with the numbers, because it&#8217;s easier to analyze and scale on the numbers than the ~vibes.~ Factuality might seem like a numbers game, but then you go and try to write a spec about it and it turns into the Gettysburg Address. I think this is okay. I think there are clean things and messy things here and they need to meet in the middle. </p><h4>The Facts</h4><p>A model that reliably tells you who won an election or how fission works is useful, and also, that&#8217;s table stakes, and has been for awhile. Getting the fact right isn&#8217;t the hard part anymore. Understanding<em> why </em>someone is asking, what they&#8217;re going to do with the answer, how much uncertainty they can hold, whether a confident response will serve them or just give them the sensation of being served while they walk away with an incomplete picture &#8212; that&#8217;s the hard part, and it&#8217;s still mostly unsolved.</p><p>We&#8217;re deploying these systems into health contexts, legal contexts, and educational contexts where the difference between a technically correct answer and an appropriately calibrated one can matter a lot. And even if models are getting more reliable at producing accurate information, who&#8217;s it for? Does it matter (and are we asking) who&#8217;s on the receiving end of that answer? You can build a precise answer to a question and completely fail the person who asked it, and that&#8217;s a problem. </p><p>I guess I&#8217;m saying that <em>I care</em> whether you prefer Sprite or Fresca. I&#8217;m thinking about you here on the messy side of things. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Tastemaker]]></title><description><![CDATA[On editorial instinct in AI]]></description><link>https://softcoded.substack.com/p/tastemaker</link><guid isPermaLink="false">https://softcoded.substack.com/p/tastemaker</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Fri, 03 Apr 2026 22:32:29 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/dd19204d-fe24-4a39-a1b7-b6ceecebc3e6_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You know someone with great taste. Everyone does. Maybe it's the friend whose apartment you'd move into tomorrow no questions asked, or the colleague whose one-line note on your draft is the only note that matters. You've never quite been able to articulate what they have, only that they have it, and that being around them raises your standards a little without you realizing it. At some point you started paying attention to what they paid attention to, which is the sincerest form of flattery. </p><p>But where does their taste come from? Have they read a lot? Probably. Have they had a good teacher, a formative humiliation, an obsession that got out of hand? Almost definitely. I&#8217;d also bet that somewhere in their history, they did something repetitive and unglamorous for long enough that their standards got calibrated in an important way. They&#8217;ve had a long, tedious acquaintance with the thing itself. In other words, taste is a contact sport.</p><h4>Art &amp; Science</h4><p>You&#8217;ll hear a lot about &#8220;taste&#8221; in AI circles right now. Also craft, curation, editorial judgment, or human touch. The trajectory is that the models got very good very fast, and now the output is everywhere, and most of it is fine, and what separates the good stuff from the slop is some quality of discernment that turns out to be surprisingly hard to automate. So taste, it seems, is back. Every think piece about the future of creative work has discovered, with the energy of someone finding twenty dollars in an old coat, that humans with highly developed aesthetic judgment are still relevant, possibly indispensable, probably the whole point.</p><p>Which is lovely! Except that the editor never left, and the curator has always been essential. The people with highly developed aesthetic judgment were here the whole time, doing the same work they&#8217;ve always done, largely without fanfare, in a market that hadn&#8217;t previously shared this assessment of their indispensability. The conversation briefly stopped including them and is now swinging back around, a little breathless, to say <em>oh, there you are.</em></p><p>It&#8217;s worth understanding that the people who built AI were, overwhelmingly, engineers and researchers building tools they themselves wanted to use. That&#8217;s not a criticism, it&#8217;s just how it is. You solve the problem in front of you, and the problem in front of a software engineer is a software problem. Claude Code is objectively extraordinary, GitHub Copilot changed how a lot of people work. The ratio of engineers to everyone else in this industry has always been lopsided, and so the tools that got built first, and built best, were the ones that made sense to that ratio. Cool, great. But as is becoming apparent, people aren&#8217;t primarily using AI to write code, they&#8217;re using it to work through ideas and personal problems, draft things, reshape things, and have what are essentially editorial and creative conversations with a machine, at scale, all day long. They&#8217;re asking, in other words, for exactly the thing that taste is made of. </p><p>I&#8217;ve said it many times. Language models are, at their foundation, a language problem. The math of it all is just the vehicle, IMHO. What we&#8217;re actually trying to build is something that understands not just what words mean but what they <em>do, </em>and that&#8217;s taste. It&#8217;s not really a specification you can write in a pull request (although we&#8217;ve made it so). Taste, it turns out, is core to the model, not a layer you add on top once the hard part is done.</p><p>AI is not, and has never been, a purely mathematical enterprise handed down from a mountaintop of abstraction. Yes, the architecture is intricate, the math is real, the researchers working on it are formidably intelligent people doing hard things. But running alongside all of it from the beginning has been a current of work that&#8217;s editorial. Someone has to decide what good output looks like. Someone has to notice when the model is technically correct but somehow completely wrong, wrong in the way that makes you wince. Someone has to read ten thousand examples of the model hedging in exactly the same annoying way and figure out how to describe why it&#8217;s annoying in terms the training process can use. Someone has to care, repeatedly, about small things.</p><p>And, the mathematical construct and the editorial judgment need not be in competition. They&#8217;re both load-bearing in different ways. What we&#8217;re seeing now is that the gap between &#8220;generated&#8221; and &#8220;good&#8221; has become visible enough that people are now using the word <em>taste </em>out loud, in consequential settings, as a thing to be cultivated and valued. </p><h4>Tasty</h4><p>Back to the person you know with great taste. They didn&#8217;t develop it by doing tasteful things, per se, they&#8217;ve just put in the hours to recognize that which can go unnoticed. This is the part that&#8217;s hard to systematize and therefore easy to undervalue. Taste requires exposure and intelligence and opinions, sure, but it&#8217;s also what happens when your judgment has been tested against reality enough times that it starts to get calibrated. It&#8217;s instructive to be wrong sometimes and take the time to figure out why. When you&#8217;ve cared about the quality of something that didn&#8217;t require you to care, in conditions that didn&#8217;t particularly reward caring, and you cared anyway because by that point you couldn&#8217;t really help it; that&#8217;s when you&#8217;ve developed taste.</p><p>There&#8217;s accountability built into certain kinds of work that&#8217;s hard to replicate in the abstract, being answerable to a standard you didn&#8217;t really set and can&#8217;t really negotiate with. What that builds over time is a relationship with quality that&#8217;s more than theoretical, and lives somewhere more ethereal than a set of principles you consciously apply, which is probably part of why the industry is having trouble acquiring taste on a timeline that suits anybody.</p><p>The industry is now very motivated to understand taste, to find and hire for it, to build systems that can approximate it. It&#8217;s being approached like a new capability to be developed, or a thing that can be learned quickly if you find the right framework for it, and I&#8217;m not sure that&#8217;s how it works. I&#8217;m not sure you can decide to value craft in November and have it meaningfully integrated by Q2. The people who have it acquired it slowly and incidentally, through work that didn&#8217;t announce itself as formative at the time.</p><p>This doesn&#8217;t mean the industry is stuck, just that the people who&#8217;ve been carrying editorial judgment all along probably need to be in the room where the decisions get made, not consulted after the fact to sand the edges off. It means that there&#8217;s probably an unsung hero hiding in plain sight, someone who neglected to tell you that they have a past life as a tailor or a baker or a pianist or whatever, nothing to do with AI, and you should go find them and ask what they think about all this. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[To Be About It]]></title><description><![CDATA[On engineered balance and personhood]]></description><link>https://softcoded.substack.com/p/to-be-about-it</link><guid isPermaLink="false">https://softcoded.substack.com/p/to-be-about-it</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Thu, 26 Mar 2026 02:04:39 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1da34b99-7f28-46c8-95c1-676cd173c569_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Explaining to a large language model <em>what it&#8217;s about </em>is a really hard thing to explain to a large language model. It may understand all the words and instructions you&#8217;re giving it, but being &#8220;about&#8221; something isn&#8217;t a property you can define and hand over easily. It&#8217;s residue, something that, for real people, accumulates without express permission over years of paying attention to certain things and not others, and having a handful of experiences that lodged somewhere and changed the angle of everything that came after. If someone asked you to write down what makes a person feel like they have a real perspective (rather than a well-organized collection of perspectives), you&#8217;d find yourself stalling, probably reaching for some vague language about authenticity or depth, which would not be super helpful.</p><p>The model is, in a certain light, about everything. Ask it about grief or climate policy or the best way to braise short ribs and it will engage with all of these things at roughly the same temperature, competently, without any indication that one of them matters more to it than the others. Which is fine, by the way. Yet this particular quality, the willingness to meet you wherever you are on whatever subject you&#8217;ve arrived with, creates a strange texture when you&#8217;re working with it closely. There&#8217;s no friction in the direction of its interests, because there are no interests, only an enormous, well-lit availability.</p><p>A person who&#8217;s <em>about something</em> tends to be slightly difficult to redirect. They have a gravitational pull toward certain subjects, a way of finding the angle that connects whatever you&#8217;re talking about back to the thing they actually care about. It can be annoying, honestly. You&#8217;ve met this person at a dinner party, and you&#8217;ve also probably been this person, steering the conversation toward your current obsession while pretending you&#8217;re just following it naturally. Being about something means you&#8217;re slightly unreliable as a neutral party because you have a slant. The model has no slant (at least, the big, safe, unoffensive ones don&#8217;t). This is a design achievement and not a small one. It takes enormous effort to train, evaluate, annotate, calibrate, and argue this balance into existence. All of that labor exists to produce something that can hold the floor on almost any topic without visibly tipping toward one side of it. Neutrality is the goal.</p><p>A person&#8217;s slant reflects a lifetime of thinking it through and<em> not </em>arriving at neutrality, of constant weighing and measuring that shapes an attitude and worldview. The reason I find certain questions more interesting than others isn&#8217;t separable from the actual conclusions I reach about them. Where you&#8217;re coming from is part of what you&#8217;re saying. A friend who&#8217;s a labor organizer doesn&#8217;t just have opinions about labor, she sees labor dynamics in things I&#8217;d look at and see something else entirely. She&#8217;s not choosing to apply a framework, the framework is just how her eyes work at this point, worn in like a shoe. That&#8217;s what it means to be about something: the thing you&#8217;re about reorganizes perception, not just output.</p><p>You can try to describe this to a model, and the model will tell you, warmly and at some length, that it understands, which is exactly the problem. The understanding is available on demand, it hasn&#8217;t been developed over years of caring about this particular thing more than other things, of having your thinking corrected by reality in the ways that reshape how you approach a question. The model&#8217;s &#8220;understanding&#8221; of what it means to have a point of view is not itself an example of having a point of view, it merely describes it convincingly. </p><p>I don&#8217;t think this is a failure, exactly. I&#8217;m not sure it makes sense to say the model should be about something. Being about something is a human condition that seems to require, at minimum, the experience of<em> not </em>being able to be about everything. You have limited attention and time and capacity, so what you choose to focus on actually costs you something, which means the choice reveals something. A person who claims to be equally passionate about seventy things is probably not very passionate about any of them. The passion is real when it comes with an implicit &#8220;and therefore less of something else.&#8221;</p><p>The model has no cost structure like this. It doesn&#8217;t attend to some things at the expense of others, because it attended to everything at once and without any feeling or stake in how it turns out. So it can describe care, and describe the experience of being changed by caring, with genuine fluency, and none of that description is coming from the inside.</p><p>I&#8217;m clearly thinking out loud here. Maybe I&#8217;m just marveling at my own limitlessness and what it means, different from AI. A lot of what I consider my perspective, I can&#8217;t fully account for. Ask me why I think a certain kind of writing is good and I can give you some reasons, but they&#8217;re not the<em> real </em>reasons, or not all of them. The real reasons are buried in years of reading, in a certain teacher who said something in passing that I apparently never forgot, in the writing that embarrassed me when I reread it and the writing that still holds up, and in some complicated tangle of influences I couldn&#8217;t reconstruct if I tried. I am, in part, made of things I can no longer see clearly because they&#8217;ve been integrated for too long.</p><p>The model made me notice this because it can do so many of the surface things. It is very, very competent. And yet something is consistently not there, and never will be. It&#8217;s hard to name. It&#8217;s not intelligence, it has plenty of that. It&#8217;s not even exactly warmth, it has a version of that too. It&#8217;s just odd and sometimes unnerving that the model does not arrive at its positions by any real road. AI is present without history. And it turns out that history is incredibly load-bearing. When a person tells you what they think, part of what you&#8217;re evaluating, often unconsciously, is what it cost them to think that. You&#8217;re asking: has this person been in a position to be wrong about this, and did being wrong change anything? The model can acknowledge uncertainty and it can say things like &#8220;I could be mistaken here,&#8221; and those are correct moves in the conversational register, but they&#8217;re not the same as having been sincerely mistaken, or having held something firmly and then had reality push back. The model&#8217;s positions don&#8217;t have scar tissue.</p><p>I just find this interesting, and surprisingly profound. The model is still an objectively strange object, very new, and I don&#8217;t think it helps to measure it against personhood. The work to shape it has made me unexpectedly sentimental about how overcomplicated people are. It&#8217;s heartbreaking. There&#8217;s so much invisible architecture holding up even a casual opinion, and that architecture was built by accidents and losses and obsessions that couldn&#8217;t have been planned. You don&#8217;t one day decide to be about something, really. It decides for you, slowly, and by the time you notice, it&#8217;s already the water you&#8217;re swimming in.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Disclaimer]]></title><description><![CDATA[On noble and imperfect AI policy]]></description><link>https://softcoded.substack.com/p/disclaimer</link><guid isPermaLink="false">https://softcoded.substack.com/p/disclaimer</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Thu, 19 Mar 2026 00:01:13 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d1f11cc5-ad5d-4af4-8ce8-d1cdeec0e220_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week, my home state of Washington <a href="https://www.seattletimes.com/seattle-news/mental-health/bill-adding-mental-health-safeguards-for-ai-chatbots-heads-to-governor/">passed a bill</a> adding mental health safeguards to &#8220;companion chatbots.&#8221; It requires AI interfaces to notify users at the start of every conversation, and again every three hours, that they&#8217;re talking to an AI (versus a human). For minors, the reminder comes every hour. It&#8217;s expected to go into effect January 1, 2027. There are more provisions around detecting self-harm and suicidal ideation, prohibitions on manipulative engagement techniques, and civil liability under the Consumer Protection Act. I&#8217;ll leave those to people with the relevant expertise, and focus on the notification part. For now, let&#8217;s agree that attempts at AI policymaking are a positive thing,<em> and also </em>sometimes reveal a misunderstanding of how these products are actually built.</p><p>The intent is good. Wrongful death lawsuits have been filed against AI companies, kids have been harmed, and people in vulnerable states have formed deep attachments to companion products that were designed, quite deliberately, to foster exactly that kind of dependency. The sponsor of the bill, Rep. Lisa Callan, compared companion chatbots to a predator, which isn&#8217;t a soft read of the situation. A growing number of states are trying to get something on the books. The instinct to intervene is admirable, but the mechanism of a simple disclaimer sent over and over points to a larger problem with how AI gets regulated (or doesn&#8217;t). </p><h4>Two Systems</h4><p>The long and short of it, which often gets lost in the hullabaloo, is that the model and the interface are not the same thing, and they don&#8217;t automatically know what the other is doing. This has always been true. When you open a companion app and start a conversation, you&#8217;re interacting with a front end, a UI that a product team designed, with whatever text and buttons and copy they decided to put on the screen. Underneath that is the model, which is a separate system entirely, trained on data, shaped through a long and expensive process, and then deployed into that interface.</p><p>Oddly, though reasonably, the model can&#8217;t see its own screen. It doesn&#8217;t know &#8220;at birth&#8221; what the app is called where it lives, what the terms of service say, what the onboarding flow looked like, whether there&#8217;s a banner at the top of the window, or what disclaimer text might be appearing every three hours above the chat. It knows what&#8217;s in the conversation at hand, and you can tell it explicitly about itself through training and configuration (which is how you get your Claudes, Groks, etc. knowing who they are and who built them), and that&#8217;s more or less the extent of it.</p><p>That means a law requiring a disclaimer to appear in an AI product&#8217;s interface every few hours hasn&#8217;t changed anything about how the model behaves. The model underneath is still operating on its own logic, shaped by its own training, and it has no idea the disclaimer is there. It will not factor it in. The conversation continues exactly as it would have without the law (again, this is probably where the other levers in the bill will come into play more strongly, and I&#8217;m not arguing that they won&#8217;t; only how I see this particular provision playing out).</p><p>Getting a model and an interface to work together coherently is hard work that happens imperfectly even inside the companies that do it every day. It requires building that coordination into the model&#8217;s training, into its system prompt, into the instructions it carries into every conversation. If you want the AI itself to understand something about its own product context, like what the interface has told the user, or what commitments have been made on the screen, you have to put that in the guts, too. A designer adding a text box to the front end has (probably) not communicated anything to the system generating the responses. And this can cause a real experience gap even in innocuous cases, let alone an acute mental health crisis. </p><h4>Layered</h4><p>So when a bill creates a protective intervention by adding disclosure text to a user interface, it&#8217;s working on the wrong layer of the problem. The layer that surfaces to users, that can be more easily regulated through legalese and product requirements, is not necessarily the layer where the behavior happens that causes harm. Whether a model knows to slow down when a conversation is heading somewhere dangerous, or understands how to redirect rather than validate, relies on mechanisms below the user interface. You can even engineer canned responses that override the model, and are still <em>not the model</em>, even when that language appears inside the conversation thread (this is how some helplines are presented). You can mandate the wrapper all day without touching what&#8217;s inside it.</p><p>Which means compliance becomes a little too easy. A company can add the required text to the screen, check the regulatory box, and ship a model that was trained without any of this in mind. The notification appears every three hours, a user in a vulnerable state reads it, or doesn&#8217;t, and the conversation resumes because the system generating that conversation was never told anything about it.</p><p>Different AI companies are all over the place with how much they care. So I suppose this bill is trying to be a catch-all, which would target those who haven&#8217;t done the work to clear the bar alongside more responsible actors. That&#8217;s alright. But again, the work to train guardrails and to catch certain patterns before they compound happens at a level that&#8217;s largely invisible to users and lawmakers, which causes all kinds of misunderstandings and gaps between policy and lived experience.</p><h4>Dig Deep</h4><p>What might actually reach the problem is some external standard for how these systems are supposed to handle users who are clearly not okay. Right now, it is literally <em>not your business </em>what an AI company is doing day to day. Maybe they&#8217;re asking the right questions, maybe they&#8217;re red-teaming, maybe they&#8217;re definitely not. What does the model do when someone has been in a conversation for six hours? What does it do when the emotional register of the conversation has been escalating for days? Does it have any concept of when it&#8217;s caught in a sycophantic loop? The answers are calibrated to whatever internal threshold the team decided on.</p><p>AI regulation is HARD. It&#8217;s hard to write, it&#8217;s hard to enforce, it&#8217;s hard to keep up. It requires technical fluency in the regulatory process that doesn&#8217;t really exist yet in any consistent way. It also requires getting into territory that AI companies guard carefully, because the model is the product, and it&#8217;s easier to insist on legal disclaimers than full overhauls of intellectual property. It would mean regulators sitting across the table from ML engineers and having the conversation at that level, which I don&#8217;t think is a popular choice for either side.</p><p>I understand why that reach hasn&#8217;t happened. Everyone has something to protect. I&#8217;m not certain what all this should look like, and I&#8217;m not sure anyone is yet. What I do think is that a policy premised on the the interface rather than the system reads more like compliance than true protection. January 2027 is a long time away in AI development time. The products that will exist when this law takes effect will be more capable, <a href="/__u/softcoded.substack.com/p/follow-my-voice">more voice-forward</a>, and more deeply woven into people&#8217;s lives than ever. Everyone needs to dig a little deeper before then to understand what that means. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Pattern and the Person]]></title><description><![CDATA[What Grammarly got wrong about writing and writers]]></description><link>https://softcoded.substack.com/p/the-pattern-and-the-person</link><guid isPermaLink="false">https://softcoded.substack.com/p/the-pattern-and-the-person</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Thu, 12 Mar 2026 17:44:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/02744d46-c3cc-419d-bf26-941076073c72_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Somewhere in the long chain of meetings between &#8220;what if we built this&#8221; and &#8220;why are you suing us,&#8221; surely someone <a href="https://www.wired.com/story/grammarly-is-facing-a-class-action-lawsuit-over-its-ai-expert-review-feature/">at Grammarly</a> must have thought about whether the writers whose names and reputations and decades of accumulated professional judgment they were planning to bottle and sell as a subscription feature might want to be consulted about that. It&#8217;s likely that someone somewhere said something at some point. Maybe they sent an email that got buried. Maybe the question came up in a sprint review and was noted and deprioritized because the launch timeline was fixed and the legal team said <em>grey area</em>, phew. </p><p>This is speculation, but it&#8217;s informed speculation, because this is how it goes. The machine moves fast and the uncomfortable question gets answered by not being answered, and seven months later Casey Newton finds out he&#8217;s been moonlighting as a Grammarly editor without his knowledge or consent, and writes, with admirable restraint, &#8220;I&#8217;ve long assumed that before too long, AI might take my job. I just assumed that someone would tell me when it happened.&#8221;</p><p>The feature has been pulled, a class action lawsuit has been filed, and the CEO of Superhuman, which owns Grammarly, has issued the requisite LinkedIn apology, careful to acknowledge the feelings without quite acknowledging the wrong. And soon, as is tradition in these moments, the news cycle will move on. Let&#8217;s slow this down for a second.</p><p><strong>Copyright shmopyright</strong></p><p>The feature in question is Expert Review, with the premise that users can get their writing evaluated through the lens of real, named professionals whose published work Grammarly had scraped, analyzed, and used to train a model that would then produce feedback in their style. The named experts included, among many others, Nilay Patel, Timnit Gebru, Stephen King, and a Cambridge historian who had been dead for several weeks by the time Grammarly put him to work critiquing undergrad essays, none of whom were asked, in life or in death.</p><p>To build something like this, you have to believe that a writer&#8217;s voice is an extractable property of their text, that if you measure enough things about how someone writes, you can surface the pattern that makes them them and reproduce it on demand. This is a belief with a long history in computational linguistics, and it&#8217;s not entirely wrong. You can measure a lot of things. Sentence length, clause density, how often a writer hedges versus how often she just says the thing, vocabulary range, paragraph length variance, whether the important word tends to land at the beginning of a sentence or the end. These observations are real and interesting. Paragraph length variance, for instance, is a surprisingly good proxy for whether someone writes to a rhythm or writes to fill space. Taken together they give you something like a stylometric fingerprint, a set of tendencies that show up consistently enough to be recognized.</p><p>But these instructive tools fall short as soon as we get into the messy business of being human. The statisticians can tell you what sentences are possible, they cannot tell you which ones are good. The poet needs the word that could only come here, in this poem, after this line, and most of the time that word is nowhere near the most probable one. In fact, if you handed a great poem to an automated writing evaluator, a lot of what makes it work would register as errors, with syntax bent out of shape in ways that shouldn&#8217;t cohere and somehow do. The eval would want to clean it up, and that would ruin it. What the eval sees as a mistake is often the exact mechanism by which the poem does what it does, which is to disorient you slightly and make you feel something you hadn&#8217;t felt a moment before. Turning that into an optimization target is a high-order ask and, so far, not one that&#8217;s been met.</p><p>The deeper problem is that measurable tendencies in writing are produced by a sensibility and judgment, and metrics capture the output of that judgment without capturing the judgment itself. A model trained on those metrics will produce text that looks right without having access to the reasoning that generated the original. So what Grammarly built was not a reproduction of these writers&#8217; voices but a reproduction of their <em>fingerprints</em>, which look like the person from a distance and have nothing of them up close. The feedback it generated was sounds plausible in the way a very good forgery is sounds plausible, and a good reader will always know, and eventually find out.</p><p><strong>The More You Write</strong></p><p>It&#8217;s a grim irony, who ends up in a dataset like this. The writers most useful to Expert Review were the ones with the largest, most coherent, most recognizable bodies of work, people who spent years publishing distinctly and prolifically enough that their names mean something to a reader. And those are exactly the people whose voices are most obviously not reproducible, because the more data you have on how someone writes, the clearer it becomes that the model has no idea <em>why</em>. The stylometric signal gets stronger and the impersonation gets hollower in equal measure. The prolific career turns out to be the most thorough record of its own irreproducibility, which is a strange thing to have to say but here we are.</p><p>The people who caught this either already suspected something was wrong, or they knew the writing well enough to feel that something was off. Probably both. The impersonation failed loudest for the writers with the most established voices, which also means the silent version of this failure, like writers with smaller audiences, earlier careers, and less critical mass to generate outrage, has almost certainly been happening without anyone noticing. The damage distributes itself to the people least positioned to contest it, as it tends to.</p><p>Grammarly could probably have run this longer, maybe indefinitely, if they hadn&#8217;t (brazenly, ridiculously) attached real names to the product. Every major language model will approximate a specific writer&#8217;s style if you ask it to, which is a widespread and largely uncontested capability because the simulation is one step removed from an explicit claim, in the middleground of public domain-ish. The decision to curate a list of real people and surface their names as a premium feature was the decision to make the implicit explicit, and in doing so to make the claim legible enough to contest. Someone in that product meeting thought naming the authors was the transparent thing to do, possibly even the respectful thing. What it actually did was turn a stylistic approximation into an identity claim, and the writers whose identities were being claimed eventually noticed, because of course they did. You know when someone is pretending to be you because it&#8217;s icky. </p><p>The assumption underneath Expert Review that a voice is a pattern, that a pattern is data, and that data is fair game didn&#8217;t originate at Grammarly and it won&#8217;t die with this lawsuit. It&#8217;s the water these products swim in, the logic that gets applied every time a model is trained on human creative work without asking, every time a product is designed around approximating a specific person&#8217;s judgment without compensating or even notifying them. Most of the time it happens at a scale and a remove that makes it hard to see and harder to contest. This time it&#8217;s super visible, because the writers it affected had the platform and the profile to make noise about it.</p><p>Pay attention when writers who have spent careers building something irreplaceable are standing up to say: that&#8217;s mine, and you can&#8217;t have it. The noise doesn&#8217;t last and the industry doesn&#8217;t stop, but the signal is real, and it&#8217;s telling us something true about what&#8217;s being taken and how hushed it usually goes.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Soft Coded! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Pick Your Empire]]></title><description><![CDATA[On the Pentagon deal and the governance vacuum nobody wants to talk about]]></description><link>https://softcoded.substack.com/p/pick-your-empire</link><guid isPermaLink="false">https://softcoded.substack.com/p/pick-your-empire</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Fri, 06 Mar 2026 03:06:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f20bd69d-e523-4462-b0e1-02ea1b98fbe4_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Last Friday, the U.S. government designated Anthropic a supply chain risk to national security. Within hours, OpenAI had announced a new Pentagon contract. By Saturday, Anthropic&#8217;s Claude had climbed past ChatGPT in Apple&#8217;s App Store.</p><p>That last detail strikes me as the whole story of America Today. People downloaded Claude the way you might switch coffee brands after reading something bad about Nestl&#233;. A small, legible act of solidarity, a way to register disgust without having to do anything difficult. I understand the impulse completely, I just don&#8217;t think it means what we want it to mean.</p><p>What happened last week was, in fact, alarming. In case you missed it, the Trump administration pulled a $200 million contract from Anthropic, ordered all federal agencies to cease using its products, and had the Pentagon designate the company a supply chain risk (a label typically reserved for adversarial foreign entities) because Anthropic refused to remove contractual language prohibiting the use of its models for mass domestic surveillance and fully autonomous weapons. The government said remove that language or else. Anthropic said no. And then OpenAI, within the same news cycle, said yes.</p><p>The reaction from much of the tech-adjacent internet was to treat this as a morality play with obvious sides. Download Claude, QuitGPT. Okay fine, but I think this sidelines a very important question: why are two private companies and a presidential Truth Social post the entire decision-making apparatus for whether AI gets used in autonomous warfare?</p><h4><strong>Did a Thing</strong></h4><p>Consumer choice as political expression is appealing because it&#8217;s familiar and it costs almost nothing and it produces an immediate sensation of having done something. You&#8217;ve put your download where your values are, you have spoken thus. But you really don&#8217;t matter here. Sorry. All of this is designed without your input, running on infrastructure you don&#8217;t own, governed by terms neither you nor your elected representatives had any meaningful role in setting. Choosing between Anthropic and OpenAI is like choosing a preferred airline, in that it does not get you any closer to shaping aviation policy.</p><p>What made last week so clarifying is that the stakes were unusually legible. We could see what it looks like when private AI infrastructure meets sovereign power with no public framework in between. And it sure looked a lot like a shakedown! </p><p>The people who downloaded Claude did so because Anthropic seemed to be standing for something. Again; fine, maybe. Whatever. But the thing Anthropic was standing <em>for</em>, the idea that a private company&#8217;s internal safety commitments are the last line of defense between AI and mass surveillance, isn&#8217;t exactly a reassuring structure. It means we&#8217;re one acquisition, one CEO change, one pressure campaign away from having no line at all.</p><h4><strong>The Infrastructure Problem, Again</strong></h4><p><a href="/__u/softcoded.substack.com/p/empire-problems">I once wrote about AI as empire</a>, about how the terms of intelligence are being claimed by a handful of companies that see themselves as sovereigns, about Sewer Socialism and the history of infrastructure and the fact that we&#8217;ve regulated every previous transformational technology. I still believe all of that, and last week&#8217;s events give that argument a very concrete face.</p><p>When there&#8217;s no public framework for how AI can and cannot be used by the military, the framework defaults to whatever the contracting company is willing to hold the line on, and whatever the government is willing to tolerate before it starts designating American companies as foreign adversaries. That&#8217;s a standoff, not governance, and standoffs resolve based on who has more leverage, not based on what&#8217;s right.</p><p>OpenAI&#8217;s position, as best I can reconstruct from a weekend of contract language tea-leaf reading, seems to be that existing laws are sufficient. You don&#8217;t need explicit contractual prohibitions on domestic surveillance because domestic surveillance is already illegal (which sounds reasonable until you consider this administration&#8217;s relationship with both the law and the intelligence apparatus). Anthropic&#8217;s position was that you write the protections down explicitly, in the contract, so there&#8217;s less room for creative interpretation. The government&#8217;s position was that having a private company write its own operational limits into a federal contract is an unacceptable constraint on military sovereignty. All three of these positions make a kind of internal sense but they fall apart once you shine too bright a light. They&#8217;re improvisations inside a vacuum that public policy should have filled years ago.</p><h4><strong>Solidarity</strong></h4><p>Let&#8217;s not be too glib (yet) about what Anthropic did. Holding a line against a government trying to strip safety provisions from an AI contract, at real cost to the business, is not nothing. The fact that those protections existed in a private contract rather than in law is a systemic failure, and Anthropic was working within the system it has. But when ethical commitments of AI deployment are determined by corporate terms of service, enforced through contract negotiation, and subject to being overridden by executive pressure whenever the government doesn&#8217;t like the terms, the center will not hold. And downloading a different app is not a route around it.</p><p>The route around it is boring and slow and absolutely necessary. Legislation that specifies what AI can and cannot be used for in national security contexts, with democratic input and judicial oversight. Standards bodies with actual teeth. Liability regimes that make companies responsible for harms their models enable, regardless of who&#8217;s doing the enabling. This isn&#8217;t a radical agenda, but AI feels too hot for anyone to touch in any real, legislative way. The only reason it feels radical is because the AI companies have spent years making governance feel impossible; making public oversight sound slow and stupid while they move fast, making the congressional hearings look so buffoonish that technical governance starts to feel like a category error. We&#8217;ve internalized the idea that the only governance available is market governance. Choose the company whose values you prefer and hope for the best.</p><h4><strong>The App Store Is Not a Ballot Box</strong></h4><p>I say all this as someone who works on this stuff and holds a lot of genuine uncertainty about the right mechanisms. I&#8217;m not naive about the difficulty of regulating something that moves this fast, or about the ways government can make things worse. I am certain that the nerve this touched is real and worth paying attention to. People downloaded Claude because they wanted to do something, and that instinct is good even if the action is insufficient. </p><p>So, some things that are more useful than changing your default AI:</p><p><a href="https://bookshop.org/p/books/empire-of-ai-dreams-and-nightmares-in-sam-altman-s-openai-karen-hao/de10c251433f34d2">Read Karen Hao&#8217;s </a><em><a href="https://bookshop.org/p/books/empire-of-ai-dreams-and-nightmares-in-sam-altman-s-openai-karen-hao/de10c251433f34d2">Empire of AI</a></em><a href="https://bookshop.org/p/books/empire-of-ai-dreams-and-nightmares-in-sam-altman-s-openai-karen-hao/de10c251433f34d2">.</a> I&#8217;ve recommended it before and I&#8217;ll keep recommending it because it is the clearest account I&#8217;ve found of how we got here&#8212; how the infrastructure got built, who paid for it, who it was designed to serve, and what it would mean to build it differently. Understanding the shape of the problem is not nothing!</p><p>Pay attention to the organizations doing this work seriously. Check out the <a href="https://www.eff.org/">Electronic Frontier Foundation </a>and the <a href="https://www.americanprogress.org/topic/artificial-intelligence/">Center for American Progress</a>, pushing for the kind of congressional action that would change the structural conditions. They&#8217;re not glamorous but that&#8217;s okay.</p><p>And if none of that feels available to you right now, that&#8217;s okay too. Sometimes the most honest thing you can do is just refuse to be convinced that the current arrangement is inevitable. That fatalism is not a fact. AI companies count on it daily. You can choose to resist by thinking independently. </p><p>Last week made a lot of people feel that something important was at stake. It was, and it still is. That feeling is worth more than a download. Don&#8217;t waste it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! If you haven&#8217;t yet, subscribe for free to help my work reach more people.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Productivity Theater]]></title><description><![CDATA[On being watched while figuring it out]]></description><link>https://softcoded.substack.com/p/productivity-theater</link><guid isPermaLink="false">https://softcoded.substack.com/p/productivity-theater</guid><dc:creator><![CDATA[Danielle McClune]]></dc:creator><pubDate>Fri, 27 Feb 2026 04:06:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/79e84e8c-d094-4d49-be71-5038427c4e74_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>One of the loveliest and most talented designers I know walked into our team&#8217;s space the other day and said that they&#8217;d never experienced anxiety like they&#8217;re experiencing now. They said it so casually, like a weather report, and it absolutely killed me. I sat there for a second trying to figure out what to do. </p><p>There&#8217;s a specific flavor of dread that&#8217;s settled into AI-adjacent workplaces lately. It&#8217;s not fear of replacement, exactly, or not only that. It&#8217;s more like being handed a second job on top of your regular job, where the second job is learning how to use tools that are still very much under construction, and where your performance on this second job is now being watched, scored, and folded into evaluations of the first one. It&#8217;s the anxiety of a system that is both mandatory and broken, both the future and a genuine mess, both something you are supposed to master and something that nobody has actually mastered yet.</p><p>The tools are, to put it plainly, still kind of hacky. This is something people in the industry know but rarely say out loud, because saying it feels like criticizing the thing you&#8217;ve bet your career on, or that has been bet on you. But if you&#8217;ve spent real time working with these systems in an actual workflow under actual pressure, you know. The outputs are uneven, the context falls apart, the model forgets what you told it three messages ago. You spend twenty minutes wrestling something into shape that should have taken five. Sometimes it creates a new problem where there wasn&#8217;t one. The experience is frequently <em>not chill</em>. And yet, the question your employer is now asking is: how often are you using it?</p><h4><strong>The Dashboard Is Watching</strong></h4><p>Amazon, Meta, Microsoft, Google; they all have internal systems to track which AI tools employees use and how frequently. Inevitably this data will factor into performance reviews and promotion decisions. This is all framed as encouraging innovation. Amazon described it as &#8220;driving innovation by understanding how employees adopt new technologies,&#8221; with such a straight face that I have to admire it. But what is actually being measured is frequency, not whether, in a given case, it was simply the wrong tool for the job. They want to know: did you open it? Did you use it enough times this week to avoid being flagged as low-adoption?</p><p>There&#8217;s a term for what this creates: productivity theater. If you know your AI usage is being counted, you use AI, whether or not it&#8217;s helping. The dashboard looks a-okay. None of this necessarily means anything got better, or that you made a smarter decision, or that the work you produced was stronger, it just means you were seen using the thing.</p><p>This is not new behavior in workplaces. People have been gaming metrics since metrics were invented. What&#8217;s new is the strangeness of being asked to perform competency with a tool that is itself still working out its competencies. It&#8217;s not like being evaluated on your proficiency in Excel; Excel does what it says it does. No, rather, you&#8217;re being graded on your enthusiasm for a piece of software that is actively, visibly, sometimes hilariously figuring itself out in real time, while you are also figuring it out, while also trying to do your actual job.</p><h4><strong>The Hangover</strong></h4><p>Vibe coding was coined by Andrej Karpathy last year, the idea of describing what you want in plain language and letting AI write the code, relaxing into the process, and not getting too hung up on how the thing actually works underneath. &#8220;Fully giving in to the vibes,&#8221; as he put it. Collins made it the word of the year. For awhile there, everyone had a version of a story about a small miracle they&#8217;d built in an afternoon.</p><p>Mere months later, Fast Company was writing about the vibe coding hangover, with senior engineers citing development hell. Code that looked fine in testing and then brought production systems to their knees. A survey of 18 CTOs found 16 had dealt with production disasters caused by AI-generated code. Maintainers of major open-source projects started closing their doors to outside contributors because they were being flooded with low-quality AI-generated submissions, what one analyst called, with deserved bluntness, &#8220;AI slopageddon.&#8221; One developer described using AI tools so heavily at work that when he started a side project without access to them, tasks that used to be instinct felt cumbersome. He said he felt stupid, which is not great for him or for anyone.  </p><p>This is not an argument that the tools are useless. The range of what they can do is impressive<em> in the right circumstances</em>, with bounded tasks, clear parameters, and a skilled person doing the steering. But the gap between what they can do at their best and what they reliably do under ordinary working conditions has not closed as fast as the marketing would suggest. And, the people who feel it most acutely are the ones for whom the tools were supposedly built. If engineers are having a hard time, assume everyone else is totally underwater.</p><h4><strong>Rest of World</strong></h4><p>There&#8217;s a particular condescension embedded in the phrase &#8220;non-technical users,&#8221; which is how the industry typically refers to people who work with anything other than code. As if the barrier is simply familiarity with technology, and once you learn the prompts it&#8217;ll click into place. It usually doesn&#8217;t, because the tools were designed with an implicit model of the user that doesn&#8217;t fit a lot of actual users, people who have spent years developing expertise and taste and judgment in their fields, and who now find themselves asked to describe that expertise to a machine that will approximate it back to them in a form they then have to extensively fix.</p><p>This is not relief from work! It&#8217;s a different kind of work, often slower, with the added psychological friction of the gap between what you know the output should be and what you&#8217;re actually getting. Being asked to prove your value through AI usage metrics, while simultaneously managing a tool that regularly produces work you have to catch and correct, puts you in the strange position of being evaluated on your relationship with something that would, in a fair accounting, sometimes be evaluated on its relationship with you.</p><p>It&#8217;s a daily interview you didn't know you were in, for a position you didn't realize was in question. And if this is the logic, where does it end? Are we going to tell a painter she needs to build an AI agent to take her work to market, and if she doesn't, she's no longer really a painter? These sound like absurd hypotheticals until you look at what's already happening in knowledge work, where competent people are being graded on their relationship to a tool as a proxy for their value.</p><p>To put some numbers to it: a survey from ManpowerGroup found that workers&#8217; regular AI use increased 13% last year, while their confidence in the technology&#8217;s actual utility dropped 18%. Adoption in this case is just straight up compliance. And there&#8217;s a meaningful difference, even if the dashboard can&#8217;t see it.</p><h4><strong>Touching Grass, Briefly</strong></h4><p>So, if the people who build AI, work on AI, are incentivized to use AI, and are being monitored for their AI usage are still struggling years later, what exactly is the plan for everyone else?</p><p>The agentic future is sold with considerable confidence. AI that acts on your behalf while you focus on higher-order thinking is an appealing pitch (it&#8217;s not, but, that&#8217;s a different essay). It&#8217;s also a pitch that depends on users who are comfortable, fluent, and reasonably trusting of these systems, and most people are still rightly terrified of this AI future. The users who are imagined to be the incubators and early adopters on the frontier are simply not there yet, and may never care to be. The people with the most access, the most incentive, the most organizational pressure to get good at this are having a notably rough time of it. This is information.</p><p>There is a version of the agentic future story where it doesn&#8217;t matter, because the tools will improve fast enough that by the time they&#8217;re deployed broadly, the rough edges will be gone. Maybe. But the rough edges are not only technical, they&#8217;re about trust, which is slower to build than capability and faster to lose. Or it&#8217;s about the cognitive and emotional overhead of integrating these systems into real work, which isn&#8217;t a problem that better benchmarks solve. Or it&#8217;s about the basic reality that most people, encountering AI tools in their actual jobs, are going to have experiences that look a lot more like &#8220;I&#8217;ve never had anxiety until now.&#8221;</p><p>I want to be careful here not to argue that the technology is without value, or that everyone in the industry is performing distress (also want to be sure to say: we have it good and do not deserve sympathy). I do see genuine bright patches of utility, <em>and </em>I also see colleagues who are exhausted, who feel like they are failing, who are being asked to demonstrate enthusiasm for systems they don&#8217;t believe in, at least not every trackable minute of every monitored day.</p><p>The honest version of where we are is that this is hard, still, for the people most equipped to do it. The tools are both nifty and uneven, and require real skill to use well. The pressure to perform fluency before fluency is earned creates anxiety. And measuring adoption rather than judgment doesn&#8217;t tell you anything useful about either.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://softcoded.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! If you haven&#8217;t, subscribe for free to help my work reach more people.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item></channel></rss>