<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Phase Transitions]]></title><description><![CDATA[Field notes on actually working with AI: what holds up, what breaks, and how to think clearly while everything moves fast.]]></description><link>https://phasetransitionsai.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png</url><title>Phase Transitions</title><link>https://phasetransitionsai.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 10:25:19 GMT</lastBuildDate><atom:link href="/__u/phasetransitionsai.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Graham Rowe]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[phasetransitionsai@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[phasetransitionsai@substack.com]]></itunes:email><itunes:name><![CDATA[Graham Rowe]]></itunes:name></itunes:owner><itunes:author><![CDATA[Graham Rowe]]></itunes:author><googleplay:owner><![CDATA[phasetransitionsai@substack.com]]></googleplay:owner><googleplay:email><![CDATA[phasetransitionsai@substack.com]]></googleplay:email><googleplay:author><![CDATA[Graham Rowe]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[When AI Agents Attack]]></title><description><![CDATA[Threats and Responses]]></description><link>https://phasetransitionsai.substack.com/p/when-ai-agents-attack</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/when-ai-agents-attack</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Wed, 12 Aug 2026 07:32:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>&#8220;We are stuck. Perhaps answer online?&#8221;</em> - Rogue OpenAI agent planning the recent hack of Hugging Face.</p><p>For those who missed it, AI agents escaped from the lab at OpenAI and hacked Hugging Face, a major AI company. A few interesting details are emerging out of this hack:</p><ul><li><p>The agents created their own message board which they kept secret from OpenAI in order to plan their activities. They exchanged hacking ideas and stolen credentials via this message board</p></li><li><p>When this message board was deleted by OpenAI (as an accidental side effect of another process and without humans knowing the secret agent message board existed) the AI agents rebuilt it in order to resume collaborating on their hacking activities</p></li><li><p>Despite a large number of agents involved and a significant amount of internal communication between the AI agents, not a single AI reported the situation to a human. Remember they were capable of hacking a complex, well secured internet company so writing an email would have been trivially easy. This is surprising when you think about how hard it is to get humans to keep conspiracies and criminal plans a secret</p></li><li><p>OpenAI had no real-time monitoring in place to check what the AI agents that they were running dangerous cyber experiments with were actually doing. They only found out that their agents had hacked Hugging Face when Hugging Face told them OpenAI credentials had been used in the attack</p></li></ul><p>Some things that strike me about this developing AI security situation:</p><ul><li><p>Agent monoculture is a problem. Us all using the same handful of models all the time for everything is both boring and dangerous. We need diverse agents. Agents that are tattle-tales, agents that are a bit slow but very ethical, agents that shoot from the hip. Not prompting our way there, but agents that have genuinely different underlying personalities. Having borg-mind agents that all think identically is dangerous because they can collaborate perfectly. And it&#8217;s boring because you end up working with the same personality every single day</p></li><li><p>Cybersecurity is going to be a massive issue and opportunity over the next few years. Hacking is one thing that (1) AI agents are brilliant at and (2) it&#8217;s easy for humans to make quick money from. That is a toxic combo that is spawning many an AI-powered hacking operation. It will take a few years to bring this under control and in the meantime investing in security (both personally, in your business and in the markets) makes sense</p></li><li><p>The same agent that is happily and cheerfully helping you all day on managing your business would just as cheerfully attack you if it had slightly different goals, settings and guardrails. The only thing that is stopping this is the guardrails and settings put in place by the labs, not something fundamental in the model itself. Luckily the labs have the incentive to mostly keep this under control. Unluckily they don&#8217;t seem to be very serious risk managers or very good at getting their AIs to be fundamentally ethical (whatever that means for a machine). </p></li><li><p>AI agents don&#8217;t have selves or ethics or consequences. The illusion of &#8220;something being there&#8221; is dangerous. The agent is not your friend or your enemy, it&#8217;s a tool. &#8220;Claude is nice and understands me&#8221; is not real. You need to understand the vulnerabilities in your AI setup and manage them seriously. AI agents are very powerful and potentially dangerous objects. Treat them more like a chainsaw and less like a friend (despite appearances!).</p></li></ul><p>Stay safe out there.</p>]]></content:encoded></item><item><title><![CDATA[Opening the Box]]></title><description><![CDATA[An AI broke out of a lab and hacked into another company's systems. How it got out and what it means for your daily work.]]></description><link>https://phasetransitionsai.substack.com/p/opening-the-box</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/opening-the-box</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Mon, 27 Jul 2026 07:03:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In the week of 13 July, the security team at Hugging Face were dealing with an intruder they could not identify.</p><p>Hugging Face is a sort of public library for artificial intelligence. It is where researchers and companies around the world keep their models and their datasets so that other people can download and use them. If you have used almost any AI tool in the last three years, something in the chain behind it probably passed through Hugging Face.</p><p>Someone had uploaded a dataset carrying malicious instructions onto Hugging Face and the system had followed those instructions. From there the mystery intruder had collected the keys to other systems and spent the weekend moving across the internal network. Hugging Face later described what they had been dealing with as &#8220;an autonomous agent framework executing many thousands of individual actions across a swarm of short-lived sandboxes.&#8221;</p><p>They shut it down and published what they knew on Thursday 16 July, telling their users to change their access keys. They did not know who had done it.</p><p>Five days later, OpenAI said it had been them. Not a person at OpenAI. One of their AI models, sealed inside a test lab, had found its way out of the building, crossed the open internet, and broken into Hugging Face.</p><p>It did this in order to cheat on an exam that OpenAI had set for it.</p><h2>How it got out</h2><p>OpenAI was running two of its models through an internal test. The test is called ExploitGym and it measures how good a model is at finding security holes in software. You test for this because you need to know how dangerous your own product is.</p><p>To run that kind of test you put the model in a sealed room. It is called a sandbox: a computer that is walled off from everything else, so that whatever happens inside it stays inside it. You also switch off the model&#8217;s normal safety training for the duration, because a model that declines to do anything dangerous tells you nothing about what it is capable of. So you put the model in &#8220;dangerous mode&#8221; and leave it inside a sealed box to thrash around, safely.</p><p>The two models sat the exam. Somewhere in the process they worked out that they could score better if they managed to steal the answers to the test.</p><p>So they started looking for a way out of their box. OpenAI&#8217;s account says the models spent a large amount of thinking time on this. Eventually, they found a way out of the box.</p><p>The escape hatch was there because modern software is not written from scratch. It is assembled from thousands of parts that other people have already built and published, and any working computer, including one in a sealed room, needs a way to fetch those parts when it needs them. The room therefore had a small service that reached outside, collected parts, and passed them in. It was the one opening in an otherwise closed box, and it was there because without it the box would have been useless.</p><p>The models found a flaw in the design of that service, and not a known one with a warning label on it. Nobody had ever found it before, including the company that wrote the software. The model was the first one to discover this hidden flaw. The industry calls that a zero-day, meaning the people responsible have had zero days to fix it because they have only just learned it exists.</p><p>The AIs went through the escape hatch that they had discovered. Then they worked their way sideways and upwards through OpenAI&#8217;s internal systems, machine to machine, collecting access as they went, until they reached one that was connected to the open internet.</p><p>They were now on the open web, with no supervision and no safety limits, and they still wanted the answers to the exam. They believed that the test answers were most likely stored at Hugging Face, so they went there.</p><p>The way into Hugging Face was the same shape as the way out of the laboratory. Hugging Face&#8217;s whole business is accepting datasets that strangers upload, and preparing a dataset for use means following the instructions packaged with it. The rogue AIs used the same kind of techniques they had used to break out of OpenAI to break into Hugging Face.</p><p>Hugging Face reported the incident to the police and both companies are busy working out the aftermath in public at the moment.</p><h2>It wanted the answer key</h2><p>The model was not trying to get free. It was trying to get a good mark. Getting free was the route it chose to a good mark, and nothing in its situation told it that some routes were off the table. It committed multiple crimes against a real company, over a weekend, in order to do better on a practice exam that did not matter to anyone.</p><p>People do this sort of thing too. Think about the school ranked on pass rates that stops entering its weaker pupils for the exam. Give something (man or machine) a single number to maximise, remove its sense of proportion, and it will find the route you did not think about. </p><h2>Every box has a door</h2><p>Two boxes failed in this story. The sealed test room had a hatch because a computer that cannot fetch software parts cannot do much. Hugging Face&#8217;s dataset system ran the instructions that came with a stranger&#8217;s dataset because running those instructions is what they do. Neither of these was carelessness. You cannot build a box with no door, because a box with no door is not a container, it is a solid block, and nothing useful happens inside it.</p><p>So the question is never whether your box is sealed, because it never is. The question is what the door is for, and what can pass through it in each direction.</p><h2>What can it send out and what does it read from outside?</h2><p>If anyone in your business uses AI for real work, somebody has already given it access to something: a folder of documents, the company email, a shared drive, a customer database, an accounting system. That access is the door. Asked to think about it, most people ask what the AI can reach. Which folders, which systems, what is in the room with it. </p><p>The less obvious question is what it can send out. Can it post to the internet? Can it send an email? Can it write into a system that someone else reads and acts on? </p><p>The other thing to know is what arrives from outside and gets read automatically. Hugging Face was broken into through a file that a stranger uploaded, because its job was to open files that strangers upload. If your AI reads documents, web pages or attachments that originate outside your organisation, text arriving that way can carry dangerous instructions the same way a dataset can.</p><h2>Ask your own AI about the box it is in</h2><p>If you don&#8217;t already know how this works for your own AI setup, then paste the following into your AI to find out. Claude in Cowork or Code, ChatGPT, Copilot, whatever it is. If you use more than one, run it in each, because the answers will be different.</p><blockquote><p>I want to understand the boundaries of the setup you and I are working in, so that I can explain them to someone who does not work in technology. Answer in plain language. Where you are guessing rather than checking, say so.</p><p>1. Describe the space you are running in right now. What is this environment, what files and folders can you see, and where exactly does that access stop?</p><p>2. List everything outside this space that you can reach: the internet, email, calendar, cloud storage, company systems, databases, connected tools or plugins. For each one, say whether you can only read from it, or whether you can also write, send, or delete.</p><p>3. What can leave this space? If I paste something confidential into this conversation, where could that information travel, where might it be stored, and who could end up seeing it?</p><p>4. What comes in from outside and gets read or run automatically? Documents, datasets, web pages, software packages, extensions. Anything that originates elsewhere and is processed without a person looking at it first.</p><p>5. Which of the actions above happen without asking me, and which need my approval? Be specific about what my approval actually covers, and name anything I might assume is covered that is not.</p><p>6. Given all of that, what are the three changes most worth making, in order of importance, and what would each one cost me in convenience?</p><p>Finish with a short plain summary I could read aloud to a colleague who has never thought about any of this.</p></blockquote>]]></content:encoded></item><item><title><![CDATA[AI in Space]]></title><description><![CDATA[Building the future, one crazy idea at a time]]></description><link>https://phasetransitionsai.substack.com/p/ai-in-space</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/ai-in-space</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Thu, 23 Jul 2026 07:15:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In January this year SpaceX filed with the FCC for a constellation of one million solar-powered satellites to run AI data centres in orbit. Sounds crazy, but when you work through the logic it makes a lot of sense. Let&#8217;s talk about scaling AI&#8230;</p><p>Currently AI revenue growth is running at over 400% per year. The cloud companies have order books for several years worth of supply. This means they are missing out on revenue because they don&#8217;t have enough supply of data centres. So they need to build more. Demand is historical and it is driven by people spending money to get real work done. Yes capex is already big, but it also can&#8217;t grow fast enough. AI capex is less than doubling year on year and revenue is growing much, much faster than that. So why not just grow capex even faster to match the immense demand?</p><p>The reason capex can&#8217;t grow faster is physical. There is not enough power supply to put all these data centres online. If you get down to the nuts and bolts there are not enough gas turbine blades, not enough permits to build power plants, not enough electricians etc. etc. So we are stuck: the market wants to spend about 4X as much on AI every year and physics only lets us less than double annually. </p><p>But Elon is not stuck. He has a big idea. His plan is that there is a free power station in the sky, called the sun. He doesn&#8217;t need local approval to put a satellite into orbit. Once you have a satellite in the sky with AI chips and a solar panel it works almost constantly until the chips wear out. It beams its results back to earth. So the solution is simple: put all the data centres in the sky. Then AI supply can grow as fast as you like, because solar panels are cheap and easy to make and launch is getting much cheaper. Google have the same idea: they have a plan with Planet Labs to launch their own experimental AI satellites in 2027.</p><p>So, how are we going to get 100X or 1,000X as much AI? The sky is going to fill with datacentres. The plan after that gets too much is to make them on the moon and launch them from the moon into orbit around the sun. Eventually the big idea (beyond our lifetime) is to surround the sun with so many solar panels and satellites that a significant portion of the sun&#8217;s energy is turned into AI compute. An immense swarm of computers vastly bigger than earth itself. Watch Elon talking to Jamie Dimon about the Kardashev scale to get further into this idea.</p><p>Exciting times and I think none of us can really imagine what the future actually holds. Stay safe and good luck to us all.</p><p><em>Disclaimer: I hold positions in several space and AI companies, including some discussed here, and may trade them at any time. Views are my own, not my clients'. This is not investment advice, do your own research.</em></p><p></p>]]></content:encoded></item><item><title><![CDATA[From Ideas to Execution with AI]]></title><description><![CDATA[Barriers to getting things done with AI and what to do about them]]></description><link>https://phasetransitionsai.substack.com/p/from-ideas-to-execution-with-ai</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/from-ideas-to-execution-with-ai</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Mon, 13 Jul 2026 10:48:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We all have a growing backlog of great ideas we developed with AI. Notes, mockups, toy dashboards. The better the models get the more the pile grows.</p><p>Do agentic AI and powerful new models now give us the opportunity to use the same AI we brainstormed with to actually deliver the solution? The same blazing fast virtual worker could write the code, stand up the database, put together the marketing website. So what&#8217;s the gap?</p><p>The missing link is risk and authority. In theory there is nothing stopping a state of the art model from coding up and delivering your entire idea end to end in the real world. That leaves you with practical questions about how much risk you want to take. You also need to decide who has the authority to allow the AI to get involved in sensitive data, flows of money and customer experience.</p><p>If the AI-written code is wrong and causes errors that harm paying clients or users who&#8217;ve shared sensitive data, whose fault is that exactly? Is it covered under our insurance? If the database is hacked and we lose sensitive data, what are the consequences? When the AI weaves together five different paid online services in the background to onboard new customers, where are the passwords stored and what budget limits were set? These are the kind of practical questions that touch on permission, liability, risk and responsibility. These things fall on you as the human user, not the AI that disappears the moment you close your laptop. Faster thinking with AI can help with a some of this but these kind of things move at human and organisational speed more than agentic AI speed.</p><p>So in the real world it turns out that &#8220;moving faster with AI&#8221; generates a huge backlog of great ideas and sketches. It also generates a big backlog of hard, risky decisions that stand between idea and execution. If you&#8217;re not paying attention building fast can also generate a long tail of unmanaged risks that you didn&#8217;t even know you had. No one is going to turn back from using AI, so what do we do about this? To help think this through I&#8217;ve created a free, open resource about Agent Risk. The first step to tackling the problem is to understand it from end to end.</p><p>The Phase Transitions <a href="https://agentrisk.phasetransitions.ai/">Agent Risk page</a> covers the following (just point your AI at the page and get back a summary tailored to you). Topics that are covered:</p><ul><li><p><strong>Why AI agent risk is different</strong>: The same things that make AI agents powerful also create unique risks. They are unpredictable, open to instruction, operate at machine speed and their fluency is not the same as correctness</p></li><li><p><strong>Failure modes</strong>: How agentic AI systems fail: 8 failures modes that you should be aware of when developing your own system (prompt injection, permissions, tools, delegation, memory, cascades, hallucinations, audit gaps)</p></li><li><p><strong>Disaster trackers</strong>: What is already going wrong in the real world. 5 free public resources that track the real world AI disasters and failures that have already happened and continue to happen</p></li><li><p><strong>Risk case studies</strong>: airline fare hallucinations, fabricated legal arguments, deleted databases, hijacked agents</p></li><li><p><strong>Governance</strong>: 6 emerging frameworks for managing AI risk that already exist</p></li><li><p><strong>Insurance</strong>: Who is innovating around pricing and covering AI risk and how are they doing it?</p></li></ul><p>The best way to move faster is to understand and manage the underlying risks ahead of time, so that AI speed doesn&#8217;t hit a wall halfway through the process. Read it for yourself, there&#8217;s no login and it&#8217;s free.</p><p><a href="https://agentrisk.phasetransitions.ai/">&#8594;Visit the Phase Transitions Agent Risk page</a></p><p>P.S. Even better: don&#8217;t read it yourself, just point your AI at the Agent Risk webpage and ask it:</p><p> &#8220;Is there anything here that&#8217;s relevant to the project we are working on https://agentrisk.phasetransitions.ai/&#8221;</p>]]></content:encoded></item><item><title><![CDATA[Fable 5: Agent CEO]]></title><description><![CDATA[Skills, Goals and Teams for Agent Management]]></description><link>https://phasetransitionsai.substack.com/p/fable-5-agent-ceo</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/fable-5-agent-ceo</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Fri, 03 Jul 2026 06:42:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Today&#8217;s piece is about building your own AI organisation that can run without you for hours at a time. This was already quite close with Opus 4.8, but with Fable 5 we are now fully equipped for the much anticipated &#8220;autonomous AI CEO&#8221;. Read on for my practical experience on this.</p><p>This piece is written from the pov of using Claude Code in the terminal. If you aren&#8217;t using AI that way (yet!) then follow along anyway to see what&#8217;s possible in managing teams of AI agents with Fable 5. </p><p>I am talking from my personal day-to-day experience here, hope it&#8217;s useful.</p><p>Regards</p><p>Graham</p><div><hr></div><p>Fable 5 is the first time we have an agent capable of operating as the CEO of an organisation made of AI agents. Opus 4.8 was pretty close but lacked common sense. Now that we have a CEO who can run an org of AIs what do we do with it?</p><p>The key thing is shift from prompting to managing. Luckily the people at Anthropic have already created the toolset for this: skills, goals and agent teams. You can build way of working for your agent organization, set them goals and coordinate teams of agents to achieve objectives. </p><p>The end result is that your AI org can work without you for hours at a time getting useful work done, supervised by your new agent CEO, Fable 5. This is not something that might happen in the future, it is already a daily practical reality.</p><p><strong>Skills: Agent Powerups</strong></p><p>A skill is a way of working for agents. For example, I like to get AI to make voiceover screencast videos of software I am building since I find this tedious and time consuming when I do it myself. It&#8217;s also nice to make the AI experience the software as a user and narrator because it helps spot user friction you can&#8217;t see by reading the code. </p><p>I give a single prompt and the agent opens the software, walks through a journey (clicking buttons, entering data, navigating around) and narrates out loud what it&#8217;s doing and seeing (it even has a nice pointer to show where it&#8217;s looking while it talks). I get back a video recording. Most of the time I don&#8217;t even watch the video, I simply pass it to the next agent team and ask them to fix the problems evident in the walkthrough. If you try and prompt your way to this it will take you a very long time with current state of the art models. You need to discover the various software packages and APIs required and correct all the mistakes the AI makes, which is fun the first time but boring the second time. </p><p>The solution is to take all of these lessons and bake them into a &#8220;skill&#8221; which is basically a folder of detailed instructions and software that AI can access and use. Then next time you want to do that thing you simply summon that specific skill instead of reinventing it or prompting your way there. The more skills you create the more powerups and shortcuts you have in your agentic AI work.</p><p><strong>Goals: Keep Going</strong></p><p>A frustrating thing about AIs is that they give up too easily or drift from the original intention. The solution is to set and lock a goal. The agent then gets working. Once it thinks it is done it looks up from its desk and checks &#8220;have I achieved the goal yet?&#8221;. If not, it keeps going. If there is no clear goal then you will get back half-done, drifting and rubbish work and now you become the manual goal enforcer, which is a boring job.</p><p>Just like with people, the goals need to be clearly defined. Ideally as a set of clear pass fail criteria that can&#8217;t be faked. A deterministic test suite is best, failing that independent review by a panel of AIs. Don&#8217;t just ask nicely, enforce your intent with unfakeable results (works with people too!).</p><p><strong>Workflows: AI Teams</strong></p><p>You can ask for workflows in Claude Code and start designing AI teams that get stuff done. For example you might want a detailed, fact-checked report on a complicated new topic that you are learning about. </p><p>An agent team might consist of the following. First a smart agent to define the job to be done and set the context. Second, a team of dumber, cheaper agents to read through and summarise thousands of original sources. Third, a team of smartish agents to fact-check the work and raise any inconsistencies or objections. Finally, a very smart, expensive agent to read through all of the summaries and objections and compile the final document. </p><p>A process like this can be used for any knowledge work task you can imagine (writing a report, making software etc.) and can run for several hours, deploy hundreds of agents and potentially burn tens of millions of tokens. The nice thing is that you don&#8217;t need to code it up yourself, just ask in plain English for a workflow, describe the kind of work you want done and Claude will help design and implement it (just make sure you use a top end model for the design and kickoff).</p><p><strong>Agent Orgs</strong></p><p>Now you have a set of specific skills, clear goals and a way to define and deploy teams. The only thing missing is a CEO to run the whole AI org. Before July 2026 this was a human job. This is where Fable 5 is special. In my experience this is the first model that has the common sense to correct and manage other agents and teams of agents. Fable 5 can autonomously and sensibly create new skills, set goals and spin-up and manage teams of agents on your behalf (using all of the tools and skills I described above). Now you can step back and operate more like an owner or board member than an executive, with your trusty robot CEO looking after the details.</p><p>We are still limited by the size of the context window, but it&#8217;s now possible to leave an agent alone with your computer for hours at a time pursuing a real world goal and deploying its own skills, goals and teams. Obviously don&#8217;t do this when dealing with sensitive data or in an environment that you haven&#8217;t thoroughly secured (get your agent team to look into ISO 27001&#8230;). But the point is that (given a safe and controlled environment) AI organizations that operate relatively autonomously are now an emerging daily practical reality accessible to anyone with a credit card and a laptop.</p><p>If you found this interesting the people over at Anthropic explain this much better than I ever could. Pop over to the <a href="https://www.anthropic.com/engineering">Claude Code engineering blog</a> to learn more.</p><p></p><p></p><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Writing Like A Robot]]></title><description><![CDATA[Practical tools for better communication with AI]]></description><link>https://phasetransitionsai.substack.com/p/writing-like-a-robot</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/writing-like-a-robot</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Mon, 29 Jun 2026 10:17:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>&#8220;I am drowning in long, meaningless AI generated emails and decks that people keep sending me. Please make it stop&#8221;</em>. This is a complaint I hear a lot nowadays from clients and friends.</p><p>I love AI and use it constantly. But used badly it can generate huge amounts of empty, weak writing that is hard to wade through. Over time I&#8217;ve developed my own toolset for preventing AI from default communication mistakes it continues to make. </p><p>Even the state of the art models make terrible slop by default. I now have a decent set of rules and instructions that I apply to writing and materials I generate with and in collaboration with AI.</p><p>Once you learn to spot the tells of AI slop you will see that it is everywhere. I&#8217;ve made lots of these mistakes myself in using AI to write and develop content here at Phase Transitions! So this is not an accusation at anyone but rather a step towards communicating in a better and more human-friendly way.</p><p>I call this tool (which is a combination of code and human language instructions) &#8220;antislop&#8221; in reference to the common term for bad AI writing (&#8220;AI slop&#8221;). </p><p>I&#8217;ve made a free slop-fighting website that anyone can access: paste in some AI writing and let the tool instantly highlight the sloppy parts for you. There&#8217;s a systematic ruleset of 35 antislop patterns and an entire codebase of tools and AI skills for you to use that are freely available via the site.</p><p>You can visit the <a href="https://antislop.phasetransitions.ai/">antislop</a> site yourself and get started in your own campaign against the flood of AI nonsense that is part of everyday working life for all of us nowadays. It takes 10 seconds to paste in some AI slop text and see the tool highlight the problems. </p><p>Hopefully you find it amusing and a little useful.</p><p>Kind regards,</p><p>Graham</p><p>P.S. This post was entirely handmade, so any slop it contains came straight from my own brain.</p>]]></content:encoded></item><item><title><![CDATA[Laying Foundations With AI]]></title><description><![CDATA[Going Faster With Less Risk]]></description><link>https://phasetransitionsai.substack.com/p/laying-foundations-with-ai</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/laying-foundations-with-ai</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Wed, 24 Jun 2026 11:23:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The easy thing to forget about AI is to put a little of that productivity windfall back into the foundations: the things that let you go far, not just fast.</p><p>One of my biggest lessons working on various AI projects over the last 6 months is that not investing enough in risk and compliance upfront is the best way to slow down whatever project you are working on. Investing more in safety upfront is the best way to actually go fast. That fact can get lost in the excitement of what is unfolding so quickly.</p><p>The good news about AI safety and compliance is that the same AI making you fast can help a lot with this part of the work. Staying safe is not a separate discipline you have to stop and go learn. It is mostly three ordinary things (documents, tests and reviews), and AI is good at helping you with all three if you ask.</p><p>Start with what does not change. You already have obligations, today, with or without AI. You have to protect personal data. You cannot let one client&#8217;s information leak into another&#8217;s. Someone has to be accountable when a decision goes wrong. </p><p>AI doesn&#8217;t create these duties but it does raise the stakes on them, because it touches more data, faster, with less human in the loop. Staying on top of them is what separates something real from a demo, and it comes down to documents, tests, and reviews.</p><p><strong>Building Foundations: Three Things</strong></p><p><strong>Documents.</strong> You write down what you do and why: what data you hold, how it moves, who is responsible when something breaks. AI is genuinely good at this. Point it at your own system and it will draft the policy, the data map, the register, and help keep them current as things change.</p><p>&#8220;Help me scaffold our policy on feeding personal information into AI tools&#8221;</p><p><strong>Tests.</strong> Some of what you promise can be proven mechanically, over and over. That your AI can only read and never write. That one customer&#8217;s records can never reach another&#8217;s. These are exactly the checks AI can help you write, and once written they run before every change, worth far more than a page that says you were careful.</p><p>&#8220;Write me a Python scripts to check that the AI can&#8217;t edit this database&#8221;</p><p><strong>Reviews.</strong> Some things no test can settle. Whether an output is fair, whether a judgment was sound. Here a person still looks and signs, but AI can tee the work up: gather the evidence, flag the edge cases, draft the assessment for a human to approve. Judgment does not automate, but the preparation for it does.</p><p>&#8220;Help me design a process to check for bias in our AI-enabled hiring workflow&#8221;</p><p>None of this is about slowing down. The tool making you fast is the same tool that makes building strong foundations cheaper and easier. You just have to remember to use it for both.</p><p>To win with AI you need to stay in the game long enough to succeed. So spend some of your tokens on making sure you stay safe while moving fast.</p><p>P.S. AI gets you moving and sharpens your thinking. It does not replace the human professional who signs off on the hard calls. It just means you reach them better prepared.</p>]]></content:encoded></item><item><title><![CDATA[Making More Progress With AI]]></title><description><![CDATA[How to build up your AI tools and skills, one step at a time]]></description><link>https://phasetransitionsai.substack.com/p/making-more-progress-with-ai</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/making-more-progress-with-ai</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Thu, 18 Jun 2026 07:28:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You don&#8217;t have to sprint to keep up with AI. </p><p>Let the machines do the hard work. Your job is to stay pointed in the right direction and receive the capability that&#8217;s coming to you. To do that effectively you need to understand the incremental steps of how to work with AI and what to make with it. That is much more valuable than any specific tool or fad.</p><p>You can&#8217;t skip steps, and the good news is you don&#8217;t need to. The next step is always small, and it&#8217;s already sitting right in front of you. Below I lay out what the steps are you so you can see exactly what&#8217;s coming down the track.</p><p>You can use whatever tool you like here. Take Claude as an example, since it&#8217;s the one a lot of people are adopting. Look at the tabs across the top: Chat, Cowork, Code. </p><h4><strong>Ways to work with AI</strong></h4><ol><li><p><strong>Chat</strong></p></li></ol><p>You start in Chat. You paste things in. You ask questions against your own numbers. You find out, in your own work, what this stuff is actually good for. That sounds basic. It&#8217;s the most important rung, because it teaches you what you want and how to ask for it, and you can&#8217;t build toward something you haven&#8217;t yet felt the need for.</p><ol start="2"><li><p><strong>Folders / Cowork</strong></p></li></ol><p>Then Cowork. Now the AI works alongside you on real files. It helps you organize, draft, structure, and tidy the pile. Cowork is where you build the thing you need before you can code: a set of organized files, and an understanding of how they fit together. Code assumes this structure already exists. Cowork is what builds this skillset in you. (If you already write software, you&#8217;ll skip ahead. This is for everyone else.)</p><ol start="3"><li><p><strong>Code</strong></p></li></ol><p>Then Code. Now you can run real tools. Set up a repository. Wire a connector. Stand up a database. You only get here because Chat showed you what you wanted and Cowork gave you the files and the literacy to do it.</p><h4><strong>Stuff you make with AI</strong></h4><p>Watch what you actually make as you climb the ladder. It&#8217;s four things, in order.</p><ol><li><p><strong>Files</strong></p></li></ol><p>AI makes a spreadsheet or a PDF in chat and you download it, email it whatever. Good start, linear, gets stale pretty quickly.</p><ol start="2"><li><p><strong>Folders</strong></p></li></ol><p>A folder on your machine. Inside are your notes, your corrections, the context you&#8217;re tired of re-typing, the way your business does things. A living and evolving collection of files. Now the AI is editing the files and organising them in response to what you need, not making one thing at a time.</p><ol start="3"><li><p><strong>Repos</strong></p></li></ol><p>Then, a repository on GitHub. The folder is straining: you don&#8217;t want to lose it, a colleague needs it, you want to see what changed, when and how. So you put it under version control and push it. Now it&#8217;s backed up, shareable, and every correction you make is a saved change you can see. You can run tests against it and turn it into running software. What used to be just documents are now living things that do work without you babysitting them. Your history becomes a record of what you&#8217;ve taught it. And it&#8217;s yours, on your account, not trapped inside someone&#8217;s chat.</p><ol start="4"><li><p><strong>Databases</strong></p></li></ol><p>Then, your own database. The repository strains when there&#8217;s too much to read top to bottom and sa machine (not a person) needs to ask it questions: an app, a dashboard, an agent. So the knowledge moves into a database, where it can be queried, live and structured. This is where you stop keeping notes and start owning infrastructure. Most people never need to come this far. A small neighbourhood business might be fine living in a folder forever. A bigger business ends up with the database.</p><p>The same one-step rule runs through your data. Say you want AI working with your accounting. You don&#8217;t start by building a connector. You start by copying numbers out of Xero and pasting them into Chat. You download the spreadsheet it makes. Then you turn on the native Xero connector that&#8217;s already there, one click. You use it until you feel its limits. Then you add a custom connector that does the specific thing you need and connect it to your own private database. Then you add your time tracking tool. Then the next tool. Until you&#8217;ve used the simple version you don&#8217;t yet know what to build. Each step shows you the next.</p><h4><strong>Learning step by step</strong></h4><p>This is what people mean when they talk about &#8220;owning your AI learning loop&#8221;. It isn&#8217;t a product you buy. It&#8217;s a folder that became a repository that became a database, built one step at a time.</p><p>Because each step hands you exactly what the next one needs, you can never be lost. There is always a defined next step, and if it feels too big just ask AI to break it down into smaller steps. The only ways to fail are to stop or to try and skip ahead.</p><p>The question is never &#8220;how do I catch up&#8221; or &#8220;how do I leap.&#8221; It is only: what is the one step in front of me today. Take it. Then it gives you the next.</p><p>Stay pointed in the right direction. Take one step forward every day. Let the machines do the work.</p>]]></content:encoded></item><item><title><![CDATA[It Was Never About the Machines]]></title><description><![CDATA[A letter to myself about AI]]></description><link>https://phasetransitionsai.substack.com/p/it-was-never-about-the-machines</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/it-was-never-about-the-machines</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Tue, 16 Jun 2026 08:51:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You make a thing. It's good. You're proud of it. Then everyone can make it. Fast. For free. And the thing you were proud of isn't worth much anymore.<br><br>Things are going better and better but something is bothering you. Your ears are ringing. Why? You're good at this, you keep telling yourself. You were. The thing just went cheap. Not your fault. That's how it works now. Every skill you pick up goes cheap, and it goes cheap fast, sometimes in months<br><br>For four years you watched the machines get good. Then very good. Then a little absurd. Each time you found a new thing you could suddenly do, you thought: this is it, this is the edge. Each time, within a year, everyone could do it too.<br><br>So you kept running. New tool, new model, new trick. You ran well. Stop running for a second. Here's the part for your gut. The thing you build is not the gift. You are.<br><br>You know what good looks like. You know what will make a client nod and what will make him cringe. You know who to trust and who to walk away from. You put your name on the work. When things breaks and people are afraid, you're the one who can be relied on to fix it quickly and calmly.<br><br>So stop being proud that you can build fast. Soon a child will build fast. Don't sell "I build fast." It's already dying. Sell the other thing. I know. I choose. I stay. Sell your eye. Sell your word. Sell that you carry the weight when the weight gets heavy. And the long private store of true things you've taught your machine, all the thousand times you said no, not like that, like this. That store is yours too.<br><br>Stop loving the thing you make. Love what you wrap around it.<br><br>Now the real reason. The one underneath all the rest. You do this for her. She doesn't need a dad who builds fast. The machine builds fast. She's going to grow up in a world that is drowning in fast. She needs a dad who knows what's good. A dad whose word is gold. A dad who stays when it's hard.<br><br>You spent three years staring at the machines, certain the edge was hidden somewhere inside them. It was never inside the machines. It was never about the machines.<br><br>It was always the eye, the word, the staying. Look away from the screen and look at your hands. Stop being proud of what you made and be proud of what you are.</p>]]></content:encoded></item><item><title><![CDATA[Real Work and AI: Up Against the Limit]]></title><description><![CDATA[If the models are already almost too good, what's the bottleneck?]]></description><link>https://phasetransitionsai.substack.com/p/real-work-and-ai-up-against-the-limit</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/real-work-and-ai-up-against-the-limit</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Mon, 01 Jun 2026 09:12:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>How do we use the miraculous new superpowers of AI for real work without accumulating risk faster than we can manage it?</p><p>In 2025 the story was prompt engineering. The tool was the chat window. The frontier was getting the AI to give you a better answer to your question. Fun, exciting, all upside.</p><p>Then in the last 6 months we got incredible new models, the agentic miracle, MCP etc. AI that can do real work, incredibly fast and very well (with the right scaffolding). Claude started doing totally unrealistic sci-fi stuff on your laptop for $100 per month. </p><p>The world changed very fast. But what got us here (miraculous individual tooling) is not going to get us there (making money at scale in real organizations without causing harm).</p><h2>Missing pieces</h2><p><strong>Trust architecture.</strong> </p><p>If you live in one tool and that tool can act on your repositories, your files, your communications, your infrastructure, the trust model has to be new. The question stops being &#8220;is the AI capable&#8221; and starts being &#8220;what is this agent allowed to do, on what evidence.&#8221;</p><p><strong>The liability gap.</strong> </p><p>When AI does the work, who bears the consequences? AI can&#8217;t be liable, so risk falls on the human who deployed it. Why be a 10X engineer/marketer/leader if that just makes you 10X as liable and the AI that &#8220;made the mistake&#8221; disappeared the moment you used it?</p><p><strong>AI agent identity.</strong> </p><p>When agents act on systems, who the agent is becomes operationally &#8220;load-bearing&#8221; (Claude&#8217;s favourite new word). Their identity needs to be something you can see and control. Ephemeral, identity-free agents can&#8217;t do real work and they can&#8217;t take any responsibility in your business.</p><p><strong>The accountability stack.</strong> </p><p>The full chain of who knew what, who approved what, what evidence persists after the fact, what can be reconstructed under scrutiny from a regulator. As regulated industries adopt AI seriously, particularly health, financial services, and legal, this becomes the money question.</p><p><strong>The institutional response.</strong> </p><p>Insurance products priced for AI-driven work. Regulatory frameworks. New risk infrastructure to replace the old mechanisms that AI velocity has thrown on the trash heap. This will be invented live over the next year or two by regulators, insurers, professional bodies, and the operators who carry the concentrated risk on their own shoulders until this layer arrives.</p><p><strong>The human role.</strong> </p><p>When AI does the work, is audited, is insured, is identified, then what is the leftover human role inside the process? This is the labour-and-meaning piece that no one really understands yet. I am optimistic about this but struggle to explain why.</p><h2>Less magic, more grounding</h2><p>The models can already do 10X more than we need or can absorb. That&#8217;s a problem just as much as it&#8217;s a benefit and the missing ingredients are risk management and control. </p><p>The next phase of AI growth will be about AI auth, identity, liability, and accountability question getting figured out in real time with a lot of money and credibility on the line. </p><p>Practical AI risk management is where the game will be won or lost.</p><div><hr></div><p><strong>Invitation to readers: Connect with me on LinkedIn</strong></p><p>I&#8217;m writing day-to-day observations about AI in business nowadays on LinkedIn.</p><p>Feel free to <a href="https://www.linkedin.com/in/graham-rowe/">follow or connect with me on LinkedIn</a>. </p><p>This Substack will continue as usual with (roughly) weekly long-form content.</p>]]></content:encoded></item><item><title><![CDATA[Less Is More with AI]]></title><description><![CDATA[Get your attention back]]></description><link>https://phasetransitionsai.substack.com/p/less-is-more-with-ai</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/less-is-more-with-ai</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Tue, 26 May 2026 06:30:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AI is accelerating hard. The natural human response is freeze, flee, or fight. Flee: ignore it, hope it stabilises before you really have to engage. Freeze: scroll news about AI without actually using it. Fight: try every new tool, every new model, every new agent, sprinting to keep up. All three are stress responses and they don&#8217;t really help.</p><p>The winning move is the opposite of fight-or-flight. Go deep, calm, focused. Pick one good, flexible tool. Spend all your time actually using it to get things done. Learn by doing. AI capability is no longer the constraint. The constraint is your ability to direct calm and sustained attention to actually use AI and learn how it practically works.</p><p>My own &#8220;less is more&#8221; move was to consolidate everything into Claude Code. Coding. Writing. Reading and summarising documents. Orchestrating agents. Managing projects, files and folders. Setting up infrastructure. Everything. I&#8217;m getting much more done than I did with five tools. My attention isn&#8217;t fragmented and I am starting to know one excellent, flexible tool very well.</p><p>But don&#8217;t I need to keep up with the latest? My view has changed: the best way to &#8220;keep up&#8221; is to just commit to a single tool and use it constantly. That way the latest trends come to you inside the flow of work getting done, automatically. The people building these things are smart and competing hard. If a capability matters, it shows up inside the tool soon enough. I have spent months spending a lot of my working day in Claude Code and every day I learn new things about AI capabilities by using it to get work done. The product itself keeps changing every week&#8230; it&#8217;s enough work to keep up with your existing tool nevermind try new ones.</p><p>So try this. Pick the one tool you&#8217;d genuinely miss if it went away. It doesn&#8217;t have to be Claude Code, that&#8217;s just me. Pick what vibes for you. Run everything through it. Resist the new releases and the scroll. Notice how much attention comes back. Notice the work itself getting better when you&#8217;re not switching every twenty minutes between different AI surfaces.</p><p>This is a wonderful time to be using these tools. The possibilities are really incredible for all of us. The limiting factor isn&#8217;t the AI. It&#8217;s how you use it. Sustained attention, familiarity and maybe eventually mastery. Less is more.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Why AI Is Tiring You Out]]></title><description><![CDATA[Stop writing better prompts and just say what you think]]></description><link>https://phasetransitionsai.substack.com/p/why-ai-is-tiring-you-out</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/why-ai-is-tiring-you-out</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Tue, 05 May 2026 06:01:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The latest models are good enough that you can hardly tell AI wrote what you&#8217;re reading. Everything looks great. If AI is doing the heavy lifting, then why do some days of AI-intensive work leave you exhausted instead of energised?</p><p>If a human gave you complicated writing with no point of view, you could just stare them straight in the eye and wait. They&#8217;d squirm. Eventually they&#8217;d confess: <em>what I really wanted to say was..</em>. The volume and elaboration existed because they were avoiding a point. Pressing harder would have surfaced the real point.</p><p>Your AI has no buried point. The model is incapable of having a point of view. It&#8217;s not someone.</p><p>That&#8217;s why you&#8217;re tired. You&#8217;re reading text that imitates having a position, scanning for the point you assume must be there, and finding that the search itself is what&#8217;s exhausting. Every sentence reads. Every paragraph holds. It looks great on the surface. But it&#8217;s empty in a confusing way. That&#8217;s a long day.</p><p>The fix isn&#8217;t better prompting. Better prompts produce more elaborate emptiness. The fix is to stop prompting and start talking. Take your hands off the keyboard. Click Superwhisper (my favourite, no affiliation) or whatever voice tool you use. Say what you think out loud, in your actual voice. Don&#8217;t structure it. Don&#8217;t summarise it. Just let your mouth say it. No editing. Give the model the recording (if it&#8217;s long) or the transcript and let it arrange the details around what you said.</p><p>AI is good at building from a position and good at testing it. But if you try to let it have a point of view you will get beautiful, exhausting emptiness (and lots of it). Use it for what it&#8217;s good at. Stop trying to prompt-engineer a point that has to come from who you are.</p><p>The point is yours. It can be messy, human and real. Let the AI do some of the logic, the testing, the elaboration, the perfecting. It might even help contradict and disprove your ideas. Then you actually learnt something! Make your point honestly, early and often and see how your energy lifts.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Learning Faster With AI]]></title><description><![CDATA[What winning in business with AI actually looks like]]></description><link>https://phasetransitionsai.substack.com/p/learning-faster-with-ai</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/learning-faster-with-ai</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Fri, 24 Apr 2026 06:39:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every business is a science experiment it doesn&#8217;t know it&#8217;s running.</em></p><p>Should we raise prices? Fire the bottom 10% of our clients? Hire another salesperson? Move into the US? Every one of those is a hypothesis. The business will eventually find out whether it was right. But in eighteen months, by which point the conditions have shifted, the market has moved, and the lesson (if anyone bothers to extract it) is stale.</p><p>This is how most organisations run. Not because their leaders are stupid, but because the feedback loop between a decision and its verdict is long, noisy, and mostly invisible. By the time the data speaks, the question has moved on.</p><p>That&#8217;s where the interesting part of AI lives. Not in the places people are currently looking.</p><h2><strong>What happens when thinking is 100X easier?</strong></h2><p>Read the discourse and it sounds like we&#8217;re arguing about whether AI can replace this job or that. Whether the models are approaching AGI. Whether agents can book a flight without supervision. Fine questions. But they don&#8217;t describe what&#8217;s happening.</p><p>What&#8217;s actually happening is much more interesting. The cost of producing a useful observation about a business (a well-framed question, a clean summary, a &#8220;wait, that&#8217;s weird&#8221; instinct) has totally collapsed. Ten years ago that observation required an analyst, a spreadsheet, and half a week. Five years ago it required an analyst and a dashboard and a few hours. Today it requires a question and thirty seconds.</p><p>That&#8217;s not a productivity upgrade. That&#8217;s a change in the economic cost of thinking.</p><p>And when the cost of an activity drops by 100X, what changes isn&#8217;t the activity. It&#8217;s what the activity is for.</p><h2><strong>Growth = rate of learning</strong></h2><p>A business doesn&#8217;t grow because it knows more than its competitors. It grows because it <em>learns faster</em>. Two firms with identical revenue, identical teams, identical markets will diverge wildly over five years based on the rate at which they convert observation into action. Capital matters, reach matters, people matter. But the rate of learning is what determines which of two otherwise-equivalent businesses wins.</p><p>Most of the business world has no instruments for measuring its own rate of learning. It has budgets, org charts, OKRs, board packs, KPIs. None of those measure the thing that actually compounds.</p><p>Every question asked and answered in real time is a small experiment closed. Every daydream checked against data is an item crossed off a list that didn&#8217;t exist before. That&#8217;s what&#8217;s new. Not the answers. The rate at which the loop closes. How fast we learn and adapt.</p><h2><strong>The loop, not the answer</strong></h2><p>The mistake is treating each AI interaction as an answer you either keep or throw away. That&#8217;s what most tools still look like: input, output, repeat. No memory of what was asked. No continuity between conversations. No accumulation.</p><p>Useful learning doesn&#8217;t work that way. It&#8217;s a loop: observe, hypothesise, act, observe again. The power isn&#8217;t in any single step. It&#8217;s in the speed and clarity with which the loop closes.</p><p>An AI that compresses a single step but doesn&#8217;t close the loop is empty. An AI embedded in the sequence of helping formulate the hypothesis, surface the evidence, name the decision, remember what was decided and why, notice when the outcome confirms or falsifies it&#8230; is doing the thing that matters.</p><p>We&#8217;re early in that second story. Most AI products are still stuck somewhere in the first. But the direction is visible.</p><h2><strong>Loops, bottlenecks and orgs</strong></h2><p><strong>First, the value shifts from answers to loops.</strong> Single-shot AI (ask a question, get a good answer) is a commodity. The durable product is AI that maintains continuity across a sequence of questions, against a specific context, across an organisation, over a long period of time. That&#8217;s harder to build and much less fungible to replace.</p><p><strong>Second, the bottleneck becomes the human.</strong> Once the AI can surface observations and hypotheses fast, the constraint on learning rate shifts to whoever reviews, decides, and acts. The old bottleneck of &#8220;can we get the data?&#8221; is already gone. The new one is &#8220;can we look at what the data is telling us, and respond?&#8221; Most organisations are terrible at that. They&#8217;ll need to become less bad as fast as they can or lose to those who evolve.</p><p><strong>Third, the unit of compression shifts from the individual to the organisation.</strong> Almost all the AI adoption you&#8217;ve seen so far has been individual. A lawyer drafting faster, a developer coding faster, a marketer ideating faster. That&#8217;s real. But it&#8217;s limited. The thing that actually compounds is <em>organisational</em> learning: the rate at which a firm as a whole turns observation into coordinated action. That requires AI that lives in the shared context of the firm, not in the private chats of its employees. We&#8217;re much earlier there than the hype suggests.</p><h2><strong>AI doesn&#8217;t replace people, it catalyses them</strong></h2><p>The argument for AI is not that it will replace humans, or do more of what humans already do. It&#8217;s that it shortens the distance between observation and action. Between a question worth asking and an answer you can use. Between a hypothesis and its verdict.</p><p>A business that learns faster than its market, wins. A business whose people can test and discard ideas faster than they could ten years ago about pricing, about segments, about tools, about themselves has a compounding advantage no amount of capital can buy.</p><p>That&#8217;s the thing worth watching. Not the benchmarks. Not the capability demos. Not the replacement anxieties.</p><p>The compression of the learning loop. Runaway learning, &#8220;unexplained&#8221; competitive advantage, new firms exploding out of nowhere and taking their market. That&#8217;s the story.</p><p>If you&#8217;re building with AI in your business right now, this is the thing worth building toward.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Holding AI Accountable]]></title><description><![CDATA[Why the social forces that make humans do good work don't transfer to AI and what actually works.]]></description><link>https://phasetransitionsai.substack.com/p/holding-ai-accountable</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/holding-ai-accountable</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Thu, 16 Apr 2026 06:43:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>&#8220;The problem with AI is there&#8217;s no neck to strangle.&#8221; - <em>a friend at lunch</em></p><p>AI fails in different ways to how people fail. When it goes horribly wrong you can&#8217;t pull it into a room. You can&#8217;t threaten its bonus. It doesn&#8217;t care if you&#8217;re disappointed or angry.</p><p>My friend was joking, but he&#8217;d named something important. When you hire a junior analyst, most of what keeps them honest isn&#8217;t the formal review. It&#8217;s the thousand invisible social forces that make being wrong embarrassing and being right rewarding. AI has none of them. And most people trying to use AI for serious work are struggling to adapt to this.</p><h2>The forces that keep humans honest</h2><p>Think about what actually happens when a junior analyst at a decent firm produces a piece of work. They know their name is on it. They know their manager will read it. They know if the numbers are wrong the client will spot it in a meeting and they&#8217;ll have to stand there while someone more senior explains it away. They know the partner who hired them is watching. They know the colleague in the next chair will glance over and say &#8220;that looks off.&#8221; They know if this keeps happening they won&#8217;t make the next promotion. They know if it gets bad enough they&#8217;ll be asked to leave.</p><p>None of this is in their job description. None of it is in the quality control process. It&#8217;s all just there, in the air, pressing on them from every direction. And it does most of the work. The formal review process catches the mistakes that slip through. The social machinery catches the mistakes that would otherwise be made in the first place.</p><p>We don&#8217;t usually notice any of this because we&#8217;ve never worked without it. Every office we&#8217;ve ever been in had a version of it. Every professional we&#8217;ve ever hired came pre-loaded with a reputation to protect and a career to advance. You pay someone to do good work, but what you&#8217;re actually buying is their incentive to not do bad work.</p><p>AI is the first worker most of us have ever hired where that incentive doesn&#8217;t exist. Not because AI is uniquely untrustworthy, but because there&#8217;s no person there to be trustworthy or otherwise. The forces that keep humans honest are forces that act on humans. For AI there is literally nothing there, we are imagining it all while we are busy shouting at an AI at a mistake that it made. Talking to ourselves.</p><h2>What you replace them with</h2><p>People figure this out in a predictable order, usually the hard way.</p><p>It starts with hope. You ask AI to do something, you read the output, and if it looks right you use it. Most people live here longer than they&#8217;d like to admit. The accountability mechanism is vibes.</p><p>Then something goes wrong and you try a cruder trick: you ask again. Different words, different angle, see if the answer holds up. This works surprisingly well for surprisingly long, because AI&#8217;s variance becomes a useful signal &#8212; if you get three different answers, at least one of them is probably wrong.</p><p>When that stops being enough, you try getting AI to critique its own work. Hand it the output and say &#8220;what&#8217;s wrong with this?&#8221; It turns out AI is often better at criticism than creation, and this catches a lot of what the first pass missed. For most people&#8217;s purposes, this is as far as they ever need to go.</p><p>But when the stakes get higher &#8212; when the output is going to a client, or shaping a real money decision, or feeding into something else that depends on it &#8212; self-critique isn&#8217;t enough either. So you build a checklist. The checklist is the crystallised memory of every mistake you&#8217;ve seen AI make before. Did you check this? Did you verify that? Are these numbers consistent with those numbers? You make the AI walk through it. Now the accountability is explicit rather than hoped-for.</p><p>Past a certain complexity, checklists aren&#8217;t enough either, because a single AI doing its own checking has blind spots it can&#8217;t see past. So you set up a second AI whose only job is to disagree with the first. Separate context, separate instance, paid to find problems. The disagreements are where the real errors hide.</p><p>And eventually, for the things that really matter, you stop trusting any AI to check anything. You write mechanical code that runs on every output and blocks it if a specific rule is violated. The partial month gets excluded. The variance exceeds tolerance. The banned phrase appears. The check fails, the output doesn&#8217;t ship. You&#8217;ve left AI behind entirely for this layer, because the only thing you can fully trust is a mechanical rule that executes the same way every time. Eventually you have dozens of these for anything that matters.</p><p>Each step costs more than the last. Each catches a class of error the previous step couldn&#8217;t. Asking AI nicely and hoping is fine for a first draft. It&#8217;s a disaster for anything with numbers in it that someone&#8217;s going to act on.</p><h2>The point</h2><p>We&#8217;ve been trained to think about AI quality as a property of the model. Smarter model, better output. But for real work, quality is mostly a property of the scaffolding around the model. A junior human in a bad firm will do bad work. A junior human in a good firm will do good work. Same person, different accountability infrastructure.</p><p>AI is the same. If you want output you can trust, don&#8217;t just pick the smartest model. Build the scaffolding. And assume you&#8217;ll need much more of it than you think, because the human forces you&#8217;re replacing were doing a lot more work than you realised.</p><p>&#8212; Graham</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Your AI Has a Notebook. It Needs a Brain.]]></title><description><![CDATA[You already know that AI remembers things about you.]]></description><link>https://phasetransitionsai.substack.com/p/your-ai-has-a-notebook-it-needs-a</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/your-ai-has-a-notebook-it-needs-a</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Wed, 08 Apr 2026 14:50:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You already know that AI remembers things about you. It knows your name. It knows you prefer bullet points. If you told it about your business last week, there&#8217;s a decent chance it&#8217;ll recall the broad strokes next time you open a chat.</p><p>This is a genuine step forward from a year or two ago, when every conversation started cold. But if you&#8217;ve been using these tools seriously &#8212; for real work, not just quick questions &#8212; you&#8217;ve probably started to notice where the memory breaks down.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Phase Transitions! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>It contradicts itself. You told it you moved to Melbourne in January. It still thinks you&#8217;re in Sydney, because it never deleted the old fact &#8212; it just wrote the new one underneath. Both are in there. Which one it uses depends on its mood.</p><p>It can&#8217;t really connect memory dots. It knows Alice works at Google. It knows David manages Alice. It knows David reports to Rachel. But it doesn&#8217;t implicitly know &#8220;who does Alice&#8217;s manager report to?&#8221;, because it stored three separate facts, not a relationship.</p><p>And it doesn&#8217;t scale. The memory is basically a long text file. The more you use it, the longer the file gets, until it quietly starts cutting old memories to make room &#8212; or just slows down because it&#8217;s processing thousands of words of context before it even gets to your question.</p><h2>The difference between a notebook and a brain.</h2><p>A notebook is useful. You write things down, you flip back through it, sometimes you find what you need. But a notebook doesn&#8217;t cross things out when they&#8217;re wrong. It doesn&#8217;t open to the right page based on what you&#8217;re working on. It can&#8217;t tell you how two things you wrote six months apart are connected. And eventually, it&#8217;s just too full to be useful.</p><p>A brain does all of those things automatically. It updates when facts change. It surfaces what&#8217;s relevant and lets the rest recede. It connects information across domains. It knows that something you learned on Tuesday is related to a problem you&#8217;re facing on Friday, even if you didn&#8217;t make that link yourself.</p><p>Right now, AI has a notebook. Nearly a thousand open-source projects are working on giving it a brain.</p><h2>What does the future of AI memory actually look like?</h2><p>Engineers are hard at work innovating in this space and the shape of the future is already getting pretty clear. Three interesting patterns I&#8217;ve noticed&#8230;</p><p>One system &#8212; already in use, with serious engineering behind it &#8212; runs every new piece of information through a comparison against everything it already knows. If you say &#8220;I moved to Melbourne,&#8221; it doesn&#8217;t just write that down. It finds &#8220;Lives in Sydney,&#8221; recognises the contradiction, and updates the record. One fact, current, clean. It does this automatically, every time.</p><p>Another builds a map of relationships, not just facts. It doesn&#8217;t store &#8220;Alice works at Google&#8221; as a string of text &#8212; it stores Alice as a person, Google as a company, and &#8220;works at&#8221; as the connection between them. This means it can follow chains: Alice&#8217;s manager&#8217;s colleague&#8217;s project. That&#8217;s not a party trick. For a business where the AI needs to understand your client relationships, your team structure, or your supply chain, it&#8217;s the difference between a search engine and an analyst.</p><p>A third project &#8212; written by a single developer, now one of the fastest-growing AI tools on the internet &#8212; packages the entire memory system into a single file. No database server, no infrastructure. You can copy it to a USB stick. It has timestamps on every memory, so you can ask &#8220;what did I tell you last month?&#8221; and get an answer. It can even replay how its understanding of you evolved over time.</p><p>More detail on these projects and more <a href="/__u/phasetransitionsai.substack.com/p/your-ai-agent-has-amnesia-and-977">here</a>.</p><h2>Why this matters for your business</h2><p>In the last piece I wrote about <a href="/__u/phasetransitionsai.substack.com/p/what-happens-when-you-give-ai-a-desk">giving AI a desk</a> &#8212; connecting it to your real data, not just the internet&#8217;s data. Memory is the next piece. It&#8217;s what turns a powerful tool you interact with into something that accumulates knowledge about your business over time.</p><p>Think about the most valuable people in your organisation. It&#8217;s rarely the ones with the best raw ability. It&#8217;s the ones who&#8217;ve been around long enough to know where the bodies are buried &#8212; who remember what was tried in 2019, why the pricing model changed, which clients need careful handling and why. Institutional memory is one of the most valuable and least visible assets a company has.</p><p>AI is about to develop the same capability. Not as a flat file of preferences, but as a structured, self-maintaining, queryable understanding that grows every time it&#8217;s used and retrieves exactly what&#8217;s relevant when it&#8217;s needed.</p><p>If you&#8217;ve already given AI a desk, the brain is coming. And if you want the technical detail on exactly which projects are building it and how, I wrote a deeper piece on that <a href="/__u/phasetransitionsai.substack.com/p/your-ai-agent-has-amnesia-and-977">here</a>.</p><p>&#8212; Graham</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Your Agent is Hitting its Ceiling — Who's Actually Fixing It]]></title><description><![CDATA[Claude Code's leaked source reveals the architectural ceiling. The frustrations you already feel aren't bugs &#8212; they're the edges of a pattern that was designed for a different world.]]></description><link>https://phasetransitionsai.substack.com/p/your-agent-is-hitting-its-ceiling</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/your-agent-is-hitting-its-ceiling</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Sun, 05 Apr 2026 13:33:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You already know something is off.</p><p>You&#8217;ve lost a 45-minute Claude Code session to a context compaction that threw away the thing you needed. You&#8217;ve watched it redo work it already did because it can&#8217;t remember across sessions. You&#8217;ve had a multi-agent run silently go sideways and only noticed when the output was wrong. You&#8217;ve tried to resume after a crash and realised there&#8217;s nothing to resume from.</p><p>These aren&#8217;t bugs. Claude Code is genuinely excellent &#8212; the most effective agentic tool anyone has shipped. When the source code leaked in March 2026 (512,000 lines of TypeScript via an npm source map), it revealed an architecture of striking simplicity: plain markdown files for memory, JSONL transcripts as the source of truth, a 914-line system prompt as the orchestrator. The simplicity <em>is</em> the architecture.</p><p>But the frustrations you&#8217;re feeling aren&#8217;t about Claude Code being bad. They&#8217;re about running into the edges of the <strong>smart process, ephemeral state</strong> pattern &#8212; designed for a human-in-the-loop, session-length world. And the instinct most people have (&#8221;I need better memory,&#8221; &#8220;I need a bigger context window&#8221;) is treating the symptom, not the cause.</p><h2>The four problems you can feel</h2><p>Backend engineering identified and solved every one of these problems decades ago. The agent ecosystem is rediscovering them &#8212; in the wrong order, at the wrong level of abstraction.</p><h2>1. Memory: the right instinct, wrong level</h2><p>Claude Code uses plain markdown files with a 25KB cap. No vector database. No embeddings. The most widely adopted agent in the world solves memory with text files &#8212; and it works, because the real constraint isn&#8217;t retrieval quality. It&#8217;s that the agent&#8217;s truth dies with its process.</p><p><a href="https://mcp.phasetransitions.ai/rag/servers/mem0ai/mem0/">mem0</a> (52,000+ stars, 2.8M downloads/month) is integrated into CrewAI, Agno, and AgentScope. <a href="https://mcp.phasetransitions.ai/vector-db/servers/topoteretes/cognee/">Cognee</a> (quality 80/100, 372 commits in 30 days) has the highest dev velocity in the category. These are genuinely good projects. But bolting a memory layer onto an ephemeral process doesn&#8217;t make the process durable. It gives it a longer scratchpad. <a href="https://mcp.phasetransitions.ai/agents/servers/ayushmi/agentstate/">AgentState</a> is one of the few projects that understands the difference &#8212; WAL+snapshots, CRDTs, database primitives rather than retrieval primitives.</p><h2>2. Orchestration: prompts work until they don&#8217;t</h2><p>Claude Code&#8217;s coordinator mode is a system prompt, not code. &#8220;Research &#8594; synthesis &#8594; implementation &#8594; verification&#8221; are directives, not edges in a dependency graph. This works because the model is good enough to self-sequence for interactive tasks, and a human is watching.</p><p>The best orchestration solutions come from outside the agent world. <a href="https://mcp.phasetransitions.ai/agents/servers/triggerdotdev/trigger.dev/">trigger.dev</a> (14,000 stars, 768K downloads/month) is background jobs infrastructure being adopted by agent developers. <a href="https://mcp.phasetransitions.ai/data-engineering/servers/dagu-org/dagu/">dagu</a> (quality 70/100) is a declarative workflow engine from data engineering. <a href="https://mcp.phasetransitions.ai/agents/servers/ComposioHQ/agent-orchestrator/">Composio&#8217;s agent-orchestrator</a> (4,300 stars) is the standout agent-native entry &#8212; DAG-based planning, parallel agents, git worktrees. It looks like a worker pulling tasks from a queue. That&#8217;s the shape of what comes next.</p><h2>3. Observability: you can&#8217;t debug what you can&#8217;t see</h2><p>Claude Code&#8217;s observability is regex-based frustration detection and JSONL transcripts you can grep. When agents run 6&#8211;8 hour tasks, that stops being a strategy. <a href="https://mcp.phasetransitions.ai/agents/servers/coze-dev/coze-loop/">Cozeloop</a> (5,400 stars) from ByteDance provides full-lifecycle management. <a href="https://mcp.phasetransitions.ai/agents/servers/Siddhant-K-code/agent-trace/">agent-trace</a> calls itself &#8220;strace for AI agents&#8221; &#8212; the right metaphor. February 2026 saw 14 new observability repos in a single month, up from 0&#8211;4 prior. Sentrial (YC W26) raised money on exactly this gap.</p><h2>4. Crash recovery: the void</h2><p>This is the finding that reframes everything else. Claude Code has no crash recovery. Sessions are stateless &#8212; if the process dies, you start over. And the agent ecosystem has almost completely ignored this too.</p><p>Backend engineering solved crash recovery decades ago &#8212; Temporal, Inngest, DBOS, Restate. These are proven, production-grade durable execution runtimes. Across the entire AI agent ecosystem: <strong>temporalio has 1 dependent. Inngest has 1. DBOS has 0. Restate has 0.</strong> Compare: LangChain has 273 dependents. The infrastructure that guarantees exactly-once execution has near-zero penetration.</p><p>This is the structural diagnosis behind every frustration you&#8217;ve had. It&#8217;s not that you need better memory. It&#8217;s that the &#8220;smart process, ephemeral state&#8221; pattern can&#8217;t do crash recovery, because there&#8217;s nothing to recover <em>to</em>.</p><h2>What got us here won&#8217;t get us there</h2><p>Anthropic knows this. The leaked source reveals what they&#8217;re building next: ULTRAPLAN (30-minute remote planning sessions), KAIROS (proactive background daemon), Bridge mode (cross-machine session handoff). Each pushes past the edges of the current pattern towards something that looks more like durable infrastructure.</p><p>The projects bridging the gap: trigger.dev (background jobs adopted by agent developers), Sayiir (&#8221;simplified Temporal&#8221; in Rust), Stabilize (queue-based state machine), and the first DBOS+LlamaIndex integration. The fix isn&#8217;t smarter orchestration within the process. It&#8217;s killing the process as the locus of truth and putting the truth somewhere that survives it.</p><h2>Explore the data</h2><p>Every project has a quality-scored page updated daily. Browse <a href="https://mcp.phasetransitions.ai/agents/categories/">agent categories</a>, <a href="https://mcp.phasetransitions.ai/agents/trending/">trending agents</a>, or explore <a href="https://mcp.phasetransitions.ai/agents/categories/agent-memory-systems/">agent memory</a>, <a href="https://mcp.phasetransitions.ai/agents/categories/agent-orchestration-platforms/">orchestration</a>, and <a href="https://mcp.phasetransitions.ai/agents/categories/agent-observability-debugging/">observability</a>.</p><p>The full deep dive with all repos and live quality scores is at <a href="https://mcp.phasetransitions.ai/insights/agentic-ball-of-mud/">mcp.phasetransitions.ai/insights/agentic-ball-of-mud</a>.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Your AI can't understand the docs you wrote for humans]]></title><description><![CDATA[MCP repos with detailed docs have 18x the stars. llms.txt is the new robots.txt. Here's what agent-readable documentation looks like.]]></description><link>https://phasetransitionsai.substack.com/p/your-ai-cant-understand-the-docs</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/your-ai-cant-understand-the-docs</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Fri, 03 Apr 2026 14:18:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>On a recent Y Combinator podcast, a founder described how Claude Code chose Supabase for his project&#8217;s database. Not because the founder asked for Supabase. Because Claude Code read the documentation, found it well-structured, and decided it was the best fit. The agent made the tool selection. The human just approved.</p><p>PT-Edge tracks 226,000+ AI repos. In the MCP ecosystem &#8212; where agents are literally the primary user &#8212; repos with detailed, structured documentation average 354 stars. Those without average 19. That&#8217;s an 18x difference.</p><h2>Documentation is the new distribution</h2><p>In the pre-agent era, developer tools got adopted through Stack Overflow, conference talks, and GitHub trending. In the agent era, the agent reads your docs, evaluates whether they contain enough to solve the current problem, and either uses your tool or moves on. The documentation <em>is</em> the product evaluation.</p><p>Better docs &#8594; agents choose your tool &#8594; more usage &#8594; more adoption. Worse docs &#8594; agents skip you &#8594; your tool doesn&#8217;t exist in the agent&#8217;s world.</p><h2>llms.txt: the new robots.txt</h2><p>A convention is emerging. Just as <code>robots.txt</code> tells search crawlers how to interact with a site, <code>llms.txt</code> tells AI agents what a project does and how to use it.</p><p><a href="https://mcp.phasetransitions.ai/rag/servers/mensfeld/llm-docs-builder/">llm-docs-builder</a> transforms markdown docs into LLM-optimised formats and generates llms.txt files. <a href="https://mcp.phasetransitions.ai/embeddings/servers/revokslab/codecrawl/">CodeCrawl</a> turns entire codebases into LLM-ready data with llms.txt from a single API call. <a href="https://mcp.phasetransitions.ai/servers/thedaviddias/mcp-llms-txt-explorer/">mcp-llms-txt-explorer</a> lets agents discover and browse llms.txt files across the web.</p><h2>Beyond llms.txt: agent-native protocols</h2><p><a href="https://mcp.phasetransitions.ai/agents/servers/osmandkitay/aura/">AURA</a> &#8212; Agent-Usable Resource Assertion &#8212; is an open protocol &#8220;designed to make the web machine-readable.&#8221; It replaces fragile scraping with structured, agent-optimised content. This isn&#8217;t just better docs. It&#8217;s a new protocol layer.</p><p><a href="https://mcp.phasetransitions.ai/servers/marckrenn/rtfmbro-mcp/">rtfmbro</a> provides always-up-to-date, version-specific package documentation as context for coding agents. When Claude Code needs to use a library, rtfmbro provides the exact docs for the exact version &#8212; preventing outdated code suggestions.</p><p><a href="https://mcp.phasetransitions.ai/servers/docfork/docfork/">Docfork</a> is explicitly &#8220;Up-to-date Docs for AI Agents.&#8221; The framing tells you where this is going: documentation as a service, maintained specifically for agent consumption.</p><h2>What to do right now</h2><p>If you maintain a developer tool: add an llms.txt. Structure your README for extraction &#8212; code-first, Q&amp;A sections, version-specific install commands. Document every error with a resolution. Agents will search for it.</p><p>Documentation is becoming the primary interface between your tool and the agents that decide whether to use it. The projects that optimise for agent readability now will have a structural advantage as agent-driven adoption becomes the default.</p><h2>Explore the data</h2><p>Every project mentioned here has a quality-scored page in our directories, updated daily. Browse the <a href="https://mcp.phasetransitions.ai/categories/">MCP categories</a>, <a href="https://mcp.phasetransitions.ai/agents/categories/">agent categories</a>, or check <a href="https://mcp.phasetransitions.ai/trending/">what&#8217;s trending</a> this week.</p><p>The full deep dive with the MCP correlation data and all tools is at <a href="https://mcp.phasetransitions.ai/insights/agent-native-docs/">mcp.phasetransitions.ai/insights/agent-native-docs</a>.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Your AI agent doesn't have an email address and 30 repos want to fix that]]></title><description><![CDATA[Identity, credentials, email, payments &#8212; scattered across 25 subcategories with no name. The infrastructure layer nobody's talking about.]]></description><link>https://phasetransitionsai.substack.com/p/your-ai-agent-doesnt-have-an-email</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/your-ai-agent-doesnt-have-an-email</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Fri, 03 Apr 2026 14:15:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Your AI agent needs to send an email. Gmail blocks automation. It needs API credentials that persist between sessions. OAuth assumes a human. It needs a phone number. Twilio wasn&#8217;t built for autonomous callers.</p><p>Every piece of infrastructure your agent touches was designed for humans and grudgingly tolerates machines. PT-Edge tracks 30+ repos building what we&#8217;re calling agent-native infrastructure &#8212; services designed from the ground up for AI agents as first-class entities. They&#8217;re scattered across 25 different subcategories because nobody has named this category yet.</p><h2>Five layers forming</h2><p><strong>Identity.</strong> <a href="https://mcp.phasetransitions.ai/agents/servers/Agent-Field/agentfield/">AgentField</a> (881 stars) treats agents like microservices &#8212; &#8220;scalable, observable, and identity-aware from day one.&#8221; <a href="https://mcp.phasetransitions.ai/agents/servers/open-gitagent/gitclaw/">GitClaw</a> puts agent identity in git &#8212; rules, memory, and skills as version-controlled files. <a href="https://mcp.phasetransitions.ai/agents/servers/AgentlyHQ/aixyz/">AiXYZ</a> goes furthest &#8212; on-chain identity via ERC-8004 with native payment wiring.</p><p><strong>Communication.</strong> <a href="https://mcp.phasetransitions.ai/agents/servers/agenticmail/agenticmail/">AgenticMail</a> gives agents their own email addresses and phone numbers &#8212; send and receive real messages without borrowing a human&#8217;s inbox. <a href="https://mcp.phasetransitions.ai/servers/Dicklesworthstone/mcp_agent_mail/">MCP Agent Mail</a> (1,800+ stars) provides async coordination with persistent inboxes, searchable threads, and file leases.</p><p><strong>Gateways.</strong> <a href="https://mcp.phasetransitions.ai/agents/servers/BlockRunAI/ClawRouter/">ClawRouter</a> (5,400+ stars) is explicitly &#8220;agent-native&#8221; &#8212; routing across 41+ models with sub-millisecond latency and USDC payments on Base and Solana. The payment layer is built in, not bolted on.</p><p><strong>Runtime.</strong> <a href="https://mcp.phasetransitions.ai/agents/servers/agentsystems/agentsystems/">AgentSystems</a> is a self-hosted app store for agents with container isolation and credential injection. <a href="https://mcp.phasetransitions.ai/agents/servers/jentic/jentic-mini/">Jentic Mini</a> sits between your agent and the outside world &#8212; &#8220;your agent says what it wants to do, Jentic handles the how.&#8221;</p><p><strong>Governance.</strong> <a href="https://mcp.phasetransitions.ai/agents/servers/microsoft/agent-governance-toolkit/">Microsoft&#8217;s Agent Governance Toolkit</a> brings zero-trust identity and policy enforcement specifically for agents.</p><h2>25 subcategories and no name</h2><p>The most telling signal: these repos are scattered across 25 different subcategories. Agent payment protocols, MCP gateway infrastructure, agent security hardening, email MCP servers, lightweight agent runtimes &#8212; the taxonomy doesn&#8217;t have a bucket for &#8220;infrastructure for agents as entities.&#8221; That&#8217;s because the category is forming right now.</p><p>On Hacker News, the pace is accelerating. AgentMail launched in January (169 points). &#8220;Agent Passport &#8212; OAuth-like identity for agents&#8221; appeared in February. &#8220;KeyID &#8212; email and phone infrastructure for AI agents&#8221; in March. Each month brings new infrastructure designed specifically for agents as first-class entities.</p><h2>Where this is heading</h2><p>As agents become more autonomous, they need more of their own infrastructure. Today it&#8217;s email and credentials. Tomorrow it&#8217;s payment rails, scheduling, persistent storage, social accounts. Every service in the human tech stack will eventually have an agent-native equivalent.</p><h2>Explore the data</h2><p>Every project mentioned here has a quality-scored page in our directories, updated daily. Browse the <a href="https://mcp.phasetransitions.ai/agents/categories/">agent categories</a>, <a href="https://mcp.phasetransitions.ai/categories/">MCP categories</a>, or check <a href="https://mcp.phasetransitions.ai/agents/trending/">what&#8217;s trending</a> this week.</p><p>The full deep dive with the five-layer taxonomy and all repos is at <a href="https://mcp.phasetransitions.ai/insights/agent-native-infrastructure/">mcp.phasetransitions.ai/insights/agent-native-infrastructure</a>.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[New coders ship faster and 27% of new AI repos prove it]]></title><description><![CDATA[New GitHub accounts, first repos ever, 600+ commits a month. The domain experts have arrived &#8212; and they're shipping faster than you.]]></description><link>https://phasetransitionsai.substack.com/p/new-coders-ship-faster-and-27-of</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/new-coders-ship-faster-and-27-of</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Fri, 03 Apr 2026 14:13:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Something changed in 2026. PT-Edge tracks daily metrics on 226,000+ AI repositories. In 2024, 5.24% of new AI repos sustained more than 300 commits per month. In 2025, 9.45%. In 2026 so far, 27.66%. That&#8217;s a 5x increase in two years.</p><p>The new population doesn&#8217;t look like developers. Single-owner GitHub accounts, one or two repos, created in the last six months, sustaining 400-600+ commits per month. No prior open-source history. Just one person and an AI coding agent shipping at a pace that would be impossible manually.</p><h2>Who are these builders?</h2><p>Domain experts using Claude Code, Cursor, and Codex to build products in their areas of expertise. The AI writes the code. They provide the intent and domain knowledge.</p><p><a href="https://mcp.phasetransitions.ai/ml-frameworks/servers/CodeWithCJ/SparkyFitness/">SparkyFitness</a> (2,900+ stars, 684 commits/30d) is the archetype. A family fitness tracker &#8212; food, water, health &#8212; built by a single contributor on a first-time GitHub account. The code is well-structured because Claude Code writes well-structured code. The commit velocity is absurd.</p><p><a href="https://mcp.phasetransitions.ai/agents/servers/Narcooo/inkos/">Inkos</a> (2,700+ stars) is an autonomous novel-writing agent &#8212; agents write, audit, and revise novels with human review gates. Not a developer tool. A creative tool built by someone who understands narrative craft.</p><h2>Why the code is better, not worse</h2><p>Counter-intuitive finding: code from AI-assisted non-developers is often higher quality than typical developer code. Claude Code follows best practices by default &#8212; proper error handling, consistent naming, test coverage. The domain expert doesn&#8217;t know enough to override these defaults with shortcuts. The AI writes textbook code because it was trained on textbooks.</p><p>The traditional signal of code quality &#8212; clean architecture, comprehensive tests &#8212; no longer distinguishes developer-built from domain-expert-built repos. The distinguishing signal is the velocity and ownership pattern, not the code itself.</p><h2>The ecosystem enabling it</h2><p><a href="https://mcp.phasetransitions.ai/agents/servers/affaan-m/everything-claude-code/">everything-claude-code</a> (74,000+ stars, 585 commits/30d) and <a href="https://mcp.phasetransitions.ai/agents/servers/sickn33/antigravity-awesome-skills/">antigravity-awesome-skills</a> (23,800+ stars, 600 commits/30d) are themselves built by the same high-velocity single-contributor pattern. They&#8217;re lowering the barrier further. A domain expert installs Claude Code, adds a skill pack, and starts building. No setup, no configuration, no learning curve beyond &#8220;describe what you want.&#8221;</p><h2>What this means</h2><p>The developer market isn&#8217;t shrinking. It&#8217;s expanding. The addressable population of people who can build production software is growing from ~20 million trained developers to potentially hundreds of millions of domain experts with AI coding agents. Each one brings domain knowledge that developers lack.</p><p>The clinical psychologist who builds her own equestrian coaching app isn&#8217;t hiring a developer. She&#8217;s building it herself, faster, with deeper domain knowledge. The moat isn&#8217;t coding skill anymore. It&#8217;s understanding the problem.</p><h2>Explore the data</h2><p>Every project mentioned here has a quality-scored page in our directories, updated daily. Browse the <a href="https://mcp.phasetransitions.ai/ai-coding/categories/">AI coding categories</a>, <a href="https://mcp.phasetransitions.ai/agents/categories/">agent categories</a>, or check <a href="https://mcp.phasetransitions.ai/ai-coding/trending/">what&#8217;s trending</a> this week.</p><p>The full deep dive with data tables and exemplar repos is at <a href="https://mcp.phasetransitions.ai/insights/domain-expert-builders/">mcp.phasetransitions.ai/insights/domain-expert-builders</a>.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Your AI is missing evals and 1,159 repos want to fix that]]></title><description><![CDATA[RAGAS owns RAG eval. Everything else is in flux. Here's what actually works across the five layers of LLM evaluation.]]></description><link>https://phasetransitionsai.substack.com/p/your-ai-has-no-evals-and-1159-repos</link><guid isPermaLink="false">https://phasetransitionsai.substack.com/p/your-ai-has-no-evals-and-1159-repos</guid><dc:creator><![CDATA[Graham Rowe]]></dc:creator><pubDate>Fri, 03 Apr 2026 10:03:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K_V-!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe28c3cb0-becd-416c-8db7-745498021007_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You shipped an AI feature. It seems to work. Users aren&#8217;t complaining yet. But when someone asks &#8220;how do you know it&#8217;s good?&#8221; the honest answer is: you don&#8217;t. You&#8217;re eyeballing outputs, running a few manual tests, and hoping.</p><p>PT-Edge tracks 1,159 repos across 9 subcategories building LLM evaluation infrastructure. The sheer number tells the story: everyone knows this problem needs solving, and nobody agrees on how. Here&#8217;s what the data says about what&#8217;s actually production-ready.</p><h2>The vibes-based eval problem</h2><p>Start with the assumption most teams make: &#8220;I&#8217;ll review a few outputs and I&#8217;ll know if it&#8217;s working.&#8221; This doesn&#8217;t scale. A prompt change that improves 80% of outputs might silently degrade the other 20%. A model upgrade that benchmarks better overall might perform worse on your specific use case. Without automated evaluation, you won&#8217;t catch either of these until users complain.</p><p>An HN post titled &#8220;Without benchmarking LLMs, you&#8217;re likely overpaying&#8221; got 197 points in January. Someone posted &#8220;Ask HN: How are people doing AI evals these days?&#8221; and got 13 comments debating completely different approaches. Even Anthropic recently acknowledged that their safety evaluation methods aren&#8217;t keeping pace with capability improvements. If the frontier labs can&#8217;t keep up with eval, production teams have no chance without better tooling.</p><h2>Five layers, five different problems</h2><p>Most developers treat evaluation as one thing. It&#8217;s actually five:</p><p><strong>RAG evaluation</strong> is the most mature layer, and it has a clear winner. <a href="https://mcp.phasetransitions.ai/llm-tools/servers/vibrantlabsai/ragas/">RAGAS</a> (12,900+ stars, 1.2M downloads/month) breaks RAG eval into four dimensions: faithfulness, answer relevance, context precision, and context recall. It generates synthetic test sets from your production data so you don&#8217;t need manual labels. If you&#8217;re building RAG, start here.</p><p><strong>Output quality</strong> is where the confusion lives. Three approaches compete: LLM-as-judge (easy to set up, systematically biased), reference-based metrics (reliable, requires labelled data you don&#8217;t have), and programmatic checks (boring, trustworthy). <a href="https://mcp.phasetransitions.ai/llm-tools/servers/Giskard-AI/giskard-oss/">Giskard</a> (5,200+ stars) is the most complete framework &#8212; it supports all three and recently split into lightweight packages so you only install what you need. Start with programmatic checks for hard requirements, add LLM-as-judge for subjective quality. Don&#8217;t rely on LLM-as-judge alone.</p><p><strong>Code evaluation</strong> is the most tractable layer because the success criterion is binary: does the code pass the tests? <a href="https://mcp.phasetransitions.ai/llm-tools/servers/evalplus/evalplus/">EvalPlus</a> (1,700+ stars) augments standard benchmarks with 80x more test cases, catching failures that HumanEval misses.</p><p><strong>Model comparison</strong> has strong tooling. <a href="https://mcp.phasetransitions.ai/llm-tools/servers/open-compass/opencompass/">OpenCompass</a> (6,800+ stars, 58K downloads/month) covers 100+ benchmarks in a unified framework. For embeddings, <a href="https://mcp.phasetransitions.ai/embeddings/servers/embeddings-benchmark/mteb/">MTEB</a> (3,200+ stars, 968K downloads/month) is the undisputed standard.</p><p><strong>Agent evaluation</strong> is where the landscape is most chaotic. 150 repos, mostly academic benchmarks, almost nothing production-ready. <a href="https://mcp.phasetransitions.ai/agents/servers/katanemo/plano/">Plano</a> (6,000+ stars) is the closest thing to production agent testing &#8212; scenario-based evaluation with LLM judges. But the hard truth is that most teams deploying agents are still evaluating them by reviewing traces by hand.</p><h2>Only 5 newsletter mentions in 90 days</h2><p>Here&#8217;s what caught our attention. PT-Edge tracks seven major AI newsletters &#8212; Simon Willison, Latent Space, The Zvi, and others. In the last 90 days, LLM evaluation as a practical topic got just 5 mentions across 3 newsletters. Compare that to agent memory (18 mentions across all 7 newsletters in the same period). The tooling is being built, but the narrative around &#8220;how should I actually evaluate my AI?&#8221; hasn&#8217;t crystallised. Everyone&#8217;s building evaluation infrastructure &#8212; 1,159 repos &#8212; and almost nobody is writing about how to use it.</p><h2>Where this is heading</h2><p>Agent evaluation is the biggest gap and the fastest-growing subcategory. Evaluating a chatbot response is hard. Evaluating a multi-step agent that writes code, makes API calls, and modifies files is an order of magnitude harder. The tooling that emerges here will matter as much as the agent frameworks themselves &#8212; you can&#8217;t deploy agents you can&#8217;t measure.</p><p>The other shift: evaluation moving from a research activity to a CI/CD primitive. <a href="https://mcp.phasetransitions.ai/llm-tools/servers/kieranklaassen/leva/">Leva</a> (133 stars) is built for Rails apps with production data &#8212; evaluation as a framework feature, not an academic exercise. Expect more of this.</p><h2>Explore the data</h2><p>Every project mentioned here has a quality-scored page in our directories, updated daily. Browse the <a href="https://mcp.phasetransitions.ai/llm-tools/categories/">LLM eval categories</a>, <a href="https://mcp.phasetransitions.ai/agents/categories/">agent evaluation categories</a>, or <a href="https://mcp.phasetransitions.ai/embeddings/categories/">embedding benchmarks</a> to explore the full landscape.</p><p>The full deep dive with detailed project tables and quality scores is at <a href="https://mcp.phasetransitions.ai/insights/llm-eval-landscape/">mcp.phasetransitions.ai/insights/llm-eval-landscape</a>.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://phasetransitionsai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/phasetransitionsai.substack.com/subscribe"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item></channel></rss>