<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Building Autonomous Coding Agents]]></title><description><![CDATA[Build, evaluate, and ship a real coding agent — every lesson ends in pass or fail, not a demo. A 40-Lesson, Evaluation-First Course in Building Autonomous Coding Agents]]></description><link>https://codingagent.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!LDV_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f5b3248-393c-4554-8683-cdf8abcefc7e_1024x1024.png</url><title>Building Autonomous Coding Agents</title><link>https://codingagent.substack.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 02 Sep 2026 01:52:39 GMT</lastBuildDate><atom:link href="/__u/codingagent.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[InFin]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[codingagent@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[codingagent@substack.com]]></itunes:email><itunes:name><![CDATA[AI Roadmap]]></itunes:name></itunes:owner><itunes:author><![CDATA[AI Roadmap]]></itunes:author><googleplay:owner><![CDATA[codingagent@substack.com]]></googleplay:owner><googleplay:email><![CDATA[codingagent@substack.com]]></googleplay:email><googleplay:author><![CDATA[AI Roadmap]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Doors Open — Lesson 0 Is Live]]></title><description><![CDATA[New lesson, most weeks, starting now. If you want the honest version of what building an AI agent takes.]]></description><link>https://codingagent.substack.com/p/doors-open-lesson-0-is-live</link><guid isPermaLink="false">https://codingagent.substack.com/p/doors-open-lesson-0-is-live</guid><dc:creator><![CDATA[AI Roadmap]]></dc:creator><pubDate>Sun, 30 Aug 2026 14:37:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/578c90df-038f-4d40-9918-337c7caef8fc_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Nine posts ago I told you about an agent that solved a real GitHub issue, cleanly, on the first try. Then I ran it again on a second issue and it failed &#8212; and I asked you to explain why. Most of you couldn&#8217;t. Neither could I, the first time I saw it happen. Nothing in the transcript tells you whether the model reasoned badly, the patch didn&#8217;t apply, the test was already broken before the agent touched anything, or the agent found the hidden test and cheated. All four look identical from the outside.</p><p>That gap &#8212; between an agent that looks like it works and an agent you can prove works &#8212; is the entire course. It starts today.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://codingagent.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Building Autonomous Coding Agents is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><strong>What&#8217;s live:</strong> Lesson 0. Free, no signup wall. Clone the repo, run one command, watch the same agent succeed once and fail once, and write down your own guess for why &#8212; sealed in a file the course checks against your answer nine lessons from now. Thirty minutes. No gate, no grading. Just the question.</p><p><strong>What comes after:</strong> 54 lessons across ten phases, each one shipping something that either passes or doesn&#8217;t &#8212; a <code>make verify</code> you run yourself, not a quiz. Ten gates mark the real turns: gold patches reproducing at 100% before any agent exists, a swarm forced to beat a single agent at equal cost or lose the argument, a held-out benchmark you never tuned against, a stranger reproducing your own result from your public repo with no help from you.</p><p>Three ways to take it, depending on what you actually want:</p><ul><li><p><strong>Core</strong> &#8212; build it, measure it, ship it.</p></li><li><p><strong>Full</strong> &#8212; the above, plus retrieval, multi-agent, and production operations.</p></li><li><p><strong>Evaluation-only</strong> &#8212; just the harness. If you want to be the person who can tell whether a number is real, this is the fast path.</p></li></ul><p>Every lesson lives in two places on purpose: the article here, the full runnable code on GitHub, tagged per lesson, frozen once published. A lesson number, once it ships, never moves again &#8212; link to it in six months and it still points at the same thing.</p><p>New lesson, most weeks, starting now. If you want the honest version of what building an AI agent takes &#8212; including the parts where the plan was wrong and had to be corrected in public &#8212; Lesson 0 is below.</p><p><strong>[Run <a href="/__u/codingagent.substack.com/p/watch-it-work-then-watch-it-fail">Lesson 0</a> &#8594;]</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://codingagent.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Building Autonomous Coding Agents is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Watch It Work. Then Watch It Fail.]]></title><description><![CDATA[Before Phase 0. Thirty minutes. The only lesson in this course with no gate.]]></description><link>https://codingagent.substack.com/p/watch-it-work-then-watch-it-fail</link><guid isPermaLink="false">https://codingagent.substack.com/p/watch-it-work-then-watch-it-fail</guid><dc:creator><![CDATA[AI Roadmap]]></dc:creator><pubDate>Sat, 29 Aug 2026 10:17:25 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/894d0b71-675b-44df-8ee9-50e6618f5c4a_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You didn&#8217;t sign up to build a scoreboard. You signed up to build an agent, and the next ten lessons don&#8217;t let you &#8212; they make you build the instrument that grades one instead. That&#8217;s the correct order for reasons Lesson 2 argues at length. It&#8217;s also, for the first thirty minutes of a course you just started, a hard sell.</p><p>This lesson is the hard sell. No harness, no gate, no code you&#8217;ll keep. Just watch something work, then watch the exact same kind of thing fail in a way you cannot explain &#8212; and carry that unanswered question into everything that follows.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://codingagent.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Building Autonomous Coding Agents is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><strong>Status:</strong> this lesson is specified but its code hasn&#8217;t been built yet &#8212; unlike Lessons 1&#8211;10, there&#8217;s no <code>apr demo</code> on GitHub to actually run today. What follows describes the intended experience and interface; the commands below are the design, not yet a working build.</p><h2><strong>Why this exists at all</strong></h2><p>Every lesson from here through Lesson 10 postpones the thing you actually want to build. That deferral is correct, but correct isn&#8217;t the same as motivating, and a reader who doesn&#8217;t feel the cost of <em>not</em> measuring won&#8217;t feel the relief of finally being able to. So before Phase 0 asks for patience, it shows you exactly what patience is buying.</p><h2><strong>What you&#8217;ll actually run</strong></h2><p>Clone the repo, supply your own model API key, and run one command. A pinned, dependency-frozen agent &#8212; a model with a bash shell, roughly eighty lines of control flow, nothing clever &#8212; is pointed at a real SWE-bench instance and left to work. It reads the issue, explores the repo, writes a patch, and stops. Read the diff. It&#8217;s a real fix, produced by a real model, and it&#8217;s genuinely impressive the first time you watch it happen live instead of reading about it in a paper.</p><p>Then point the same agent at a second prepared instance. Same code, same model, same amount of time. It fails &#8212; the patch it produces doesn&#8217;t resolve the issue.</p><h2><strong>The question you can&#8217;t answer, and why that&#8217;s the point</strong></h2><p>Open the second run&#8217;s transcript and try to say why it failed. You won&#8217;t be able to, and it isn&#8217;t because you&#8217;re new to this. Four explanations are equally consistent with everything in front of you: the model reasoned about the bug incorrectly. The patch it wrote doesn&#8217;t apply cleanly to the actual file. The test suite was already broken before the agent touched anything. Or the agent found and read the hidden test file, and &#8220;fixed&#8221; the issue by special-casing exactly what the test checks rather than the underlying bug.</p><p>Each of those is a different problem with a different owner. A wrong-reasoning failure needs a better model or a better localization step. A patch-application failure needs the diff-extraction discipline Lesson 6 builds. A pre-broken test needs the environment validation from Lesson 4. A leaked test needs the firewall Lesson 8 builds and then deliberately breaks to prove it catches. Right now, from the outside, all four look identical: a red X and a transcript that doesn&#8217;t distinguish between them.</p><p>That&#8217;s not a gap in your understanding. It&#8217;s a gap in your <em>tooling</em> &#8212; you have no instrument capable of telling these apart, because you haven&#8217;t built one yet. Every lesson in Phase 0 exists to close exactly this gap, one failure mode at a time, until &#8220;not resolved&#8221; stops being one bucket and becomes a diagnosis. Lesson 9 gives that diagnosis a name for each of these cases specifically.</p><h2><strong>Predict now, or find out later?</strong></h2><p><strong>Write down a guess before you know the answer, or skip straight to Lesson 1?</strong> Skipping costs you nothing today and everything three lessons from now. If you don&#8217;t commit to an explanation while the failure is fresh and the four possibilities are still open, you&#8217;ll retroactively convince yourself you always knew which one it was &#8212; the same hindsight bias this course&#8217;s claims ledger exists to prevent everywhere else. Writing one sentence now, sealed and unread until Lesson 9&#8217;s failure taxonomy actually exists to check it against, costs you thirty seconds and gives you a real, falsifiable prediction instead of a memory you can&#8217;t trust.</p><p><strong>A real model call, or a scripted stand-in?</strong> A scripted agent that fakes success and failure on cue would make this lesson safe, reproducible, and pointless. The entire experience this lesson delivers &#8212; genuine surprise at the first result, genuine confusion at the second &#8212; depends on the model actually reasoning about a real bug it hasn&#8217;t seen the answer to. That means this is the one lesson in the course that needs your own API key and produces a result nobody, including whoever wrote this course, can fully predict in advance.</p><h2><strong>What ships</strong></h2><p><strong>1. The pinned agent.</strong> A single file, dependency-frozen, deliberately unambitious: a system prompt, a bash tool, a loop that stops when the model says it&#8217;s done or a turn cap is hit. No planner, no reviewer, no retrieval &#8212; the minimal scaffold this course keeps returning to as a baseline through Phase 1 and beyond.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;005ccde8-f73c-4078-858d-c6595bacb33f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># demo_agent.py &#8212; the shape of the whole thing
while turns &lt; MAX_TURNS:
    response = model.call(system_prompt, history, tools=[bash_tool])
    if response.is_done:
        break
    history.append(run_bash_tool(response.tool_call))
    turns += 1</code></pre></div><p><strong>2. Two prepared instances.</strong> Real SWE-bench issues, chosen so that one resolves and one doesn&#8217;t under this minimal scaffold &#8212; which one does which is not disclosed anywhere in this lesson, on purpose.</p><p><strong>3. Full transcript capture.</strong> Every tool call, every model response, saved to disk &#8212; the same discipline Lesson 10 formalizes into a trace store with a schema, applied here by hand.</p><h2><strong>Run it</strong></h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;46912fab-8510-4a36-89f3-b94b7d96f4f1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">export APR_MODEL_API_KEY=...   # your own key; nothing in this lesson works without one
make setup
apr demo --instance a          # watch it resolve
apr demo --instance b          # watch it not</code></pre></div><p>Both runs write a full transcript to <code>.apr/demo/</code>. Read both before moving on &#8212; the transcript is the artifact this lesson is actually about, not the pass/fail line at the end of it.</p><h2><strong>Seal your prediction</strong></h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;91eb1de5-3109-455f-9999-33d34d2da608&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">apr demo predict "your one-sentence guess" &gt; predictions.txt</code></pre></div><p>One sentence: your best guess at <em>why</em> instance B failed, chosen from the four explanations above or a fifth you&#8217;ve spotted yourself. This file is sealed &#8212; nothing in this lesson or the next nine checks it against anything. Lesson 9 does, once its failure taxonomy actually exists to grade your guess against.</p><h2><strong>Where this actually confuses people</strong></h2><ul><li><p><strong>&#8220;It worked both times&#8221;</strong> &#8212; rerun it. This scaffold is deliberately minimal and the two instances were chosen to be right at the edge of what it can do; a different model, a different day, or a different temperature can occasionally flip one. If it happens, that&#8217;s data too &#8212; note it, and move on.</p></li><li><p><strong>&#8220;I don&#8217;t have an API key&#8221;</strong> &#8212; this is the one lesson in the course you can&#8217;t fully complete without one. If you&#8217;re not ready to spend on model calls yet, read the transcripts in <code>fixtures/reference_transcripts/</code> instead &#8212; real captured runs from the reference agent, not a substitute for running it yourself, but enough to feel the same confusion.</p></li><li><p><strong>&#8220;The transcript is too long to read&#8221;</strong> &#8212; that&#8217;s not a bug, it&#8217;s the point. A real agent transcript against a real repository is long and noisy in exactly the way a five-line summary can&#8217;t convey. Skim for the moment the strategy commits to something, and look there first.</p></li></ul><h2><strong>What Lesson 1 assumes you now believe</strong></h2><p>You cannot tell, from a transcript alone, why an agent failed &#8212; and that gap is expensive precisely because it&#8217;s invisible until someone builds the instrument to close it. Lesson 1 starts that build, with the least glamorous piece imaginable: counting tokens correctly, because every other measurement in this course depends on getting that one right first.</p><h2><strong>Get the code</strong></h2><p><strong>Not yet shipped.</strong> Every other lesson in this course ships code before its article publishes &#8212; this one&#8217;s the exception, on purpose disclosed rather than papered over: the scaffold described above (<code>demo_agent.py</code>, the two prepared instances, reference transcripts) hasn&#8217;t been built or pushed to <code>github.com/sysdr/coding-agent</code> yet, and running it end to end needs a real model API key this environment doesn&#8217;t hold. Building it is the next concrete step before this lesson can honestly claim the tree above.</p><p><strong>Next:</strong> Lesson 1 &#8212; LLM Foundations for Agentic Coding</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://codingagent.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Building Autonomous Coding Agents is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>