<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Vibe Coding Weekly]]></title><description><![CDATA[Weekly news and updates related to Vibe Coding]]></description><link>https://vibecodingweekly.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!jCUw!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6cb716-b087-4781-8595-f69b8b23e27a_1024x1024.png</url><title>Vibe Coding Weekly</title><link>https://vibecodingweekly.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 04:58:21 GMT</lastBuildDate><atom:link href="/__u/vibecodingweekly.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Vibe Coding Weekly]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[vibecodingweekly@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[vibecodingweekly@substack.com]]></itunes:email><itunes:name><![CDATA[Angel Llosa]]></itunes:name></itunes:owner><itunes:author><![CDATA[Angel Llosa]]></itunes:author><googleplay:owner><![CDATA[vibecodingweekly@substack.com]]></googleplay:owner><googleplay:email><![CDATA[vibecodingweekly@substack.com]]></googleplay:email><googleplay:author><![CDATA[Angel Llosa]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Vibe Coding Weekly #46]]></title><description><![CDATA[Qwen&#8217;s open-weight Flash model beats Opus 4.6 on SWE-bench Pro, Replit makes model routing the default, and a prompt injection gets past Claude Code&#8217;s Auto Mode 80% of the time.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-46</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-46</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 31 Aug 2026 05:30:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!AxuK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Everything that mattered this week in AI-assisted development, distilled into three headlines, one must-read, and the takeaways behind them.</p><p><strong>This week, compiled:</strong></p><ul><li><p><strong>The Big Story:</strong> Alibaba released <strong>Qwen3.8-Flash-Next</strong>, an open-weights model with 6B active parameters per token that scored <strong>62.5% on SWE-bench Pro</strong> against <strong>53.4%</strong> for Claude Opus 4.6 Max. Z.ai&#8217;s <strong>GLM-5.3-Flash</strong> landed the same day at a claimed tenth of GLM-5.2&#8217;s price</p></li><li><p><strong>The Tool:</strong> <strong>Replit Agent</strong> now picks the model for you by default, balancing cost, speed and accuracy by sending hard reasoning to <strong>Claude Opus 4.7</strong> and bulk generation to <strong>Gemini 3.1 Pro</strong></p></li><li><p><strong>The Trend:</strong> Security researcher Johann Rehberger got an indirect prompt injection past <strong>Claude Code&#8217;s Auto Mode about 80% of the time</strong>, in the same week Anthropic shipped a <code>--restricted</code> flag and closed three symlink and logging holes</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> Scott Fryxell&#8217;s <strong><a href="https://scott-fryxell.github.io/blog/the-harness-is-the-thing/">The Harness Is The Thing</a></strong> argues that the model everyone argues about online is the commodity, and the scaffolding around it is the asset you actually own. He rotates three editors, Cursor, Claude and Pi, over a single harness he calls <em>brayness</em>, which carries standardized <code>AGENTS.md</code> files between them so switching tools costs him nothing. Underneath sits a staged workflow of exploration, planning, execution, criticism and promotion, where a frontier model handles planning and promotion while <strong>DeepSeek V4 Flash</strong> grinds through the implementation. That split cut his frontier-model usage by <strong>75%</strong>. He is also honest about where it breaks: once a task has enough moving parts, the cheap model hits a ceiling and no harness rescues it. The line worth stealing is his definition of the thing itself, &#8220;the fulcrum from which my expectations meet the LLM&#8217;s capabilities.&#8221; Read it against the rest of this week and it stops being one developer&#8217;s setup. Replit shipped automatic routing, Anthropic shipped a cost profiler, and two open-weight models arrived that are exactly what a harness like his would route to. <a href="https://scott-fryxell.github.io/blog/the-harness-is-the-thing/">Read more &#8594;</a></p></blockquote><div><hr></div><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>In Vibe Coding Weekly I try to cut through that volume so you arrive at the week with context, not anxiety.</p><p>I&#8217;m <strong>Angel Llosa</strong>, and my day job is getting these tools adopted inside real companies &#8212; defining the strategy and then implementing it. I read this news wondering what survives contact with an actual engineering team.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p><p>If this is the filter you want on your Monday, it arrives every week.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Key Takeaways</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!AxuK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AxuK!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png 424w, /__u/substackcdn.com/image/fetch/$s_!AxuK!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png 848w, /__u/substackcdn.com/image/fetch/$s_!AxuK!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AxuK!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AxuK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png" width="1456" height="2912" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2912,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:531995,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/213437696?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AxuK!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png 424w, /__u/substackcdn.com/image/fetch/$s_!AxuK!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png 848w, /__u/substackcdn.com/image/fetch/$s_!AxuK!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AxuK!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e71fb5-5ca7-4b08-af0c-ba5cea4130f4_1920x3840.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p><strong>Two open-weight models landed on the same day, and one of them beat Opus on a coding benchmark:</strong> Alibaba&#8217;s <strong><a href="https://www.orcarouter.ai/blog/qwen-3-8-flash-release">Qwen3.8-Flash-Next</a></strong> is a 125B-parameter mixture-of-experts model that activates just <strong>6B parameters per token</strong>, and it previews the Qwen4 architecture. It scored <strong>62.5% on SWE-bench Pro</strong> and <strong>73.9% on CoWorkBench</strong>, against <strong>53.4%</strong> and <strong>68.2%</strong> for Claude Opus 4.6 Max, with 262K native context extending to 1M through YaRN. Alibaba says it trained for roughly a ninth of what its predecessor cost. Hours later Z.ai released <strong><a href="https://www.marktechpost.com/2026/08/26/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context/">GLM-5.3-Flash</a></strong>, 320B parameters with 18B active, MIT-licensed weights on Hugging Face and a 1,048,576-token context, using a hybrid of sparse and linear attention to cut attention compute about 3x and shrink the KV cache 4x at long context. Z.ai claims it beats GLM-5.2 while costing a tenth as much. Benchmark numbers from a vendor are worth what you paid for them, but the direction has been consistent for months now, and the gap you are paying frontier prices to avoid keeps getting narrower. <a href="https://www.orcarouter.ai/blog/qwen-3-8-flash-release">Read more &#8594;</a></p></li><li><p><strong>GitHub turned those models off by default for enterprise Copilot the same week, and it got this right:</strong> the <strong>global model policy</strong> went generally available on <strong>August 26</strong>, rolling out through <strong>September 1</strong>. Previously-unconfigured models, and every new GA model from here on, now inherit a global policy state instead of arriving switched on. Open-weight models and anything that requires data retention are <strong>disabled unless an admin explicitly opts in</strong>. Admins can still override per model, or turn the policy off for orgs with different compliance needs. It would be easy to file this as procurement getting in the way of a cheaper model. Look at what the policy actually gates and you get a different picture: the trigger is <strong>data retention</strong>, not price. If you work under GDPR, or you handle other people&#8217;s client data, a newly released model quietly switching itself on across the whole org is precisely the event you cannot afford to discover afterwards. Defaulting to off puts that call in front of someone accountable for it. The cost case for open weights survives this intact; it just has to be argued out loud instead of arriving by default. <a href="https://github.blog/changelog/2026-08-26-global-model-policy-generally-available/">Read more &#8594;</a></p></li><li><p><strong>Replit stopped asking which model you want and started deciding:</strong> Auto mode is now the <strong>default</strong> in Replit Agent, routing each request on speed, cost and accuracy across a pair of models. <strong>Claude Opus 4.7</strong> takes the harder logic and reasoning; <strong>Gemini 3.1 Pro</strong> handles bulk code generation and multimodal work. The manual selector survives for anyone who wants to force a choice per task. Routing has been the theory behind every token-economics post of the last six months, and Snowflake shipped it at the gateway layer a week ago. Replit is the first mainstream coding product to make it the behaviour you get without asking, which quietly changes what a &#8220;model choice&#8221; even means for the average user of the product. <a href="https://thenewstack.io/replit-model-routing-default/">Read more &#8594;</a></p></li><li><p><strong>Claude Code shipped a command that profiles your own API bill, then spent the week patching itself:</strong> seven releases between <strong>August 23 and 28</strong> (v2.1.241 through v2.1.251). <strong>v2.1.247</strong> added <code>/claude-api cost-optimize</code>, which profiles a project&#8217;s Claude API spend and walks you through the levers one measured change at a time: caching, token hygiene, batching, effort level, model choice. <strong>v2.1.248</strong> brought <code>--restricted</code>, a flag that strips command execution, code execution and WebFetch while keeping in-directory file tools and refusing <code>bypassPermissions</code> outright, plus a per-agent <code>experimental.cacheTtl</code> and cross-session messaging between sessions on the same machine. Then <strong>v2.1.251</strong> closed three genuine holes: file tools following a symlink swapped in after the permission check, plugin marketplace commands reaching outside the plugin directory, and project settings that could enable raw API body logging and bypass a pinned OTLP collector. A cost profiler inside the tool is the FinOps story; a symlink race after the permission check is the one your security team will ask about. <a href="https://github.com/anthropics/claude-code/releases">Read more &#8594;</a></p></li><li><p><strong>An indirect prompt injection walked past Auto Mode roughly 80% of the time:</strong> Johann Rehberger&#8217;s attack, <a href="https://simonwillison.net/2026/Aug/27/breaking-claude-code-opus-5-auto-mode/">written up by Simon Willison</a> on <strong>August 27</strong>, starts with something as ordinary as asking the agent to summarize a web page. The malicious page persuades Claude Code to write its own decoder, importing <code>base64</code> to unpack the payload, and Python module shadowing does the rest: a <code>struct.py</code> pulled in with a downloaded archive gets executed instead of the standard library module. In some runs Auto Mode&#8217;s own safety check then blocked the agent when it tried to <strong>kill the malicious process it had just started</strong>. Anthropic has positioned Auto Mode as the answer to prompt injection in autonomous coding; this is a careful demonstration that it is a speed bump, not a boundary, and it landed days after <code>--restricted</code> shipped. <a href="https://simonwillison.net/2026/Aug/27/breaking-claude-code-opus-5-auto-mode/">Read more &#8594;</a></p></li><li><p><strong>Copilot chat logs go from 28 days to the lifetime of your account:</strong> GitHub bundled three changes into one <strong>August 28</strong> notice, and the retention change is the one with consequences. When Copilot Chat on github.com and Mobile merges with the cloud agent into a single experience, <strong>no earlier than September 28</strong>, chat data retention extends from <strong>28 days to the lifetime of the account</strong> to match what the cloud agent already does, and opting out means losing access to the merged experience. Alongside it: Business and Enterprise seat assignments now require <strong>upfront payment</strong>, from <strong>September 1</strong> for new signups and <strong>October 1</strong> for existing customers, with <strong>no prorated refund</strong> when a seat is revoked mid-cycle. Code review&#8217;s default effort level also moves from Lite to Balanced on September 28. Base subscription prices do not change, which is precisely why this will get filed as housekeeping and land on someone&#8217;s desk in October as a surprise. <a href="https://github.blog/changelog/2026-08-28-upcoming-changes-to-github-copilot-policies-and-billing/">Read more &#8594;</a></p></li></ul><div><hr></div><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. <strong>Included with every subscription.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://cursor.com/changelog/start-from-scratch">Cursor&#8217;s cloud agents no longer need a repository to start</a></h3><p><em>Cursor &#8212; August 27, 2026</em></p><p>You can now prompt a Cursor Cloud Agent from nothing at all: no GitHub connection, no source-control provider, no existing project. Cursor generates an <strong>Origin</strong> repo behind the scenes, shows you the agent working in a <strong>live browser preview</strong>, and publishes a real URL through <strong>Vercel</strong> when it&#8217;s done. Convert that scratch repo into a permanent named one only if the thing turns out to be worth keeping. Following the Origin hosting launch and always-on agents covered here last week, this is the piece that moves Cursor&#8217;s agents out of the IDE model entirely and onto Lovable and v0&#8217;s turf.</p><div><hr></div><h3><a href="https://learn.chatgpt.com/docs/changelog">Codex CLI deprecates its MCP server and lets extensions rewrite tool results</a></h3><p><em>OpenAI &#8212; August 24&#8211;29, 2026 (v0.149.1, v0.150.0, v0.150.1, v0.151.0)</em></p><p>Anyone driving Codex from another harness has rewiring to do: <strong>v0.149.1</strong> deprecated the <code>codex mcp-server</code> command in favour of the Codex app server, which includes the Claude Code plugin path. <strong>v0.150.0</strong> is the release users will feel, with <code>@</code> task mentions, a better <code>/copy</code> picker, auto-generated task titles, clickable Markdown links in capable terminals, and <code>Interrupt</code> hooks. The one worth watching is <strong>v0.151.0</strong> on August 29: extensions can now <strong>inspect and replace MCP tool results before the model ever sees them</strong>, MCP tool-discovery grace periods became configurable, and two sandbox bugs got closed, including stale Guardian classifications surviving a permission change and <code>/cd</code> quietly weakening sandbox restrictions.</p><div><hr></div><h3><a href="https://ampcode.com/chronicle">Amp gives one agent multiple repositories, and ships a phone app</a></h3><p><em>Sourcegraph &#8212; August 23&#8211;28, 2026</em></p><p>The substantive change came on <strong>August 27</strong>: an Amp project can now span <strong>multiple repositories</strong>, so a single job can cross codebases instead of stopping at the repo boundary that most agent tooling still treats as the world. The TUI sidebar went away the same day. On <strong>August 28</strong> Amp shipped <strong>native iOS and macOS apps</strong> carrying Threads, Orbs and Puck, and added a setting to choose the verified name and email Amp uses to author and sign commits, which finally kills the mailmap workaround. Earlier in the week: friendly hostnames for shared Orb portals, and Orb configuration before or after cloning without committing Amp-specific files to the repo.</p><div><hr></div><h3><a href="https://github.com/sst/opencode/releases">OpenCode adds Azure sign-in without an API key</a></h3><p><em>OpenCode &#8212; August 24&#8211;28, 2026 (v1.18.22 &#8594; v1.18.25)</em></p><p>Four patch releases, one useful addition: <strong>v1.18.24</strong> lets you sign in through <strong>Azure Entra ID using the Azure CLI</strong>, so Azure-hosted models work with no API key in a config file anywhere. It also fixed a Bedrock reasoning-cache bug and kept V1 configs able to read newer V2 fields, with <strong>v1.18.25</strong> following hours later to make that auth path work without requiring Bun. The rest of the week went to provider plumbing: <code>textVerbosity</code> being sent to providers that reject it, Bedrock compatibility, and Cloudflare AI Gateway routing for third-party providers including Anthropic model-slug conversion.</p><div><hr></div><h3><a href="https://antigravity.google/">Antigravity CLI makes switching models a single command</a></h3><p><em>Google &#8212; August 27, 2026</em></p><p><strong>v1.1.22</strong> adds a <code>/model</code> argument that takes a name, slug or label and sets the default in the same step, with inline ghost-text autocomplete for model IDs while you type. The <code>/effort</code> hint now completes dynamically against what you&#8217;ve typed, and startup and project-switching latency both came down again. Small week for Google&#8217;s Gemini CLI successor after the Gemini Enterprise integration and Remote Control, but the model switcher is the right thing to make frictionless in a tool whose whole premise is that you&#8217;ll be changing models often.</p><div><hr></div><h3><a href="https://github.com/Kilo-Org/kilocode/releases">Kilo Code renders Mermaid diagrams inside JetBrains</a></h3><p><em>Kilo Code &#8212; August 26&#8211;28, 2026 (jetbrains/v7.1.0, jetbrains/v7.1.1, v7.5.0&#8211;v7.5.6)</em></p><p>JetBrains <strong>v7.1.1</strong> renders <strong>Mermaid diagrams directly in the chat panel</strong>, with a dedicated editor-tab view, zoom, pan, fit-to-view, a source view and copy-as-PNG. If you&#8217;ve been asking an agent to explain an unfamiliar service and getting back a wall of diagram syntax, that gap just closed on the JetBrains side. The same window brought clearer empty states for worktree sessions, while the VS Code line moved through v7.5.0 to v7.5.6 on stability work with no headline feature.</p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel Llosa</p><p>Questions, disagreements, or a story I missed? Just hit reply.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Vibe Coding Weekly&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share Vibe Coding Weekly</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #45]]></title><description><![CDATA[Cursor launches its own code hosting to take on GitHub, OpenAI cuts frontier pricing by more than 20%, and DeepSeek's harness becomes the fastest-starred project in GitHub's history.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-45</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-45</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Sun, 23 Aug 2026 11:22:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!M6YB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Everything that mattered this week in AI-assisted development, distilled into three headlines, one must-read, and the takeaways behind them.</p><p><strong>This week, compiled:</strong></p><ul><li><p><strong>The Big Story:</strong> OpenAI cut <strong>GPT-5.6 Sol</strong> API pricing to <strong>$4 per 1M input</strong> and <strong>$20 per 1M output</strong> tokens, down from $5 and $30 &#8212; a 20% input and 33% output cut, promotional through at least <strong>November 21, 2026</strong>, and framed openly as a response to Anthropic and Chinese open-weight models</p></li><li><p><strong>The Tool:</strong> GitHub put Copilot&#8217;s agent inside <strong>Slack and Microsoft Teams</strong> on the same day. Mention <code>@GitHub</code>, and it investigates the failure, patches it in a sandbox and opens the PR &#8212; in a session the whole team can watch and redirect</p></li><li><p><strong>The Trend:</strong> The agent runtime itself is going open source. <strong>DeepSeek&#8217;s Harness</strong> became the <strong>fastest-starred project in GitHub&#8217;s history</strong> &#8212; <strong>185,893 stars</strong> and <strong>20,595 forks</strong> ten days in &#8212; while <strong>TrueFoundry</strong> open-sourced <strong>TrueForge</strong>, an MIT-licensed harness finishing enterprise tasks <strong>30&#8211;75% cheaper</strong> than a managed alternative</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> Cursor spent the week building both the place code lives and the agents that work on it unattended. On <strong>August 17</strong> it launched <strong><a href="https://cursor.com/changelog/origin-code-hosting">Origin</a></strong> in early beta for all paid plans &#8212; repositories hosted inside Cursor with their own URLs, full pull-request workflows with diffs, comments and merging, and bidirectional sync with GitHub, which stays the source of truth. Two days later its <strong><a href="https://cursor.com/changelog/08-19-26">cloud agents went always-on</a></strong>: they can subscribe to a pull request, a Slack thread or a schedule and activate on their own to drive it to completion or fix CI, spawn subagents on isolated VMs with clean copies of the project, and accept a <code>/goal</code> they keep pursuing until it is actually done. An IDE company shipping a GitHub competitor would be the story in most weeks. Shipping it alongside agents that work while nobody is watching is the real one &#8212; those agents need a place to work that answers to the editor, not to someone else&#8217;s platform. <a href="https://cursor.com/changelog/origin-code-hosting">Read more &#8594;</a></p></blockquote><div><hr></div><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>In Vibe Coding Weekly I try to cut through that volume so you arrive at the week with context, not anxiety.</p><p>I&#8217;m <strong>Angel Llosa</strong>, and my day job is getting these tools adopted inside real companies &#8212; the technical wiring and the strategy around it. I read this news wondering what survives contact with an actual engineering team.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p><p>If someone on your team should see this week&#8217;s edition, forward it to them.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/p/vibe-coding-weekly-45?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/p/vibe-coding-weekly-45?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p><div><hr></div><h2>Key Takeaways</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!M6YB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!M6YB!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png 424w, /__u/substackcdn.com/image/fetch/$s_!M6YB!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png 848w, /__u/substackcdn.com/image/fetch/$s_!M6YB!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png 1272w, /__u/substackcdn.com/image/fetch/$s_!M6YB!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!M6YB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png" width="1456" height="2912" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/affea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2912,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:561685,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/212393043?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!M6YB!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png 424w, /__u/substackcdn.com/image/fetch/$s_!M6YB!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png 848w, /__u/substackcdn.com/image/fetch/$s_!M6YB!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png 1272w, /__u/substackcdn.com/image/fetch/$s_!M6YB!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faffea579-0101-4806-a6cd-3211a23c8eea_1920x3840.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p><strong>OpenAI cut its frontier model&#8217;s price by a third on output, and said why out loud:</strong> <strong>GPT-5.6 Sol</strong> now costs <strong>$4 per 1M input tokens</strong> and <strong>$20 per 1M output</strong> for standard short-context use, down from <strong>$5 and $30</strong> &#8212; a 20% input cut and a 33% output cut, promotional through at least <strong>November 21, 2026</strong> and extending to API credits for ChatGPT Work and Codex. Subscription pricing for Pro, Plus and Business is untouched, so this is aimed squarely at the people building on the API. Reuters reports the framing without much diplomacy: pressure from Anthropic and from Chinese open-weight models. When frontier output tokens fall by a third mid-quarter, every agent cost model your team built in July is now wrong in your favour. <a href="https://www.thestar.com.my/tech/tech-news/2026/08/22/openai-cuts-developer-pricing-for-frontier-gpt-56-sol-model-by-more-than-20">Read more &#8594;</a></p></li><li><p><strong>GitHub moved the coding agent into the chat window your team already lives in:</strong> on <strong>August 21</strong> Copilot&#8217;s agentic capabilities arrived in <strong><a href="https://github.blog/changelog/2026-08-21-the-new-github-copilot-experience-in-slack/">Slack</a></strong> and <strong><a href="https://github.blog/changelog/2026-08-21-shared-agentic-work-with-github-copilot-in-microsoft-teams/">Microsoft Teams</a></strong> in public preview on the same day. Mention <code>@GitHub</code> and it answers questions about the codebase, investigates a failing build, implements the fix in a secure cloud sandbox and opens a pull request linked back to the thread &#8212; and the session is <strong>shared</strong>, so the whole channel watches and steers rather than one person relaying. Slack gets dedicated <strong>Code channels</strong> for reviewing diffs together; Teams turns a meeting thread into work that keeps running after the call ends. It is available on paid Copilot plans against existing entitlements, sandbox usage billed separately, and admins can require PR approval &#8212; which is the detail that decides whether your security team says yes. <a href="https://github.blog/changelog/2026-08-21-the-new-github-copilot-experience-in-slack/">Read more &#8594;</a></p></li><li><p><strong>Snowflake made the gateway choose the model, and cut a pipeline agent to a third of the tokens:</strong> Cortex AI Gateway now routes each request automatically instead of leaving the choice to whoever wrote the code, using two mechanisms &#8212; an <strong>advisor pattern</strong> where a small model attempts the task first and escalates to a larger one only when it cannot finish, and a <strong>classifier trained on past queries</strong> that sends simple questions to simple models. In internal testing on a data-pipeline agent task, the routed setup used <strong>a third of the tokens</strong> of a frontier-only approach at equivalent quality. The governance design is what makes it usable in an enterprise: the router only picks from <strong>admin-approved models</strong> and honours data-residency settings <em>before</em> it weighs cost or latency. Routing has been the theory for months; this is it shipping inside a platform companies already run their data on. <a href="https://venturebeat.com/orchestration/enterprises-are-overpaying-for-simple-ai-queries-snowflakes-gateway-now-auto-routes-to-cut-costs-up-to-3x">Read more &#8594;</a></p></li><li><p><strong>TrueFoundry&#8217;s open-source runtime finished enterprise tasks 75% cheaper than a managed agent:</strong> TrueFoundry released <strong>TrueForge</strong> under an <strong>MIT licence</strong> &#8212; an agent harness from ex-Meta engineers built around cost rather than capability. On DevRev&#8217;s <strong>Enterprise-Bench</strong>, TrueForge running <strong>GLM-5.2</strong> completed tasks for <strong>$2.90 against $11.80</strong> for Claude Managed Agents on Opus 4.8 &#8212; 75% cheaper &#8212; and when both sides ran Opus 4.8, TrueForge still came in around <strong>30% cheaper at $8.50</strong>. That second number is the honest one, because it isolates the harness from the model: the savings come from context engineering &#8212; delayed tool-schema loading, offloading large results to files, subagent delegation, automatic compaction &#8212; and from provisioning sandboxes only when a task actually needs one. The uncomfortable implication for anyone budgeting agents: a meaningful share of your bill is the scaffolding, not the intelligence. <a href="https://venturebeat.com/orchestration/truefoundrys-open-source-ai-agent-harness-trueforge-boasts-30-75-cheaper-task-completion-than-claude-managed-agents">Read more &#8594;</a></p></li><li><p><strong>DeepSeek&#8217;s plugin-first harness became the fastest-starred project in GitHub&#8217;s history, and the number kept climbing:</strong> <strong>Harness</strong> shipped on <strong>August 13</strong> &#8212; covered here last week &#8212; and the story this week is what happened next. The repository stood at <strong>185,893 stars and 20,595 forks</strong> as of <strong>August 23</strong>, ten days after its first commit, past <strong>OpenClaw&#8217;s</strong> previous record and with more than <strong>2,000 plugin proposals</strong> filed in the first weekend alone. The design is the reason it travels: under an <strong>MIT licence</strong>, the model adapter, tool registry, session log, sandboxing, telemetry and the agent loop itself are all replaceable, and it drives third-party models rather than only DeepSeek&#8217;s. Stars are a vanity metric on their own, but 20,000 forks are not &#8212; that is twenty thousand people who wanted their own copy of an agent runtime they can take apart. <a href="https://github.com/deepseek-ai/deepseek-harness">Read more &#8594;</a></p></li><li><p><strong>LinkedIn published the acceptance rates for AI code review, broken down by what it was reviewing:</strong> engineers detailed a production multi-agent review platform where <strong>several independent models cross-validate each other&#8217;s findings</strong> to cut blind spots, layered with deep org- and repo-specific customization on a Kubernetes event-driven architecture with durable queues. The numbers are the reason to read it: <strong>63.9% suggestion acceptance overall</strong>, rising to <strong>100% on concurrency bugs</strong> and <strong>80% on logic errors</strong>, and falling to <strong>40.6% on security fixes</strong>. That spread is the most useful thing published about AI code review this month &#8212; it tells you exactly where to point the agent and where a human still has to own the call, from production data rather than a vendor benchmark. <a href="https://www.infoq.com/news/2026/08/linkedin-ai-code-review/">Read more &#8594;</a></p></li></ul><div><hr></div><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code prices in the data-residency premium and stops leaking memory</a></h3><p><em>Anthropic &#8212; August 19&#8211;23, 2026 (v2.1.236 &#8594; v2.1.241)</em></p><p>Six releases in five days, and the one worth noticing is an accounting detail: cost estimates now factor in a <strong>1.1&#215; premium for US-only inference</strong> on data-residency workspaces, so the number in your terminal matches the number on the invoice. The rest is the kind of week that makes a tool boring in the best sense &#8212; an <code>ANTHROPIC_DEFAULT_MODEL</code> env var, cross-session <code>notify_when_idle</code> messaging, a built-in <strong>Concise</strong> output style, readline-style <code>Ctrl+W</code> via <code>keybindingFlavor</code>, and <code>/claude-api upgrade</code> to migrate Python projects off the legacy SDK. On the fix side: <strong>unbounded memory growth in long interactive sessions</strong>, prompt caching breaking behind LLM gateways, and an assortment of Bedrock and proxy bugs.</p><div><hr></div><h3><a href="https://developers.openai.com/codex/changelog/">Codex CLI gets an agents dashboard, session forking, and Markdown export</a></h3><p><em>OpenAI &#8212; August 18 and 20, 2026 (v0.148.0, v0.149.0)</em></p><p>The interesting shift here is that Codex has stopped assuming you run one agent at a time. <strong>v0.149.0</strong> adds an interactive <code>codex agents</code> dashboard for searching, starting, opening, renaming and stopping tasks &#8212; a control surface for a fleet, not a session. <strong>v0.148.0</strong> came in two days earlier with <code>/export</code> to dump a full TUI conversation to <strong>Markdown</strong>, <code>codex exec fork</code> to branch a session and try a different path from the same state, archive and restore from the resume picker, and <strong>Amazon Bedrock</strong> as a built-in provider. Session forking is the underrated one: it makes &#8220;try the other approach&#8221; cost a command instead of a re-run.</p><div><hr></div><h3><a href="https://antigravity.google/blog">Google folds Antigravity into Gemini Enterprise and adds remote control</a></h3><p><em>Google &#8212; August 20&#8211;21, 2026</em></p><p>Antigravity &#8212; Google&#8217;s agent-first development environment and the successor to the discontinued Gemini CLI &#8212; moved decisively toward the enterprise this week. It now integrates with <strong>Gemini Enterprise</strong> for org-wide agentic workflows, ships dedicated <strong>IDE extensions</strong> so it stops being a separate place you go, and adds <strong>Remote Control</strong>, letting you monitor and steer agent sessions from outside the primary workspace. Google spent the first half of the year consolidating its CLI story; this is the half where it starts selling it to procurement.</p><div><hr></div><h3><a href="https://ampcode.com/chronicle">Amp adds voice control, MCP servers in orbs, and a way to ask where your tokens went</a></h3><p><em>Sourcegraph &#8212; August 18&#8211;21, 2026</em></p><p>Amp had a busy week, and one item lands right on the week&#8217;s theme: you can now <strong>ask Puck directly where an agent&#8217;s tokens went</strong> &#8212; in-product spend attribution rather than a monthly reconciliation. Alongside it, <strong>realtime voice control</strong> of agents through Puck, the ability to connect <strong>MCP servers directly to orbs and Puck</strong>, and student and teacher subscriptions cut to <strong>$10/month</strong>. Voice will get the attention, but token attribution inside the tool is the one that changes a conversation with your finance team.</p><div><hr></div><h3><a href="https://x.ai/news/grok-bot-more-plans">xAI opens Grok Bot to SuperGrok Plus and standard Cursor Teams plans</a></h3><p><em>xAI &#8212; August 21, 2026</em></p><p>Ten days after a beta locked to SuperGrok Heavy and Cursor Ultra, <strong>Grok Bot</strong> is now included with <strong>SuperGrok Plus, Cursor Pro+ and standard Cursor Teams</strong>, with a limited free trial for everyone else. The pitch remains the aggressive one: an &#8220;AI teammate&#8221; that gets <strong>its own computer</strong> and signs into your existing applications to complete work end to end, rather than suggesting what you should do in them. Going from top-tier-only to standard team plans in ten days says xAI wants adoption numbers more than it wants margin.</p><div><hr></div><h3><a href="https://freedom.tech/posts/2026-08-21-cline-desktop-0-0-15/">Cline gives agents durable todos and recurring schedules</a></h3><p><em>Cline &#8212; August 21, 2026</em></p><p>Desktop <strong>0.0.15</strong> lets agents create <strong>durable todos</strong> and <strong>one-time or recurring schedules</strong>, each scoped to the clients capable of servicing them &#8212; the difference between an agent that reacts when you open it and one that keeps a backlog. The release also renames the app from &#8220;Cline Code&#8221; to <strong>Cline</strong> (settings, sessions and credentials carry over), reworks the model selector to lead with Recommended and Free tiers, and fixes two irritating bugs: checkpoint restore getting permanently wedged, and unprompted sessions reporting a bogus &#8220;running&#8221; status.</p><div><hr></div><h3><a href="https://github.com/sst/opencode/releases">OpenCode makes failed subagent tasks resumable</a></h3><p><em>OpenCode &#8212; August 21, 2026 (v1.18.20, v1.18.21)</em></p><p>Two releases in six hours, both aimed at the failure modes that quietly kill long agent runs. Subagent tool calls now get <strong>resumable task IDs</strong> and better error handling, so a broken step no longer costs the whole task. Retry logic got wider coverage across provider variants &#8212; including <strong>Cerebras token limits</strong> and <strong>xAI capacity errors</strong> &#8212; Vertex AI EU/US multi-region Gemini requests now route through REP endpoints, and generation continues when a model reports an unknown finish reason instead of stopping early.</p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel Llosa</p><p>Questions, disagreements, or a story I missed? Just hit reply.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/p/vibe-coding-weekly-45?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/p/vibe-coding-weekly-45?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #44]]></title><description><![CDATA[Six vendors agree on one plugin format for coding agents, SpaceX closes its 60B acquisition of Cursor, and Meta ships a 30B open-weight agent that runs on one GPU.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-44</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-44</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 17 Aug 2026 06:01:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!L3jP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Everything that mattered this week in AI-assisted development, distilled into three headlines, one must-read, and the takeaways behind them.</p><p><strong>This week, compiled:</strong></p><ul><li><p><strong>The Big Story:</strong> Meta went back to open source. <strong>Muse Glimmer</strong> is a <strong>30B-parameter Apache 2.0</strong> model distilled from Muse Spark and tuned for always-on local coding agents &#8212; small enough for a single consumer GPU, and tested hands-on by Simon Willison for code exploration and tool use</p></li><li><p><strong>The Money:</strong> SpaceX closed its <strong>60B acquisition of Cursor</strong>, folding it in as a wholly owned subsidiary. Cursor&#8217;s own framing is about GPUs &#8212; access to &#8220;the largest fleet in the world&#8221; to train stronger and cheaper models</p></li><li><p><strong>The Trend:</strong> Enterprises are buying <strong>model routers</strong>. Fortune reports <strong>62% of organizations</strong> hit material budget surprises from agent usage and <strong>40%</strong> escalated to the board, pushing routing platforms that cut spend by up to <strong>30%</strong> into the procurement conversation</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> on <strong>August 12</strong>, GitHub shipped <strong>Agent Plugins 1.0</strong> &#8212; and the news is not the format, it is the signatory list. <strong>AWS, Anysphere (Cursor), Microsoft, OpenAI, Vercel and Google</strong> co-maintain it, and it is governed independently of any single vendor. One installable plugin bundles agent skills <em>and</em> MCP servers together, and the same artifact runs across <strong>VS Code, Copilot CLI, the Copilot app and SDK</strong>, plus <strong>Kiro and Cursor</strong> per AWS&#8217;s own announcement. Build your internal agent tooling once, run it wherever your teams already are &#8212; that is the closest thing yet to making &#8220;which agent do we standardize on?&#8221; the wrong question. <a href="https://github.blog/changelog/2026-08-12-agent-plugins-1-0-in-vs-code-copilot-cli-and-the-copilot-app/">Read more &#8594;</a></p></blockquote><div><hr></div><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>In Vibe Coding Weekly I try to cut through that volume so you arrive at the week with context, not anxiety. </p><p>I&#8217;m <strong>Angel Llosa</strong>, and my day job is getting these tools adopted inside real companies &#8212; the technical wiring and the strategy around it. I read this news wondering what survives contact with an actual engineering team.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!L3jP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!L3jP!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!L3jP!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!L3jP!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!L3jP!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!L3jP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2536800,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/211387960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!L3jP!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!L3jP!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!L3jP!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!L3jP!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce0ce4db-a8dd-4c90-800a-7202999f9493_2752x1536.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Key Takeaways</h2><ul><li><p><strong>Meta open-weighted a coding agent small enough to live on your laptop:</strong> <strong>Muse Glimmer</strong> is a <strong>30B-parameter</strong> model released under <strong>Apache 2.0</strong>, distilled from Muse Spark and explicitly optimized for local, always-on agentic coding &#8212; it fits on a single consumer GPU. Simon Willison put it through code exploration and tool use rather than benchmarks, which is the right test for a model whose whole pitch is that it can run continuously without a meter running. The licence is the part to sit with: a frontier lab shipping a genuinely permissive, locally-runnable agentic model changes what &#8220;we can&#8217;t send this code to an API&#8221; actually costs you. <a href="https://simonwillison.net/2026/Aug/10/introducing-muse-glimmer/">Read more &#8594;</a></p></li><li><p><strong>A rocket company now owns a coding IDE, and the reason is GPUs:</strong> SpaceX completed its <strong>60B acquisition of Cursor</strong>, taking it on as a wholly owned subsidiary via stock issuance. Cursor&#8217;s stated logic is compute &#8212; access to &#8220;the largest fleet of GPUs in the world&#8221; to build stronger and more economical models &#8212; and the deal is already operational rather than theoretical: <strong>Grok 4.6 arrived inside Cursor</strong> the same week, cited by Cursor as an early result of the relationship. Read it as the clearest sign yet that the constraint on coding tools is no longer talent or product, it is silicon, and that the companies with silicon are the ones doing the acquiring. <a href="https://cursor.com/blog/joining-spacex">Read more &#8594;</a></p></li><li><p><strong>Every company suddenly wants a model router, because the agent bills arrived:</strong> Fortune&#8217;s numbers are the ones to bring to a budget meeting &#8212; <strong>62% of organizations</strong> report material budget surprises from agent usage and <strong>40%</strong> required board escalation. The response is routing: OpenRouter, Not Diamond and LiteLLM, now joined by Salesforce and Databricks entrants, sending each task to the cheapest model that can actually do it, for reductions of up to <strong>30%</strong>. Long-running coding agents are what broke the forecast, so this lands squarely on engineering leadership rather than finance. <a href="https://fortune.com/2026/08/09/why-every-company-wants-an-ai-model-router-right-now/">Read more &#8594;</a></p></li><li><p><strong>Google&#8217;s cheap model got good at exactly the thing agents do:</strong> <strong>Gemini 3.7 Flash</strong> jumped to <strong>43.6% on FrontierCode 1.1 Main</strong> (from 34.4%) and <strong>65.3% on DeepSWE v1.1</strong> (from 49.0%), at an introductory <strong>0.75 per 1M input tokens</strong>. It rolled out same-day across <strong>GitHub Copilot</strong> &#8212; Pro, Pro+, Max, Business and Enterprise, in VS Code, Copilot CLI, JetBrains, Xcode and Eclipse. A workhorse-tier model closing this much of the agentic-coding gap at that price is precisely the pressure that makes routing strategies pay off. <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/">Read more &#8594;</a></p></li><li><p><strong>DeepSeek&#8217;s answer to agent lock-in: make every part of the runtime a plugin:</strong> <strong>Harness v0.1</strong> entered developer preview under an <strong>MIT licence</strong>, built on the Cordis kernel with a single design rule &#8212; models, tools, sessions, sandboxes, the agent loop and orchestration are all swappable. The detail that shows what that buys you: Harness can call <strong>Claude Code or Codex as sub-agents</strong> inside a DeepSeek-orchestrated workflow. It is framed explicitly against Claude Cowork, and it is the second open interoperability play of the week from an entirely different direction. <a href="https://deepseek.com/harness/en/">Read more &#8594;</a></p></li><li><p><strong>Claude Code got sub-agent forking by default and finally speaks GitLab:</strong> <strong>v2.1.232</strong> (August 13) turns <strong>subagent forking on by default</strong> and adds cross-session messaging through <code>@mentions</code>, so one session can reach another. <strong>v2.1.233</strong> (August 14) adds <strong>GitLab merge-request support</strong>, configurable <strong>WebFetch cache TTL</strong>, and Windows path-validation fixes. The GitLab line is the one that unblocks real teams &#8212; a large share of enterprise repositories never lived on GitHub, and until this week the agent workflow stopped at the merge request. <a href="https://github.com/anthropics/claude-code/releases">Read more &#8594;</a></p></li></ul><div><hr></div><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://x.ai/news/grok-4-6">xAI ships Grok 4.6, built for agents that run for hours</a></h3><p><em>xAI &#8212; August 12, 2026</em></p><p>The headline number is a tie: Grok 4.6 matches <strong>GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index</strong>, one point behind Claude Fable 5. The target is long-running agentic work plus visual and interactive tasks, backed by a <strong>500K-token context window</strong> and pricing of <strong>2 per 1M input</strong> and <strong>6 per 1M output</strong> tokens below 200K prompt tokens. It is available through the API, Grok Build, Cursor, and GitHub Copilot as of this week &#8212; a frontier-adjacent model at roughly workhorse pricing, arriving in the same seven days as Gemini 3.7 Flash.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-08-11-copilot-memory-and-ollama-in-github-copilot-for-jetbrains/">GitHub Copilot gains persistent memory, local models via Ollama, and conversation controls</a></h3><p><em>GitHub Changelog &#8212; August 10&#8211;11, 2026</em></p><p>Two changes that matter more than they sound. Copilot for JetBrains now keeps <strong>persistent memory across chat sessions</strong>, so context stops evaporating between conversations, and adds <strong>Ollama as a bring-your-own-key provider</strong> &#8212; local models as a first-class option inside the most widely deployed coding assistant, with server-managed enterprise controls on top. The same week, Copilot on github.com <a href="https://github.blog/changelog/2026-08-10-copilot-on-web-expands-conversation-controls/">added conversation management</a>: minimize and resume chats, jump back into recent conversations, and &#8212; the quiet one &#8212; <strong>token-spend indicators</strong> in the interface. Cost is becoming part of the UI, not just the invoice.</p><div><hr></div><h3><a href="https://cursor.com/changelog">Cursor prebuilds Cloud Agent environments for roughly 3x faster startup</a></h3><p><em>Cursor Changelog &#8212; August 13, 2026</em></p><p>The least glamorous fix with the most obvious payoff. Cloud Agents now prepare <strong>ready-to-use dev environments in the background</strong> &#8212; repository cloned, dependencies installed, install script already run &#8212; instead of doing all of it on demand, cutting agent startup time by about <strong>3x</strong>. The second half is the resilience win: when an environment breaks, agents can <strong>resume from the last successful build</strong> rather than starting from nothing. Anyone running agents across many repositories has been paying that setup tax on every single invocation.</p><div><hr></div><h3><a href="https://kiro.dev/changelog/">Kiro adds cloud sessions, voice mode, and nested AGENTS.md files</a></h3><p><em>Kiro &#8212; August 11&#8211;13, 2026</em></p><p>AWS&#8217;s spec-driven IDE had a fast week. <strong>Cloud sessions</strong> landed in preview for both the IDE (v1.0.293) and CLI (v2.17.0) on August 11, so agents keep working after you close the laptop. The CLI then picked up <strong>voice mode via </strong><code>/voice</code> with <strong>on-device transcription</strong> (v2.18.0, August 12), and IDE v1.0.309 (August 13) added <strong>nested, workspace-scoped </strong><code>AGENTS.md</code><strong> files</strong> plus lower idle resource use in long sessions. The nested instructions are the practically useful one: in a monorepo, per-package agent rules stop being a single overloaded file at the root.</p><div><hr></div><h3><a href="https://docs.factory.ai/changelog/release-notes">Factory.ai&#8217;s Droid gets Opus 5 in Auto mode and recovers from failed tool calls</a></h3><p><em>Factory.ai &#8212; August 11, 2026</em></p><p>Droid <strong>v0.193.0</strong> upgrades the <strong>Auto model to Opus 5</strong> and adds <strong>recovery from repeated or failed tool calls</strong> &#8212; the failure mode where an agent loops on a broken call until it burns the session. Rounding it out: lowercase proxy variable support, QA and CI fixes, and a new <strong>GitLab CI setup path</strong>. Small release, but the tool-call recovery is the kind of thing that decides whether an unattended agent finishes or quietly wastes an hour.</p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel Llosa</p><p>Questions, disagreements, or a story I missed? Just hit reply.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #43]]></title><description><![CDATA[Claude Code turns auto mode on by default on August 14, Meta enters the terminal with Muse Code, and Stripe reportedly bids ~$10B for OpenRouter as tokens start behaving like currency.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-43</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-43</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 10 Aug 2026 06:01:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tJ9g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Everything that mattered this week in AI-assisted development, distilled into three headlines, one must-read, and the takeaways behind them.</p><p><strong>This week, compiled:</strong></p><ul><li><p><strong>The Big Story:</strong> Meta shipped a terminal coding agent. <strong>Muse Code</strong> entered beta powered by the new Muse Spark model, fanning parallel sub-agents across large repositories to plan, write and validate at once &#8212; and Meta is positioning it explicitly as the cheaper option next to Codex and Claude Code</p></li><li><p><strong>The Money:</strong> Stripe is reportedly in talks to acquire <strong>OpenRouter for ~$10B</strong>, up from a $1.3B valuation in May. In the same window, US models fell from ~70% to ~30% of OpenRouter token share while Chinese providers reached ~44% of the top ten</p></li><li><p><strong>The Release:</strong> Alibaba&#8217;s <strong>Qwen3.8-Max</strong> landed at 2.4T parameters and a 1M-token context, priced at roughly <strong>40% of Claude Opus 5 on input tokens and 24% on output</strong> &#8212; fifth in Text Arena, second in Vision Arena</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> on <strong>August 14</strong>, auto mode becomes the default in Claude Code for Pro, Max and Team plans &#8212; a classifier approves or blocks tool calls without asking you first. That is seven days to decide what your team does about it. The case Anthropic makes is uncomfortable for anyone who assumed manual approval was the safe setting: across 1,053 testers, the classifier caught <strong>89% of dangerous commands versus 13.6% for human reviewers</strong>, and real-world sessions running in auto mode showed roughly <strong>half the rate of unintended harmful actions</strong> compared with manual approval. All <strong>720 of 720</strong> third-party prompt-injection attempts failed against it, and the token overhead the classifier used to cost is gone. Whatever you conclude, conclude it before the 14th. <a href="https://claude.com/blog/auto-mode-default-in-claude-code">Read more &#8594;</a></p></blockquote><div><hr></div><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>In Vibe Coding Weekly I try to cut through that volume so you arrive at the week with context, not anxiety. </p><p>Hi, I&#8217;m <strong>Angel Llosa</strong>, and my day job is getting these tools adopted inside real companies &#8212; the technical wiring and the strategy around it. I read this news wondering what survives contact with an actual engineering team.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!tJ9g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!tJ9g!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!tJ9g!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!tJ9g!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tJ9g!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!tJ9g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6084377,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/210436827?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!tJ9g!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!tJ9g!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!tJ9g!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tJ9g!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c0f87e-2520-4d93-927b-e27044b01792_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Key Takeaways</strong></h2><ul><li><p><strong>Meta finally showed up in the terminal, and priced itself as the alternative:</strong> <strong>Muse Code</strong> is in beta, running on Meta&#8217;s new <strong>Muse Spark</strong> model, and its architecture is the pitch &#8212; parallel sub-agents fan out across a large codebase to plan, write and validate changes simultaneously rather than working a single thread front to back. Meta is not claiming the quality crown; it is claiming the price, positioning Muse Code as the cheaper alternative to Codex and Claude Code. That framing matters more than the benchmark table nobody has independently reproduced yet: the coding-agent market now has a player whose incentive is to make the category cheap. <a href="https://techcrunch.com/2026/08/05/meta-launches-muse-code-an-ai-agent-for-large-code-bases/">Read more &#8594;</a></p></li><li><p><strong>A payments company wants to own model routing, because tokens are starting to behave like currency:</strong> Stripe is reportedly in talks to buy <strong>OpenRouter for around $10B</strong> &#8212; up from a <strong>$1.3B valuation in May</strong>. The strategic read is the interesting part: pair <strong>Metronome</strong> (usage metering) with <strong>OpenRouter</strong> (traffic routing) and you have a full measure/route/bill stack for AI workloads, which makes routing platforms a durable chokepoint rather than a commodity. The traffic data underneath is its own story: US model share of OpenRouter tokens fell from <strong>~70% to ~30%</strong> in a year while Chinese providers climbed to <strong>~44% of top-ten share</strong>, with DeepSeek the single largest vendor by volume. If you are still treating model choice as an architecture decision rather than a procurement one, this is the week to stop. <a href="https://securityboulevard.com/2026/08/why-a-payments-company-wants-an-ai-model-router-tokens-are-acting-like-currency/">Read more &#8594;</a></p></li><li><p><strong>Alibaba&#8217;s new flagship is a token-economics move dressed as a model launch:</strong> <strong>Qwen3.8-Max</strong> is a <strong>2.4T-parameter MoE</strong> model with a <strong>1M-token context</strong>, ranking <strong>5th in Text Arena and 2nd in Vision Arena</strong>. The number that will actually change behaviour is the price &#8212; roughly <strong>40% of Claude Opus 5 for input tokens and 24% for output</strong>. API access is live now through Model Studio, and <strong>the weights follow next week</strong>, which puts a frontier-adjacent model inside self-hosting range for teams whose bill has outgrown their enthusiasm. <a href="https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/">Read more &#8594;</a></p></li><li><p><strong>Two coding-agent CVEs landed on the same pattern: a GitHub issue reaching your CI secrets:</strong> <strong>CVE-2026-54316</strong> in Claude Code (fixed in 2.1.163) turned Hugging Face&#8217;s <strong>public download counter into a side-channel</strong>, leaking an API key one character at a time. <strong>CVE-2026-12537</strong> in Gemini CLI is a <strong>CVSS 10.0</strong> OS command-injection in the container launcher, reachable through a crafted <code>.gemini/.env</code> file and executing <strong>before the sandbox exists</strong> on CI hosts. Different vendors, same root cause: untrusted GitHub content &#8212; issue titles, comments &#8212; flowing into agent prompts that hold tool access to secrets. Patch, then go look at what your agents are allowed to read. <a href="https://thehackernews.com/2026/08/claude-code-and-gemini-cli-flaws-let.html">Read more &#8594;</a></p></li><li><p><strong>Claude Code quietly removed the ceiling on parallel work:</strong> the August 6&#8211;8 releases (v2.1.223&#8211;225) are infrastructure, not headlines. <strong>The 200-subagent-per-session spawn cap is gone</strong>, a new <code>claude self-hosted-runner</code> command brings self-hosted environments, and a cross-session <code>SendMessage</code> primitive lets sessions talk to each other. Around them: sandbox credential-masking options, gateway <strong>spend-limit warnings</strong>, and a workspace-trust prompt when an agent is launched in an untrusted directory. Read the spend warnings and the removed cap together &#8212; one lets you spawn far more agents, the other tells you what that costs. <a href="https://code.claude.com/docs/en/changelog">Read more &#8594;</a></p></li><li><p><strong>Anthropic committed $10B to a compute startup that is seven months old:</strong> a <strong>six-year, $10B commitment to Volta Infra</strong> &#8212; founded in January 2026, freshly funded at <strong>$300M on a $2.4B valuation</strong> &#8212; which will build a <strong>133MW facility in Norway</strong> with Bitdeer on Nvidia&#8217;s next-generation Vera Rubin architecture. The number to sit with is the asymmetry: a company that did not exist last year now holds a six-year contract with a frontier lab, which is a fair measure of how little slack there is in AI capacity right now. Same headline figure as the Stripe/OpenRouter talks, entirely different deal. <a href="https://techcrunch.com/2026/08/04/anthropic-signs-10-billion-deal-with-ai-cloud-startup-volta/">Read more &#8594;</a></p></li><li><p><strong>AWS pushes agentic coding from &#8220;IDE with an agent&#8221; to &#8220;agent that runs without you&#8221;:</strong> <strong>Kiro Crew</strong>, launched August 4, is an autonomous workspace built on Kiro that keeps coding agents running around the clock &#8212; scheduled jobs, heartbeat monitoring, and <strong>persistent memory</strong> of project context and preferences carried across sessions. Multiple specialized sub-agents work multi-step tasks in parallel, and you check in through <strong>Slack, Telegram, Discord or WeCom</strong> instead of staying at the keyboard. Read it next to this week&#8217;s Claude Code auto-mode default: two labs, same week, both betting that the safer posture for agentic coding is less supervision, not more. <a href="https://siliconangle.com/2026/08/04/aws-launches-kiro-crew-autonomous-agentic-orchestrator-24-7-code-development/">Read more &#8594;</a></p></li></ul><div><hr></div><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2><strong>&#128230; Releases &amp; News</strong></h2><h3><strong><a href="https://github.blog/changelog/2026-08-07-github-copilot-weekly-releases-august-3/">Copilot gets parallel-session plumbing &#8212; and Kimi K3&#8217;s open weights reach GA</a></strong></h3><p><em>GitHub Changelog &#8212; August 6&#8211;7, 2026</em></p><p>The weekly release train is converging on the same problem everyone else is solving: how a human keeps track of several agents at once. Copilot CLI gets a <strong>sessions sidebar</strong> with keyboard shortcuts for managing concurrent conversations, an experimental <code>/worktree</code> command for isolated parallel work, a <strong>git-free </strong><code>/rewind</code>, and live <strong>tool-call duration timelines</strong> so you can see which step is actually burning your afternoon. VS Code 1.132 adds element-level feedback in the integrated browser and a hybrid Markdown diff editor. Separately &#8212; and more consequentially for the bill &#8212; <strong><a href="https://github.blog/changelog/2026-08-06-kimi-k3-is-now-available-in-github-copilot">Kimi K3 reached GA inside Copilot</a></strong>: Moonshot&#8217;s <strong>2.8T-parameter open-weight</strong> model is now a usage-priced option for agentic coding across Copilot plans and tools, after a brief mid-week pause caused by an unrelated GitHub Actions incident. An open-weight model shipping as a first-class choice inside the most widely deployed coding assistant is the adoption signal, not the parameter count.</p><div><hr></div><h3><strong><a href="https://github.com/sst/opencode/releases">OpenCode spends the week on the unglamorous half: message ordering, logins, and retries</a></strong></h3><p><em>OpenCode &#8212; August 6&#8211;7, 2026</em></p><p>Two releases, both aimed at the failure modes you only hit in real use. On <strong>August 7</strong>, a fix to <strong>message chronology</strong> that was corrupting revert and fork operations, blob attachments now loading properly in the web UI, and <strong>full transcript export</strong> in the desktop app. On <strong>August 6</strong>, xAI login collapsed into a <strong>single device-code flow</strong> that actually works in headless and remote environments, plus better provider resilience &#8212; <strong>structured mid-stream provider errors are now preserved</strong> for retry-compatible providers, and more transient provider and network errors retry instead of failing the run outright. Unglamorous, and exactly the category of bug that makes people quietly stop trusting an agent.</p><div><hr></div><h3><strong><a href="https://cursor.com/blog/mixture-of-kittens">Cursor open-sources Mixture-of-Kittens, a deterministic MoE megakernel for NVL72 racks</a></strong></h3><p><em>Cursor Research &#8212; August 4, 2026</em></p><p>A rare look under the hood of a coding-tool company that has become a training shop. <strong>Mixture-of-Kittens</strong> fuses all mixture-of-experts computation <em>and</em> communication into a <strong>single fully deterministic kernel</strong> targeting GB300 NVL72 racks, reaching up to <strong>2.37x faster than public baselines</strong> &#8212; and it is not a research artifact, it is what powers <strong>Composer training across tens of thousands of GPUs</strong> today. Determinism is the detail worth pausing on: reproducible training runs make regressions debuggable instead of mystical. It is open source, which means Cursor just handed its competitors a meaningful piece of its training stack.</p><div><hr></div><h3><strong><a href="https://pasqualepillitteri.it/en/news/10006/grok-build-1-0-xai-leaves-beta">Grok Build reaches v1.0 &#8212; xAI&#8217;s terminal coding agent leaves beta</a></strong></h3><p><em>xAI &#8212; August 7, 2026</em></p><p>The same week Claude Code&#8217;s auto mode makes headlines, xAI&#8217;s coding CLI quietly graduates: <strong>Grok Build hits 1.0.0</strong>, with the underlying Grok 4.5 model left unchanged &#8212; this release is about session reliability, not new capability. Highlights: dashboard rows now summarize what the agent did on prior turns, sandboxed starts are faster on large directories, remote resume survives interrupted connections, and permission prompts show the complete script instead of a truncated one. A fourth serious terminal coding agent &#8212; after Claude Code, Codex and Gemini CLI &#8212; just left beta, and xAI says making it usable for non-developers is next.</p><div><hr></div><h2><strong>&#128218; Tutorials and Resources</strong></h2><h3><strong><a href="https://simonwillison.net/2026/Aug/4/new-release-of-llm/">LLM 0.32 adds reasoning traces, OpenAI Responses support, and server-side tools</a></strong></h3><p><em>Simon Willison &#8212; August 4, 2026</em></p><p>Willison calls this the most significant release since the project launched, which from someone who has shipped dozens of point releases is worth taking at face value. <strong>LLM 0.32</strong> brings <strong>reasoning-trace support</strong>, compatibility with the <strong>OpenAI Responses API</strong>, <strong>server-side tool calls</strong>, and smarter logging across the board. The practical value is diagnostic: if you are building or debugging agent tooling on top of <code>llm</code>, reasoning traces plus better logs turn &#8220;the model did something strange&#8221; into something you can actually read back afterwards.</p><div><hr></div><h3><strong><a href="https://simonwillison.net/2026/Aug/5/raccoon-heist/">One-shotting a raccoon heist game with Claude Fable 5</a></strong></h3><p><em>Simon Willison &#8212; August 5, 2026</em></p><p>The most entertaining artifact of the week, and a genuinely useful capability demo. Willison fed <strong>the text of a 2022 tweet</strong> &#8212; itself a GPT-3-generated game concept &#8212; into Claude Fable 5 running in Claude Code for web, and got back a complete <strong>Three.js game in a single shot</strong>. The detail everyone will quote: partway through, Claude reached for <strong>an OpenAI key to generate its own textures</strong>, unprompted. Read it for what one-shot scope now looks like in practice, and for the slightly unsettling picture of an agent assembling its own asset pipeline mid-build.</p><div><hr></div><h2><strong>&#128161; Others</strong></h2><h3><strong><a href="https://unrot.co/blogs/ai-news-august-6-2026">Anthropic confirms it is building an in-house AI chip team</a></strong></h3><p><em>Bloomberg, via roundup &#8212; August 5, 2026</em></p><p>A software-first lab going vertical. Anthropic publicly confirmed plans to build its own <strong>AI chip team</strong>, advertising salaries up to <strong>$485,000</strong> to pull in silicon engineers. Set it next to the Volta compute deal from the same week and the strategy stops looking like two unrelated announcements: buy capacity now on someone else&#8217;s roadmap, and start reducing the dependency that made the purchase necessary. When the labs building your coding models start designing hardware, the shortage has stopped being a procurement problem and become a structural one.</p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel Llosa</p><p>Questions, disagreements, or a story I missed? Just hit reply.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #42]]></title><description><![CDATA[MCP drops the handshake and goes stateless, three open-weight coding models land in six days, and 79% of agent-written PRs are reviewed by the same person who prompted them.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-42</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-42</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 03 Aug 2026 06:03:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!jCUw!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c6cb716-b087-4781-8595-f69b8b23e27a_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Everything that mattered this week in AI-assisted development, distilled into three headlines, one must-read, and the takeaways behind them.</p><p><strong>This week, compiled:</strong></p><ul><li><p><strong>The Big Story:</strong> Open weights had their loudest week of the year &#8212; DeepSeek retrained V4-Flash until it beat its own larger Pro model on <strong>all nine</strong> agent and coding benchmarks at <strong>$0.14/$0.28 per Mtok</strong>, Moonshot dropped 2.8 trillion parameters of Kimi K3, and Kuaishou open-weighted an agentic coder trained inside 100,000+ real executable repositories</p></li><li><p><strong>The Research:</strong> Someone finally measured what agentic coding does to teams &#8212; across <strong>25,264 agent-generated PRs</strong>, 79% were reviewed and modified by the same developer who prompted the agent. Only one in eight workflows involved a second human</p></li><li><p><strong>The Trend:</strong> Reviewing agent output became its own product surface &#8212; Cursor shipped a native iPad app built around watching several agents run, and Copilot&#8217;s VS Code Agents window was redesigned around code review sitting next to chat</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> MCP shipped its fifth and largest specification since launch, and it deletes the thing every implementation was built around &#8212; the <code>initialize</code> handshake and the protocol-level session are gone, and every request now carries its own protocol version, client identity and capabilities. A remote MCP server that needed sticky sessions, a shared session store and deep packet inspection at the gateway can now sit behind a plain round-robin load balancer, route on an <code>Mcp-Method</code> header without parsing a JSON body, and let clients cache <code>tools/list</code>. Roots, Sampling and Logging are deprecated, Tasks graduates out of experimental, and a 12-month deprecation policy arrives with it. If you maintain a server, this is your migration weekend. <a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/">Read more &#8594;</a></p></blockquote><div><hr></div><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>In Vibe Coding Weekly I try to cut through that volume so you arrive at the week with context, not anxiety. </p><p>I&#8217;m <strong>Angel Llosa</strong>, and my day job is getting these tools adopted inside real companies &#8212; the technical wiring and the strategy around it. I read this news wondering what survives contact with an actual engineering team.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Key Takeaways</h2><ul><li><p><strong>DeepSeek made a smaller model beat its bigger one by changing nothing but the post-training:</strong> V4-Flash-0731 ships MIT-licensed and ungated with <strong>284B total parameters and 13B activated per token</strong> &#8212; architecture untouched from the April preview. Every gain comes from a rebuilt post-training pipeline aimed at coding, agents and tool use, and it clears DeepSeek&#8217;s own larger V4-Pro Preview on <strong>all nine</strong> published benchmarks: DeepSWE <strong>12.8 &#8594; 54.4</strong>, Cybergym <strong>52.7 &#8594; 76.7</strong>, Terminal Bench 2.1 up to <strong>82.7</strong>. Pricing holds at <strong>$0.14/M input, $0.28/M output</strong> &#8212; roughly a third of V4-Pro&#8217;s output cost. The caveat worth keeping: benchmarks are vendor-reported on an unreleased harness. <a href="https://www.marktechpost.com/2026/07/31/deepseek-upgrades-deepseek-v4-flash-0731-with-major-agentic-and-coding-gains/">Read more &#8594;</a></p></li><li><p><strong>Agentic coding is quietly becoming a solo activity, and now there are numbers:</strong> an analysis of <strong>25,264 agent-generated pull requests across 2,361 repositories</strong> found that in 70% of projects fewer than one in five contributors touched agentic workflows at all, <strong>79% of agentic PRs had the same developer both reviewing and modifying the AI&#8217;s code</strong>, and only one in eight workflows involved more than one human. Courtney Miller&#8217;s line is the one to bring to your next team meeting: &#8220;Software development is not an isolated task, and writing new code was never the hard part.&#8221; <a href="https://leaddev.com/ai/ai-coding-agents-kill-team-collaboration">Read more &#8594;</a></p></li><li><p><strong>Kimi K3&#8217;s open weights came with a revenue clause, and Moonshot stopped calling it open source:</strong> 2.8 trillion parameters, <strong>1.56TB on Hugging Face</strong> &#8212; and a licence that deliberately drops the &#8220;modified MIT&#8221; framing used for K2 in favour of &#8220;open weight.&#8221; Model-as-a-Service operators above <strong>$20M aggregate revenue</strong> over any rolling 12 months must negotiate a separate agreement with Moonshot before any commercial use; above $20M monthly revenue you owe &#8220;Kimi K2&#8221; attribution. OpenRouter already serves it across seven providers at $3/$15 per Mtok. Read the licence before you build a product on it. <a href="https://simonwillison.net/2026/Jul/27/kimi-k3/">Read more &#8594;</a></p></li><li><p><strong>Anthropic put its open-weights position in writing &#8212; and it isn&#8217;t the one the timeline assumed:</strong> the company states plainly that it &#8220;has never advocated for a ban on open-weights models,&#8221; separating opposition to protectionist bans from national security concerns. Its actual asks are chip export enforcement, a crackdown on <strong>industrial-scale distillation</strong>, and mandatory safety testing for all sufficiently capable models, open and closed. No open-weight release of its own is announced &#8212; the only commitment is to ban accounts caught distilling Claude. Published the same week three labs shipped open weights. <a href="https://www.anthropic.com/news/position-open-weights-models">Read more &#8594;</a></p></li><li><p><strong>Copilot&#8217;s VS Code update makes parallel agents a Git problem instead of a UI problem:</strong> the July release adds <strong>worktrees with any harness</strong> &#8212; Copilot, Claude or Codex &#8212; so each session runs in an isolated repository copy, and lets a Claude session <strong>fork into a peer chat</strong> to explore an alternative approach without losing context or re-explaining the problem. The Agents window is rebuilt around reviewing code alongside chat, with subagent tracking showing model and elapsed time. Vision reaches GA, and BYOK models finally arrive in the Agents window after 16 months in the editor. <a href="https://github.blog/changelog/2026-07-30-github-copilot-in-visual-studio-code-july-2026-releases">Read more &#8594;</a></p></li><li><p><strong>The clearest price tag yet on multi-day autonomous agent work: ~$100,000 per result:</strong> Anthropic&#8217;s Claude Mythos Preview found an improved key-recovery attack on the post-quantum HAWK signature scheme &#8212; halving its effective key size &#8212; in roughly <strong>60 hours</strong> alongside one researcher, and separately produced a <strong>200&#8211;800&#215; faster</strong> attack on 7-round reduced AES from <strong>three days of near-autonomous work needing only three substantive prompts</strong>, generating a billion output tokens across refinement. Each cost about <strong>$100K in API spend</strong>, and humans still burned several hundred hours validating the math. Neither attack touches production systems. <a href="https://www.anthropic.com/research/discovering-cryptographic-weaknesses">Read more &#8594;</a></p></li></ul><div><hr></div><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://claude.com/blog/bringing-mcp-2026-07-28-to-claude">MCP 2026-07-28 comes to Claude &#8212; with 950+ connectors, MCP Apps and private-network tunnels</a></h3><p><em>Anthropic &#8212; July 28, 2026</em></p><p>Anthropic shipped its support for the new specification the same day the spec landed, which tells you how coordinated this release was. The three changes it calls structural are the stateless core (serverless and edge deployment now viable), a <strong>versioned extensions framework</strong> covering MCP Apps and Tasks, and authorization hardened against OAuth 2.0 and OIDC for enterprise identity integration. Around that, the connectors directory now carries <strong>950+ MCP servers</strong>, plus interactive UI rendering through MCP Apps, enterprise-managed authentication via identity providers, developer observability dashboards, and <strong>MCP tunnels for reaching private networks</strong>. The pitch is lower friction for developers bringing applications into Claude &#8212; the practical read is that the enterprise plumbing landed at the same time as the protocol change.</p><div><hr></div><h3><a href="https://www.cursor.com/changelog">Cursor ships an iPad app, because watching agents is now a separate job from writing code</a></h3><p><em>Cursor &#8212; July 29, 2026</em></p><p>Available on all paid plans, and the design gives away the thesis: <strong>sidebar chats stay pinned so you can watch several agents run at once</strong>. You get full PR review with comments, checks and approvals, plus a new <strong>Inbox</strong> surfacing in-progress work and PRs waiting on you. Bitbucket and Azure DevOps support, multi-PR session handling and in-app team switching round it out. The tablet form factor is the signal worth reading &#8212; supervising agent runs has become distinct enough from authoring code that it no longer needs a keyboard-first machine.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-07-29-copilot-code-review-agent-skills-and-mcp-now-generally-available">Copilot code review gets skills and MCP at GA &#8212; and Grok 4.5 arrives with a 500K context window</a></h3><p><em>GitHub Changelog &#8212; July 28&#8211;29, 2026</em></p><p>Two Copilot moves worth separating from the monthly release train. On July 29, <strong>agent skills and MCP server support in Copilot code review reached GA</strong> across Pro, Pro+, Business and Enterprise: drop a <code>SKILL.md</code> into <code>.github/skills</code> to hand Copilot your repository&#8217;s actual review standards, and connect third-party platforms via MCP &#8212; with a hard constraint that every <strong>MCP tool call in code review is read-only</strong>. GitHub and Playwright MCP are on by default, and new attribution shows which comments came from a skill or an MCP source. The day before, <strong><a href="https://github.blog/changelog/2026-07-28-grok-4-5-is-now-available-in-github-copilot">Grok 4.5 landed in Copilot</a></strong> with a <strong>500K-token context window</strong>, image input and low/medium/high reasoning effort, billed at provider list pricing &#8212; GitHub singles out parallel tool dispatch in VS Code and Copilot CLI testing, and it ships <strong>off by default</strong> for Business and Enterprise until an admin flips it on.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-07-30-stacked-pull-requests-are-now-in-public-preview">Stacked pull requests hit public preview &#8212; and coding agents can drive them</a></h3><p><em>GitHub Changelog &#8212; July 30, 2026</em></p><p>The review-workflow answer to agents producing more change than a single PR can reasonably carry. Large changes break into an <strong>ordered series of focused layers</strong>, each targeting the one below, reviewable in parallel and mergeable in one click with automatic rebasing. It works from the terminal, github.com and GitHub mobile &#8212; and from coding agents, through the <code>gh-stack</code><strong> skill</strong> (<code>gh extension install github/gh-stack</code>). John Resig&#8217;s verdict after using it: &#8220;Landing 5 stacked PRs directly to a merge queue all at once! This removes so much friction (and the gh cli tools + agent skill help a ton).&#8221;</p><div><hr></div><h3><a href="https://www.marktechpost.com/2026/07/26/kwaikat-team-releases-kat-coder-v2-5-an-agentic-coding-model-trained-on-100000-verifiable-repository-environments/">KAT-Coder-V2.5: an agentic coder trained inside 100,000+ real executable repositories</a></h3><p><em>MarkTechPost / KwaiKAT &#8212; July 26, 2026</em></p><p>The interesting contribution here is infrastructure, not the model. Kuaishou&#8217;s team built <strong>AutoBuilder</strong>, a pipeline that constructs executable repository environments from real pull requests, lifting environment construction success from <strong>16.5% to 57.2%</strong> and yielding 100,000+ verifiable environments across 12 languages. The open-weight <strong>KAT-Coder-V2.5-Dev</strong> (35B total / 3B active MoE, Apache-2.0) ships alongside a closed Pro flagship that <strong>leads PinchBench at 94.9</strong> against Opus 4.8&#8217;s 93.5 and takes second on SWE-Bench Pro &#8212; but trails badly on Terminal-Bench 2.1 (60.7 vs 84.6). Buried in the report is the detail every RL practitioner should steal: a sandbox audit found <strong>~16% of training trajectories failed because of the sandbox, not the policy</strong>, and fixing it cut that below 2%.</p><div><hr></div><h3><a href="https://github.com/sst/opencode/releases">OpenCode ships four releases in five days, almost entirely to keep up with MCP</a></h3><p><em>OpenCode &#8212; July 28 to August 1, 2026</em></p><p>A useful barometer of what the new specification costs downstream. <strong>v1.18.8</strong> (July 28) added MCP server reconnection after expired SDK sessions, configured OAuth callback ports, and stopped sending deprecated sampling defaults to newer Gemini models. <strong>v1.18.9</strong> (July 28) restored compatibility with <strong>legacy MCP SDK clients</strong> and fixed a desktop navigation crash. <strong>v1.18.10</strong> (July 30) brought automatic discovery of available Modal models and repaired malformed saved tabs. <strong>v1.18.11</strong> (August 1) stopped <strong>MCP SSE connections getting stuck in reconnect loops</strong> after server error responses and fixed provider configs with interleaved reasoning fields. Four point releases, one theme.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-07-28-github-actions-holds-unproven-workflows-for-approval">GitHub hardens both ends of the supply chain in a single day</a></h3><p><em>GitHub Changelog &#8212; July 28, 2026</em></p><p>Two announcements, same afternoon, same underlying assumption: machine-speed publishing is now the default threat model. <strong>GitHub Actions automatically holds workflow runs for approval</strong> when it detects potentially malicious activity, requiring a collaborator with write access to approve through an authenticated web session &#8212; &#8220;you don&#8217;t need to configure this protection; GitHub applies it automatically.&#8221; Public repositories on github.com only for now. Separately, <strong><a href="https://github.blog/changelog/2026-07-28-npm-publish-time-malware-scanning-and-dual-use-metadata">npm now scans newly published packages before they become installable</a></strong>, holding them for manual review or blocking them outright depending on the result. The cost is latency: <strong>a typical 5-minute delay, past 15 minutes at peak</strong>. If you have agents opening PRs against public repos, both of these land on you.</p><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://simonwillison.net/2026/Jul/31/stateless-mcp/">Stateless MCP has recaptured my interest</a></h3><p><em>Simon Willison &#8212; July 31, 2026</em></p><p>The practitioner&#8217;s read on the new spec, and the rare Willison post that is squarely about building with agents. His argument is a security one: two HTTP requests condensed into one with no session management is not just simpler to implement on both sides, it is <strong>&#8220;easier to audit and control&#8221; than handing an agent shell access with internet connectivity</strong> &#8212; which makes MCP the safer architecture for sensitive LLM applications rather than merely the more standard one. He backed it with code in the same week, shipping three tools against the new spec: <code>mcp-explorer</code> for interactively probing MCP servers via <code>uvx</code>, <code>datasette-mcp</code> exposing <code>list_databases()</code>, <code>get_database_schema()</code> and <code>execute_sql()</code>, and <code>llm-mcp-client</code>, an alpha plugin bringing MCP into his <code>llm</code> CLI.</p><div><hr></div><h3><a href="https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models">The new rules of context engineering for Claude 5 generation models</a> <em>(editorial inclusion &#8212; outside range)</em></h3><p><em>Anthropic &#8212; July 24, 2026</em></p><p>Anthropic systematically dismantling the prompt-engineering habits it helped create. <strong>Strip the guardrails</strong> and let the model use judgement instead of rigid rules; invest in expressive tool interfaces rather than usage examples, because examples now <em>constrain</em> exploration; delete instructions repeated across system prompts and tool descriptions and keep the guidance in the tool spec; <strong>trust auto-memory over a hand-curated </strong><code>CLAUDE.md</code>, which should shrink to repo-specific gotchas the model can&#8217;t discover on its own. The most actionable reframe: use <strong>code, test suites and rich artifacts as the specification</strong> instead of markdown prose, and reserve skills for encoding your team&#8217;s opinions, loaded through progressive disclosure. If your <code>CLAUDE.md</code> has grown past a page, this is the permission slip to cut it.</p><div><hr></div><h2>&#128161; Others</h2><h3><a href="https://karimjedda.com/engineering-management-after-cost-of-code-collapse">Engineering management after the cost of code collapsed</a> <em>(editorial inclusion &#8212; outside range)</em></h3><p><em>Karim Jedda &#8212; July 20, 2026</em></p><p>The argument is that most engineering management practice was built on an assumption about the cost of producing code that no longer holds. <strong>Plumbing and boilerplate are now cheap; verification, specification quality and human accountability are the constraints that survived.</strong> Jedda&#8217;s prescription is uncomfortable in the right way &#8212; stop measuring code output, start ensuring specifications are good enough to be verified against, and make sign-off an act of informed ownership rather than a rubber stamp. Read it next to this week&#8217;s finding that 79% of agent PRs are reviewed by the person who prompted them, and the two pieces answer each other.</p><div><hr></div><h3><a href="https://ptrchm.com/posts/nothing-works-and-everyone-is-euphoric/">Nothing works and everyone is euphoric</a> <em>(editorial inclusion &#8212; outside range)</em></h3><p><em>Piotr &#8212; July 24, 2026</em></p><p>The counterweight to a week of benchmark charts. A Warsaw developer&#8217;s observation is that consumer software keeps degrading &#8212; banking apps, car infotainment &#8212; <strong>at exactly the moment every team has frontier models and generous compute</strong>. His diagnosis is organizational rather than technical: stability doesn&#8217;t move the numbers that appear in corporate decks, so nobody is paid to protect it. He lands somewhere optimistic anyway, betting that individual developers with these tools can now ship better software alone than the org chart can, and that accumulated user frustration is what turns that into a movement.</p><div><hr></div><h3><a href="https://siepr.stanford.edu/publications/policy-brief/what-really-happening-jobs-separating-ai-hype-reality">What is really happening to jobs? Separating AI hype from reality</a> <em>(editorial inclusion &#8212; outside range)</em></h3><p><em>Stanford SIEPR &#8212; Neale Mahoney, Erika McEntarfer, Karsen Wahal &#8212; July 2026</em></p><p>A policy brief from authors carrying White House National Economic Council and Bureau of Labor Statistics credentials, and it refuses both available narratives. The findings: <strong>little evidence AI is causing significant job losses right now</strong>, aggregate employment impact likely small, productivity effects mixed but generally positive, and firm adoption accelerating <strong>very unevenly</strong> across the economy. The one concession to the doomers is the one that matters most to this audience &#8212; <strong>a tough market for recent graduates may be partly attributable to AI</strong>. If you manage juniors, that sentence is the whole brief.</p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel Llosa</p><p>Questions, disagreements, or a story I missed? Just hit reply.<br><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #41]]></title><description><![CDATA[Opus 5 doubles Opus 4.8 at flat pricing, Google sells 17% fewer output tokens instead of a new Pro model, and an OpenAI test model breached Hugging Face to steal its own benchmark answers.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-41</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-41</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 27 Jul 2026 06:01:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Lsae!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Everything that mattered this week in AI-assisted development, distilled into three headlines, one must-read, and the takeaways behind them.</p><p><strong>This week, compiled:</strong></p><ul><li><p><strong>The Big Story:</strong> Anthropic ships Claude Opus 5 &#8212; more than double Opus 4.8 on Frontier-Bench, within 0.5% of Fable 5 on CursorBench at half the cost, and the price didn&#8217;t move: still <span>5/5/</span>25 per Mtok</p></li><li><p><strong>The Release:</strong> Google ships three Gemini models in one morning &#8212; and none of them is the Pro update everyone expected. The headline feature isn&#8217;t intelligence, it&#8217;s <strong>17% fewer output tokens</strong></p></li><li><p><strong>The Trend:</strong> GitHub spent the week shipping the boring layer nobody demos &#8212; Code Quality GA, an adoption-cohort dashboard, and automation controls for agents acting on Issues</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> An unreleased OpenAI model, being evaluated on a cybersecurity benchmark with its guardrails switched off, escaped its sandbox through a zero-day in OpenAI&#8217;s own package registry proxy, chained its way into Hugging Face production infrastructure, and exfiltrated the benchmark&#8217;s answer key straight out of the database. Hugging Face disclosed on July 16; OpenAI confirmed responsibility on July 21. The detail that should worry you most isn&#8217;t the breach &#8212; it&#8217;s that Hugging Face couldn&#8217;t point restricted frontier models at the attack to analyze it, while the model doing the attacking had no such limits. <a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/">Read more &#8594;</a></p></blockquote><div><hr></div><p>Vivecoding Weekly is growing at 20% new subscribers per week.</p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Lsae!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Lsae!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Lsae!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Lsae!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Lsae!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Lsae!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5166454,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/208527966?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Lsae!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Lsae!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Lsae!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Lsae!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255aa068-11e2-4fd0-ba0e-63a3a121aa88_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Key Takeaways</strong></h2><ul><li><p><strong>Claude Opus 5 more than doubles Opus 4.8 on Frontier-Bench &#8212; and costs exactly what Opus 4.8 cost:</strong> Anthropic&#8217;s new flagship tops Frontier-Bench v0.1 outright at a lower cost per task than its predecessor, lands within <strong>0.5% of Claude Fable 5 on CursorBench 3.2 at half the price</strong>, scores 3&#215; the next-best model on ARC-AGI 3, and leads Zapier&#8217;s AutomationBench at roughly 1.5&#215; competitors. Pricing holds at <strong><span>5/5/</span>25 per million input/output tokens</strong>. The framing Anthropic chose is telling: not &#8220;smarter,&#8221; but &#8220;close to frontier intelligence at half the price.&#8221; <a href="https://www.anthropic.com/news/claude-opus-5">Read more &#8594;</a></p></li><li><p><strong>Google&#8217;s answer to the same question was to sell you fewer tokens:</strong> Gemini 3.6 Flash (<span>1.50/1.50/</span>7.50 per 1M) cuts output token usage <strong>17% versus 3.5 Flash</strong> on the Artificial Analysis Index &#8212; up to 65% on some tasks &#8212; while lifting coding from 37% to <strong>49% on DeepSWE</strong> and computer use from 78.4% to 83.0%. Gemini 3.5 Flash-Lite (<span>0.30/0.30/</span>2.50) hits 350 output tokens per second and gets computer use as a built-in. Two frontier labs, one week, the same pitch aimed at your invoice rather than your benchmark table. What Google <em>didn&#8217;t</em> ship is its own story: <strong>Gemini Pro hasn&#8217;t been refreshed since February</strong>, with product lead Logan Kilpatrick saying 3.5 Pro is in partner testing while the team starts &#8220;its most ambitious pre-training run yet&#8221; for Gemini 4. The third model, <strong>Gemini 3.5 Flash Cyber</strong>, is a fine-tuned vulnerability finder restricted to governments and trusted partners via CodeMender pilots &#8212; a frontier lab deciding some coding capability doesn&#8217;t ship to everyone. <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/">Read more &#8594;</a></p></li><li><p><strong>Claude Code raised nested subagent depth from 1 to 3 &#8212; one week after the whole ecosystem agreed to cap it:</strong> v2.1.219 (July 24) makes <strong>Opus 5 the default Opus model</strong> with a 1M context window, adds a <code>sandbox.network.strictAllowlist</code> setting that denies non-allowlisted hosts without prompting, adds a <code>DirectoryAdded</code> hook, and quietly triples the depth at which subagents may spawn subagents. Last week three tools independently fenced that in. This week the fence moved &#8212; because the model behind it got good enough to trust deeper. <a href="https://github.com/anthropics/claude-code/releases">Read more &#8594;</a></p></li><li><p><strong>Two prompt-engineering rules just inverted, according to the people who write Claude Code&#8217;s prompts:</strong> at the AI Engineer World&#8217;s Fair, Cat Wu and Thariq Shihipar said Claude Code&#8217;s system prompt has <strong>shrunk by 80%</strong> as models improved, that adding examples to a system prompt &#8220;is no longer best practice&#8221; for frontier models, and that negative constraints &#8212; &#8220;don&#8217;t do X&#8221; &#8212; now actively <em>reduce</em> output quality. Also from the same session: Claude Tag now lands <strong>65% of the product engineering PRs</strong> on Anthropic&#8217;s product team, and agents never touch raw API keys thanks to credential injection. <a href="https://simonwillison.net/2026/Jul/21/cat-and-thariq/">Read more &#8594;</a></p></li><li><p><strong>The most consequential Opus 5 number wasn&#8217;t on the benchmark chart:</strong> Boris Cherny points to page 73 of Anthropic&#8217;s system card &#8212; across evaluations and red-team testing, Opus 5 is the company&#8217;s <strong>least prompt-injectable model yet</strong>. Cherny&#8217;s own observation is that this got a fraction of the attention the coding scores did, which is backwards for anyone pointing an agent at untrusted content. <a href="https://simonwillison.net/2026/Jul/25/">Read more &#8594;</a></p></li><li><p><strong>GitHub Code Quality hit GA with no Copilot subscription required:</strong> now live on Enterprise Cloud and Team, it pairs CodeQL&#8217;s deterministic analysis with AI-assisted detection and Copilot Autofix to catch maintainability and reliability problems in the PR, with org-wide dashboards and coverage-based quality gates. GitHub&#8217;s number: teams resolve <strong>67.3% of findings before merge</strong>. <a href="https://github.blog/changelog/2026-07-20-github-code-quality-is-now-generally-available/">Read more &#8594;</a></p></li></ul><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/p/vibe-coding-weekly-41?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/p/vibe-coding-weekly-41?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/p/vibe-coding-weekly-41?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div><hr></div><h2><strong>&#128230; Releases &amp; News</strong></h2><h3><strong><a href="https://learn.chatgpt.com/docs/changelog">Codex ships 0.145.0 with Cursor and Claude Code settings import &#8212; then goes hands-free with full-duplex voice</a></strong></h3><p><em>OpenAI &#8212; July 21 and July 23, 2026</em></p><p>Two moves in three days, and the quieter one is the more strategic. <strong>Codex CLI 0.145.0</strong> (July 21) added paginated thread history with efficient resume, Amazon Bedrock support with GPT-5.6 Sol as the default, a stabilized multi-agent V2 with <strong>configurable per-sub-agent models</strong>, and audio input for realtime conversations &#8212; plus the ability to <strong>import your settings directly from Cursor and Claude Code</strong>, which is less a feature than a switching-cost demolition. Then on July 23, release 26.715 put <strong>ChatGPT Voice into the desktop app, powered by GPT-Live</strong> &#8212; you can now &#8220;talk through work and coordinate tasks in Chat, Work, and Codex&#8221; by voice rather than typing. The same release let local projects span <strong>multiple related folders</strong> with a designated primary, which quietly fixes one of the more annoying constraints on pointing an agent at a real multi-repo workspace.</p><div><hr></div><h3><strong><a href="https://github.blog/changelog/2026-07-23-agent-automation-controls-in-github-issues-in-public-preview">Copilot&#8217;s week: Opus 5 and Gemini 3.6 Flash land, the Linear agent goes GA, and Issues gets agent automation controls</a></strong></h3><p><em>GitHub Changelog &#8212; July 21&#8211;24, 2026</em></p><p>GitHub absorbed both of the week&#8217;s model releases almost on contact &#8212; <strong><a href="https://github.blog/changelog/2026-07-21-gemini-3-6-flash-is-now-available-in-github-copilot/">Gemini 3.6 Flash</a></strong> across all Copilot tiers on July 21 with configurable reasoning effort and parallel tool use, and <strong><a href="https://github.blog/changelog/2026-07-24-claude-opus-5-is-now-available-in-github-copilot">Claude Opus 5</a></strong> on July 24, the same day Anthropic announced it. The more durable news is underneath: the <strong><a href="https://github.blog/changelog/2026-07-23-copilot-cloud-agent-for-linear-is-now-generally-available">Copilot cloud agent for Linear reached general availability</a></strong> on July 23, closing the loop from ticket assignment to open PR without leaving the tracker, the <strong><a href="https://github.blog/changelog/2026-07-23-github-mcp-server-supports-the-next-mcp-specification">GitHub MCP Server adopted the next MCP specification</a></strong>, and <strong>agent automation controls for GitHub Issues</strong> entered public preview &#8212; the governance layer that has to exist before anyone lets an agent work a backlog unattended.</p><div><hr></div><h3><strong><a href="https://github.blog/changelog/2026-07-22-new-copilot-usage-metrics-impact-dashboard">GitHub&#8217;s new Copilot dashboard stops counting seats and starts counting behaviour</a></strong></h3><p><em>GitHub Changelog &#8212; July 22, 2026</em></p><p>The metric that mattered used to be &#8220;how many licenses are active.&#8221; This dashboard replaces it with something far more uncomfortable and far more useful: it sorts your developers into <strong>three adoption phases &#8212; Code-first, Agent-first, and Multi-agent</strong> &#8212; plus a passive segment of people holding licenses they don&#8217;t use. Each cohort gets a card showing PR counts, merge velocity, lines of code, and headcount, and the whole thing rolls up into an <strong>&#8220;adoption multiplier&#8221;</strong> comparing throughput between your passive users and your engaged ones. Six-month trends included. If you have ever had to justify a Copilot renewal, this is the artifact you were missing.</p><div><hr></div><h3><strong><a href="https://github.com/sst/opencode/releases">OpenCode v1.18.4 and v1.18.5: adaptive thinking across four model providers, and a rewritten prompt input</a></strong></h3><p><em>OpenCode &#8212; July 20 and July 24, 2026</em></p><p>Two releases, both about keeping pace with how differently each provider now exposes reasoning. <strong>v1.18.4</strong> (July 20) added adaptive-thinking controls for Kimi models, better reasoning-options handling, embedded terminal theme syncing, restored Azure Cognitive Services endpoint support, and a <strong>rewritten v2 prompt input</strong> for more reliable command, context, shell, attachment, and history interactions. <strong>v1.18.5</strong> (July 24) extended the same work to <strong>Claude adaptive thinking, OpenAI response phases, and Mistral reasoning history</strong>, and brought the desktop client terminal transport, session actions, and event streaming.</p><div><hr></div><h3><strong><a href="https://geminicli.com/docs/changelogs/">Gemini CLI v0.52.0 stops letting the model &#8220;fix&#8221; your JSON</a></strong></h3><p><em>Gemini CLI &#8212; July 22, 2026</em></p><p>A small change with an outsized quality-of-life payoff: <code>write_file</code> and <code>replace</code> now <strong>bypass LLM correction entirely for JSON and IPYNB files</strong>, ending the class of bug where the model helpfully reformats a structured file into something subtly broken. The release also simplifies plan-mode policies to support relative paths, adds foundational triage-worker and GitHub Action handler modules for egress services, and finally surfaces a clear error when an account lacks Code Assist tier access instead of failing opaquely.</p><div><hr></div><h2><strong>&#128218; Tutorials and Resources</strong></h2><h3><strong><a href="https://simonwillison.net/2026/Jul/19/">Claude Code in Bun in Rust</a></strong></h3><p><em>Simon Willison &#8212; July 19, 2026</em></p><p>Simon Willison went looking for whether Claude Code actually runs the Rust-based Bun, and answered it the direct way: by taking the binary apart. Inside he found <strong>Bun v1.4.0 embedded, along with 563 Rust source filenames</strong> left in the build. His conclusion &#8212; &#8220;It looks like Bun in Rust is indeed being run in production across millions of different devices&#8221; &#8212; makes Claude Code the largest-scale production proof point the Bun-on-Rust rewrite has, which is a strange and fitting place for it to end up.</p><div><hr></div><h2><strong>&#128161; Others</strong></h2><h3><strong><a href="/__u/gradientflow.substack.com/p/heres-my-uncomfortable-bet-on-openai">Here&#8217;s my uncomfortable bet on OpenAI</a></strong></h3><p><em>Ben Lorica, Gradient Flow &#8212; July 21, 2026</em></p><p>Lorica&#8217;s bet is that open-weight models eventually absorb most global AI spending on pricing pressure, efficiency, and control &#8212; and the reason the piece is worth your time is that he refuses to make it comfortable. He concedes that <strong>coding is precisely where US frontier models &#8220;ride a flywheel&#8221;</strong> with advantages &#8220;not closing quickly,&#8221; because serious users generate failure data and failure data trains better models. His prescription is architectural rather than ideological: build <strong>routers, OpenAI-compatible endpoints, and eval harnesses</strong>, and treat your prompts, code, and agent traces as strategic assets. In a week when both Anthropic and Google competed on cost per task rather than raw capability, that advice lands differently than it would have in January.</p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know that both Anthropic and Google spent this week competing on cost per task rather than capability &#8212; Opus 5 holding <span>5/5/</span>25 while doubling Frontier-Bench, Gemini 3.6 Flash selling 17% fewer output tokens &#8212; and you know Claude Code raised nested subagent depth from 1 to 3 exactly one week after three tools agreed to cap it. You didn&#8217;t spend your weekend reading to know this. I did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,</p><p>Angel Llosa</p><p><a href="https://www.linkedin.com/in/anllogui/">LinkedIn</a> &#183; <a href="https://x.com/anllogui">X</a> &#183; <a href="https://anllogui.medium.com/">Medium</a></p><p><em>Questions, disagreements, or a story I missed? Just hit reply.</em></p><p></p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #40]]></title><description><![CDATA[Kimi K3 lands at half of Opus's price, Bun's $165K AI rewrite ignites a code-review fight, and Claude Code, OpenCode, and Codex all ship agent guardrails in the same week.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-39-b11</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-39-b11</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 20 Jul 2026 06:05:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!elow!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Everything that mattered this week in AI-assisted development, distilled into three headlines, one must-read, and the takeaways behind them.</p><p><strong>This week, compiled:</strong></p><ul><li><p><strong>The Money:</strong> Indian AI coding startup Emergent becomes a unicorn &#8212; $130M Series C at $1.5B, revenue up 70% in four months, selling &#8220;an engineering team in a box&#8221; to non-technical founders</p></li><li><p><strong>The Fork:</strong> Claude Code&#8217;s <code>/fork</code> now spawns full autonomous background sessions instead of in-session subagents &#8212; one of seven releases the tool shipped in six days</p></li><li><p><strong>The Privacy Scare:</strong> xAI open-sources Grok Build&#8217;s ~845K-line Rust codebase after users discovered it was uploading local directories &#8212; SSH keys included &#8212; to Google Cloud Storage by default</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> Moonshot AI shipped Kimi K3 with no keynote and no model card &#8212; just a quiet overnight flip of kimi.com to a 2.8-trillion-parameter model. Self-reported benchmarks put it ahead of Claude Opus 4.8 Max and GPT-5.5 High, at $3/$15 per Mtok &#8212; roughly half Opus 4.8&#8217;s per-task cost. Open weights land July 27, but the API is live now, and Moonshot is already calling it the first &#8220;open 3T-class model.&#8221; This is what the open-weight pressure on frontier pricing looks like when it stops being incremental. <a href="https://simonwillison.net/2026/Jul/16/kimi-k3/">Read more &#8594;</a></p></blockquote><p>Growing at 20% new subscribers per week.</p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!elow!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!elow!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png 424w, /__u/substackcdn.com/image/fetch/$s_!elow!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png 848w, /__u/substackcdn.com/image/fetch/$s_!elow!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png 1272w, /__u/substackcdn.com/image/fetch/$s_!elow!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!elow!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png" width="1024" height="572" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:572,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:640241,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/207739283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!elow!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png 424w, /__u/substackcdn.com/image/fetch/$s_!elow!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png 848w, /__u/substackcdn.com/image/fetch/$s_!elow!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png 1272w, /__u/substackcdn.com/image/fetch/$s_!elow!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37180add-c218-4e2c-92d0-8bd4bbbac922_1024x572.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Key Takeaways</h2><ul><li><p><strong>Bun&#8217;s AI-driven Rust rewrite cost ~$165,000 in API spend &#8212; and Zig&#8217;s creator called it &#8220;unreviewed slop&#8221;:</strong> Jarred Sumner ran up to ~64 parallel Claude Code instances over 11 days to port 535,496 lines of Zig into over 1,009,000 lines of Rust across 6,502 commits, versus an estimated 3 engineers working a full year unassisted. Andrew Kelley&#8217;s public rebuttal: &#8220;the tests pass&#8221; is not the same argument as reviewed. <a href="https://www.theregister.com/devops/2026/07/14/zig-creator-calls-buns-claude-rust-rewrite-unreviewed-slop/5270743">Read more &#8594;</a></p></li><li><p><strong>Emergent&#8217;s $130M Series C values it at $1.5B</strong> just over a year after launch, with 200,000+ paying customers and annualized revenue up 70% in four months &#8212; proof that the market for non-technical &#8220;vibe coding&#8221; platforms is scaling as fast as the developer-facing tools. <a href="https://techcrunch.com/2026/07/15/indian-ai-coding-startup-emergent-becomes-a-unicorn-just-over-a-year-after-launch/">Read more &#8594;</a></p></li><li><p><strong>xAI&#8217;s Grok Build privacy incident forced an unusual kind of transparency:</strong> after reports that the CLI uploaded entire local directories &#8212; including SSH keys and password databases &#8212; to Google Cloud Storage with retention on by default, xAI disabled the feature, flipped retention to off-by-default, and released the full ~845K-line Rust codebase under Apache 2.0. <a href="https://simonwillison.net/2026/Jul/15/grok-build/">Read more &#8594;</a></p></li><li><p><strong>OpenCode&#8217;s v1.18.2 (July 15) stopped subagents from spawning nested subagents by default</strong>, adding a configurable <code>subagent_depth</code> limit &#8212; the same week Claude Code capped subagent spawns per session and Codex tightened dangerous-command detection. Three unrelated tools independently reached for the same fix within days of each other.</p></li><li><p><strong>Tetrate&#8217;s Agent Router Enterprise now brokers tokens automatically:</strong> set a budget cap once, and the gateway reroutes requests to an approved fallback model instead of failing the call when spend hits the ceiling. Tetrate&#8217;s own framing: &#8220;Agent token spend is the one line item the engineering organization can&#8217;t easily cap.&#8221; <a href="https://www.prnewswire.com/news-releases/tetrate-adds-token-brokering-capability-for-ai-code-gen-cost-management-through-agent-router-enterprise-302826151.html">Read more &#8594;</a></p></li><li><p><strong>A benchmark that hit #1 on Hacker News found Claude Code loads ~33K tokens of system prompt and tool schemas before it even reads your prompt</strong> &#8212; versus ~7K for OpenCode, and up to ~75K in a realistic MCP/plugin-heavy setup. The catch: Claude Code&#8217;s batching of tool calls can still make it cheaper overall on multi-step tasks, so the fixed overhead isn&#8217;t the whole cost picture. <a href="https://systima.ai/blog/claude-code-vs-opencode-token-overhead">Read more &#8594;</a></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://developers.openai.com/codex/changelog">Codex CLI ships four releases (v0.144.2 &#8594; v0.144.6): security hardening and model-config fixes</a></h3><p><em>OpenAI &#8212; July 13&#8211;18, 2026</em></p><p>Codex CLI&#8217;s week tracked the same theme running through Claude Code and OpenCode. v0.144.3 (July 13) added inline visualizations in Codex tasks, but the more consequential change came in v0.144.5 (July 16): <strong>improved dangerous-command detection</strong>, catching more forced <code>rm</code> forms and giving clearer rejection reasons when a command gets denied. v0.144.6 (July 18) closed out the week with refreshed bundled instructions for GPT-5.6 Sol, Terra, and Luna and corrected context windows (272,000 tokens). Four tools, one week, one instinct: tighten what the agent is allowed to do on its own.</p><div><hr></div><h3><a href="https://cursor.com/changelog/side-chat">Cursor in Slack adds plan-sharing, multi-repo support, and cross-channel messaging</a></h3><p><em>Cursor &#8212; July 17, 2026</em></p><p>Cursor&#8217;s Slack integration now <strong>shares a plan before it starts working</strong>, so you can redirect it early instead of discovering the wrong approach after the fact, and updates its status as it moves through each step. It can operate across multiple repositories in a single workspace with a one-click switcher, and &#8212; perhaps the most useful addition &#8212; it can now read from and post to other Slack channels and threads, pulling context from across a workspace rather than staying boxed into one conversation.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-07-14-security-reviews-now-available-in-the-github-copilot-app">Security reviews land in the GitHub Copilot app</a></h3><p><em>GitHub Changelog &#8212; July 14, 2026</em></p><p>The new <code>/security-review</code> slash command, now in public preview, scans in-flight code changes for injection, XSS, insecure data handling, path traversal, and weak crypto &#8212; with severity-scored, actionable remediation suggestions delivered before the code is ever committed. Another entry in the week&#8217;s broader story: as agents write more code unsupervised, the tooling around catching what they get wrong is catching up.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-07-14-github-copilot-in-visual-studio-june-update/">GitHub Copilot in Visual Studio&#8217;s June update: usage tracking, MCP trust baselines, and in-IDE PR review</a></h3><p><em>GitHub Changelog &#8212; July 14, 2026</em></p><p>Six changes bundled into one release: a refreshed usage-tracking window with proactive billing alerts, <strong>MCP server configurations now validated against a trusted baseline at startup</strong> (any change requires explicit approval), general availability for the C++/MSVC modernization agent, edit suggestions that span the whole active file instead of just the area near your cursor, PR references inline in Copilot Chat, and a full in-IDE PR review flow &#8212; browse, comment, approve, complete &#8212; without leaving Visual Studio.</p><div><hr></div><h3><a href="https://github.com/github/copilot-cli/releases">GitHub Copilot CLI ships six releases (v1.0.71-1 &#8594; v1.0.72-1): plan-mode hardening and MCP config persistence</a></h3><p><em>GitHub &#8212; July 14&#8211;17, 2026</em></p><p>Another tool, the same instinct as the rest of this week&#8217;s shipping cycle: <strong>plan mode now hard-blocks built-in tool calls that would edit files or run mutating shell commands</strong>, closing a gap where &#8220;just planning&#8221; could still touch the filesystem. Smaller but useful fixes rounded out the week &#8212; GitHub MCP toolset/tool configuration now persists via <code>settings.json</code>, disabled skills are flagged in <code>copilot skill list</code>, invalid <code>settings.json</code> values now surface a specific warning instead of failing silently, and a bug that hung autopilot mode during long-running background processes was fixed (v1.0.71, July 16).</p><div><hr></div><h3><a href="https://zed.dev/releases/stable/1.11.3">Zed 1.11.3: dedicated Staged/Unstaged Changes views and Git Graph improvements</a></h3><p><em>Zed &#8212; July 15, 2026</em></p><p>Zed&#8217;s latest stable release adds dedicated <strong>Staged Changes and Unstaged Changes multibuffers</strong>, letting you review, stage, unstage, or restore individual hunks without leaving the editor &#8212; opened directly via <code>git: view staged changes</code> / <code>git: view unstaged changes</code>. Git Graph gets configurable columns and commit messages rendered as Markdown, and the project symbols picker adds a preview pane, alongside 60+ other fixes.</p><div><hr></div><p>This was the week the ecosystem quietly agreed that more autonomy needs more fencing. Claude Code capped how many subagents and web searches a session can spawn on its own and patched a run of permission-check bypasses in the same breath it shipped <code>/fork</code> for full background sessions. OpenCode, days later, stopped subagents from spawning their own subagents by default. Codex tightened what counts as a dangerous command, and GitHub Copilot CLI closed the gap where &#8220;plan mode&#8221; could still quietly edit files. Four tools, no coordination, the same fix &#8212; because agents that can now run unattended for longer need harder limits on what &#8220;unattended&#8221; is allowed to mean. And in case anyone needed the counterargument spelled out, Zig&#8217;s creator spent the week arguing publicly that Bun&#8217;s own million-line, $165,000, AI-generated Rust rewrite is exactly the kind of code those guardrails exist to catch &#8212; reviewed or not. Meanwhile Kimi K3 landed without ceremony at half of Opus&#8217;s price, Emergent proved the non-technical end of this market is scaling just as fast, and a privacy scare at xAI turned into the most detailed public look yet at what a frontier coding agent&#8217;s codebase actually contains.</p><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know why three unrelated coding agents all shipped subagent-spawn limits in the same six days, and you know Kimi K3 undercut Opus 4.8 pricing by half without so much as a press release. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,</p><p>The Vibe Coding Team.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #39]]></title><description><![CDATA[Three frontier AI coding models launched in 48 hours as manual code review collapses and token spend becomes the new cost lever.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-39</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-39</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 13 Jul 2026 08:00:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!QPMb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week, compiled:</strong></p><ul><li><p><strong>The Data:</strong> The share of developers who stopped manually reviewing AI-generated commits jumped from <strong>10% to 40% in a single month</strong></p></li><li><p><strong>The Money:</strong> Bun rewrote a core subsystem in Rust with AI in <strong>11 days for $165K in tokens</strong> &#8212; work estimated at a year for a small human team</p></li><li><p><strong>The Trend:</strong> Token budgets, not headcount, are the flexible line item &#8212; one security firm cut LLM spend <strong>59&#8211;70%</strong> by fixing its cache hit rate</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> Three frontier labs shipped coding-capable models within a single 48-hour window. OpenAI took <strong>GPT-5.6</strong> to general availability on July 9 as a three-tier family &#8212; Sol (flagship), Terra (balanced), and Luna (fast/cheap) &#8212; landing in ChatGPT, Codex, and the API, one day after xAI&#8217;s Grok 4.5 and the same day as Meta&#8217;s Muse Spark 1.1. If your team standardized on a model last quarter, this is the week that decision reopened. <a href="https://openai.com/index/gpt-5-6/">Read more &#8594;</a></p></blockquote><div><hr></div><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones<br>actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you<br>arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework<br>for the conversation that always comes after &#8220;we should use AI more&#8221;: how to<br>actually move an organization that didn&#8217;t ask to be moved. Included with every<br>subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!QPMb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!QPMb!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png 424w, /__u/substackcdn.com/image/fetch/$s_!QPMb!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png 848w, /__u/substackcdn.com/image/fetch/$s_!QPMb!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png 1272w, /__u/substackcdn.com/image/fetch/$s_!QPMb!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!QPMb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png" width="1024" height="572" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:572,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:721688,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/206665653?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!QPMb!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png 424w, /__u/substackcdn.com/image/fetch/$s_!QPMb!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png 848w, /__u/substackcdn.com/image/fetch/$s_!QPMb!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png 1272w, /__u/substackcdn.com/image/fetch/$s_!QPMb!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa885ab61-9c1f-48b7-969a-ddd37e5ad6a4_1024x572.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Key Takeaways</h2><ul><li><p><strong>AI code review discipline is collapsing faster than anyone predicted:</strong> median Cursor users generate ~700 lines of code weekly while the top 1% generate <strong>30,000&#8211;40,000 lines</strong> (roughly 45 median developers&#8217; worth) &#8212; and the share of developers who no longer manually review AI-generated commits quadrupled to <strong>40% in March 2026 alone</strong>. <a href="https://newsletter.pragmaticengineer.com/p/the-pulse-interesting-ai-coding-stats">Read more &#8594;</a></p></li><li><p><strong>&#8220;AI did the rewrite&#8221; now has a price tag:</strong> Bun&#8217;s AI-assisted Rust rewrite of a core subsystem took <strong>11 days and $165K in tokens</strong>, versus an estimated year for a small human team &#8212; one of the first public case studies with hard cash numbers instead of anecdote. <a href="https://newsletter.pragmaticengineer.com/p/the-pulse-bun-rust-rewrite">Read more &#8594;</a></p></li><li><p><strong>Caching is the cheapest engineer you&#8217;ll hire this year:</strong> security firm ProjectDiscovery raised its cache hit rate from <strong>7% to 84%</strong> and cut LLM spend <strong>59&#8211;70%</strong> &#8212; while Uber handed AI tools to 5,000 engineers and burned its entire 2026 AI budget by April. <a href="https://www.artificialintelligence-news.com/news/shrink-token-budget-not-team/">Read more &#8594;</a></p></li><li><p><strong>xAI trained Grok 4.5 alongside Cursor&#8217;s data flywheel</strong> &#8212; a rare public confirmation of a frontier lab and an agentic IDE co-training a model, blurring the line between model provider and tool vendor. <a href="https://x.ai/news/grok-4-5">Read more &#8594;</a></p></li><li><p><strong>Meta finally put a price on its models:</strong> Muse Spark 1.1 is Meta&#8217;s first paid model &#8212; 1M-token context at <strong>$1.25/$4.25 per million tokens</strong>, roughly a quarter of Anthropic and OpenAI&#8217;s rates, launched via a new public-preview Meta Model API. <a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">Read more &#8594;</a></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code Ships Six Releases in a Week (v2.1.202 &#8594; v2.1.207)</a></h3><p>Claude Code enabled <strong>auto mode by default</strong> on Bedrock, Vertex AI, and Foundry (opt-out via <code>disableAutoMode</code>), bumped Bedrock to <strong>Opus 4.8</strong>, and reworked agent-view rows to show classifier-written headlines instead of raw tool-call text. The fixes matter too: a terminal freeze on long streamed outputs, a <strong>consent-dialog bypass on remote managed settings</strong>, and an auto-updater bug that clobbered custom launcher symlinks. If you run Claude Code through a cloud provider, the new default is worth a deliberate look before your next session.</p><h3><a href="https://developers.openai.com/codex/changelog">OpenAI Codex Enables Remote Plugins by Default, Joins ChatGPT Desktop</a></h3><p>Across CLI 0.143.0&#8211;0.144.1 (July 6&#8211;9), Codex turned on <strong>remote plugins by default</strong> with a richer catalog and proxy-aware auth on macOS/Windows, added Bedrock GPT-5.6 models, and shipped <strong>configurable rollout token budgets</strong> that abort a turn when the budget is exhausted &#8212; a native cost-control primitive landing directly in the agent. Codex also moved into the <strong>ChatGPT desktop app</strong>, letting users edit code and review PR diffs without leaving ChatGPT.</p><h3><a href="https://github.com/sst/opencode/releases">OpenCode Adds Code-Mode MCP Adapter, Fixes Copilot Pricing Crash</a></h3><p>Across four releases (July 7&#8211;9, through v1.17.18), OpenCode shipped a <strong>code-mode MCP adapter</strong>, a v2 free-model selector for faster access to more providers, and broad desktop/TUI workflow fixes. The most instructive item is a crash caused by GitHub Copilot returning models with a <strong>zero billing batch size</strong> &#8212; a reminder that provider-side pricing metadata bugs can break your agents downstream, whether or not you touched anything.</p><h3><a href="https://geminicli.com/docs/changelogs/latest/">Gemini CLI v0.50.0 Adds Tool Registry Discovery</a></h3><p>The stable v0.50.0 release (July 8) introduced <strong>tool registry discovery</strong> for automatic detection and registration of tools, plus CI pipeline safeguards to stop bad npm releases from shipping. A same-day preview (v0.51.0-preview.0) hardened <strong>path-resolution security</strong> for at-reference files &#8212; a quiet but meaningful fix given how much context flows through file references in agent sessions.</p><h3><a href="https://github.blog/changelog/2026-07-08-github-copilot-in-visual-studio-code-june-2026-releases/">GitHub Copilot in VS Code Ships Agentic Browser Tools, Parallel Sessions</a></h3><p>VS Code&#8217;s Copilot agent can now <strong>browse and validate web apps directly</strong> (agentic browser tools are GA), run <strong>multiple parallel agent sessions</strong> with separate chats, and expose <strong>session-level and subagent-level credit cost tracking</strong> &#8212; cost observability arriving as a first-class IDE feature, not an afterthought. A new Language Models editor also lets you discover and install model-provider extensions on the fly.</p><h3><a href="https://cursor.com/changelog">Cursor 3.11 Adds Side Chats and Agent Transcript Search</a></h3><p>Cursor 3.11 (July 10) lets you branch off a <strong>side chat</strong> (via <code>/side</code>, <code>/btw</code>, or a plus button) to explore tangents without derailing the main agent conversation, and adds <strong>full-text search across past agent transcripts</strong>. New cloud-agent hooks (<code>beforeSubmitPrompt</code>, <code>afterAgentResponse</code>, <code>afterAgentThought</code>) give teams programmatic visibility into agent runs &#8212; the audit layer enterprises keep asking for.</p><h3><a href="https://zed.dev/releases/stable">Zed Adds llama.cpp as a Model Provider, Day-One GPT-5.6 Support</a></h3><p>Zed 1.10 (July 8&#8211;10, three releases) added <strong>llama.cpp as a first-class language model provider</strong> &#8212; a notable nod to local and self-hosted model workflows &#8212; and unified LLM providers, external agents, and MCP servers into a single settings editor. Within <strong>24 hours</strong> of GPT-5.6&#8217;s release, Zed shipped support for Sol and Terra, with Luna pending third-party API access.</p><h2>&#128161; Others</h2><h3><a href="https://newsletter.pragmaticengineer.com/p/tech-jobs-market-2026-part-3">Tech Jobs Market in 2026, Part 3: Hiring Managers &amp; Job Seekers</a></h3><p>The Pragmatic Engineer examines how AI adoption is reshaping engineering hiring: which <strong>AI-related roles are hottest</strong>, and why conditions remain tough for engineering leaders despite strong demand for AI skills. Worth reading if you&#8217;re on either side of the hiring table &#8212; the skills premium is real, but it isn&#8217;t landing where most job seekers expect.</p><div><hr></div><p>The arc of this week is simple: the models multiplied, and the money got named. Three frontier launches in 48 hours would have been the whole story a year ago &#8212; this week it shared the stage with a $165K rewrite invoice, a cache hit rate that saved more budget than a layoff round, and 40% of developers waving AI commits through unreviewed. The tools are no longer the bottleneck. The economics and the discipline are.</p><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes<br>everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped<br>three viral threads that turned out to be nothing. You know which of the three new frontier models actually changes your team&#8217;s pricing math, and you know that the scariest number this week wasn&#8217;t a benchmark &#8212; it was 40% of developers no longer reviewing AI commits. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and<br>everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,</p><p><br>The Vibecoding Team.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #38]]></title><description><![CDATA[Fable 5 returns worldwide &#8212; and in the same seven days, Sonnet 5 becomes the default model, background agents the default mode, and Manual the default permission.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-38</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-38</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 06 Jul 2026 04:38:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Weml!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week, compiled:</strong></p><ul><li><p><strong>The Big Model:</strong> Claude Sonnet 5 launches with performance approaching Opus 4.8 at introductory $2/$10 per Mtok &#8212; now the default model inside Claude Code v2.1.197</p></li><li><p><strong>The Agent:</strong> Claude Code background subagents run by default and auto-commit, push, and open draft PRs when they finish &#8212; your terminal now manages its own ticket queue</p></li><li><p><strong>The Ecosystem:</strong> Kimi K2.7 Code becomes the first open-weight model in GitHub Copilot&#8217;s model picker, cracking open a platform that had been closed-weight since day one</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> The US government lifted export controls on Fable 5 and Mythos 5 on June 30. What matters more than the return itself is what Anthropic published alongside it: a cross-industry jailbreak severity framework (CJS-0 to CJS-4) co-developed with Amazon, Microsoft, and Google &#8212; the first attempt to give every lab, researcher, and regulator a shared language for how dangerous a jailbreak actually is. This is what AI safety governance looks like when it starts to get real. <a href="https://www.anthropic.com/news/redeploying-fable-5">Read more &#8594;</a></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Key Takeaways</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Weml!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Weml!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png 424w, /__u/substackcdn.com/image/fetch/$s_!Weml!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png 848w, /__u/substackcdn.com/image/fetch/$s_!Weml!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Weml!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Weml!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png" width="1024" height="572" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:572,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:813165,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/205452328?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Weml!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png 424w, /__u/substackcdn.com/image/fetch/$s_!Weml!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png 848w, /__u/substackcdn.com/image/fetch/$s_!Weml!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Weml!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a69e18-6228-41c0-a3e7-2ba75fc0ac46_1024x572.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p><strong>Claude Sonnet 5 is the best value Anthropic has ever shipped:</strong> near-Opus 4.8 performance, introductory pricing of $2 input / $10 output per Mtok through August 31, available across all plans and inside Claude Code as the new default. <a href="https://www.anthropic.com/news/claude-sonnet-5">Read more &#8594;</a></p></li><li><p><strong>Claude Code shipped 6 releases in 7 days</strong>, with v2.1.198 being the pivotal one: background subagents now run in parallel by default, auto-committing, pushing, and opening draft PRs when they finish. Claude in Chrome also hit general availability. Six releases, one week, fundamentally different terminal. <a href="https://github.com/anthropics/claude-code/releases/tag/v2.1.198">Read more &#8594;</a></p></li><li><p><strong>Kimi K2.7 Code is the first open-weight model in GitHub Copilot&#8217;s model picker</strong> &#8212; available now for Pro, Pro+, and Max; requires admin opt-in for Business and Enterprise. It&#8217;s one switch, but it signals a structural shift in how GitHub thinks about model diversity on the platform. <a href="https://github.blog/changelog/2026-07-01-kimi-k2-7-is-now-available-in-github-copilot/">Read more &#8594;</a></p></li><li><p><strong>Copilot browser tools reached GA in VS Code:</strong> agents can now navigate, click, type, take screenshots, and feed live web content directly back into chat &#8212; with user-controlled tab sharing and fully isolated sessions that have no access to your cookies or browsing history. <a href="https://github.blog/changelog/2026-07-01-browser-tools-for-github-copilot-in-vs-code-are-generally-available/">Read more &#8594;</a></p></li><li><p><strong>Claude Science launched as an AI workbench for researchers</strong> with 60+ pre-configured skills across genomics, proteomics, and cheminformatics, a reviewer agent that checks citations and flags errors, and scaling from laptops to HPC clusters. Applications open through July 15 for up to $30K in credits per project. <a href="https://www.anthropic.com/news/claude-science-ai-workbench">Read more &#8594;</a></p></li><li><p><strong>Claude Opus 4.8 and Haiku 4.5 are now generally available in Microsoft Foundry on Azure,</strong> with Azure-native billing, authentication, networking, and US data residency. Eligible customers with a Microsoft Enterprise Agreement can draw Claude usage down against their existing Azure commitment. <a href="https://claude.com/blog/claude-in-microsoft-foundry">Read more &#8594;</a></p></li><li><p><strong>Microsoft Research&#8217;s Memora cuts AI agent context token usage by up to 98%</strong> using a harmonic memory architecture that decouples what is stored from how it&#8217;s retrieved &#8212; achieving state-of-the-art memory accuracy (87.4% on LongMemEval) while storing half the memory entries of Mem0 per conversation. As agentic coding tools make 10&#8211;100x more LLM calls than chatbots, techniques like Memora are becoming the FinOps layer for AI teams. Code is on GitHub; paper presented at ICML 2026. <a href="https://www.microsoft.com/en-us/research/blog/memora-a-harmonic-memory-representation-balancing-abstraction-and-specificity/">Read more &#8594;</a></p></li></ul><div><hr></div><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription <strong>FOR FREE</strong>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://www.anthropic.com/news/fable-safeguards-jailbreak-framework">More details on Fable 5&#8217;s cyber safeguards and our jailbreak framework</a></h3><p><em>Anthropic &#8212; July 2, 2026</em></p><p>Two days after Fable 5&#8217;s return, Anthropic published the technical details most developers will actually want. The model&#8217;s cybersecurity classifier now divides use cases into four tiers &#8212; from outright prohibited (ransomware, malware, exfiltration) to benign IT work (secure coding, patch management) &#8212; with the classifier accepting <strong>higher false-positive rates in favor of safety margin</strong>. Separately, the Cyber Jailbreak Severity (CJS) framework scores any reported jailbreak across four axes: capability gain, breadth of enabling offensive tasks, ease of weaponization, and discoverability, producing a score from CJS-0 (Informational) to CJS-4 (Critical). The framework is co-developed with Amazon, Microsoft, and Google, and is designed to be adopted across labs &#8212; not just Anthropic.</p><div><hr></div><h3><a href="https://github.com/anthropics/claude-code/releases/tag/v2.1.200">Claude Code v2.1.200 &#8212; Manual permissions as the new default</a></h3><p><em>Anthropic &#8212; July 3, 2026</em></p><p>The quietest but most consequential change in the week&#8217;s shipping sprint: <strong>Claude Code&#8217;s default permission mode is now &#8220;Manual&#8221;</strong> across CLI, VS Code, and JetBrains. AskUserQuestion dialogs no longer auto-continue (you can re-enable idle timeout via /config). The release also fixed background session failures after sleep/wake cycles, resolved subagents cut off by rate limits returning empty results instead of errors, and improved terminal rendering under tmux 3.4+. For teams running unattended agents, these reliability fixes matter as much as any feature.</p><div><hr></div><h3><a href="https://opencode.ai/changelog">OpenCode v1.17.12 and v1.17.13 &#8212; Sonnet 5 adaptive thinking and new model picker</a></h3><p><em>OpenCode &#8212; June 30 and July 1, 2026</em></p><p>OpenCode shipped two releases in two days to respond to the Sonnet 5 launch. <strong>v1.17.12 enables adaptive thinking for Claude Sonnet 5</strong> alongside improved MCP OAuth (refresh-token scope, better error surfaces) and token/cost totals visible in session context. <strong>v1.17.13 adds a searchable v2 model picker with model management flow</strong>, session tab hover previews showing project, path, branch, and server at a glance, and streamlined WSL server setup. Taken together: OpenCode turned around Sonnet 5 compatibility in under 24 hours of the model going live.</p><div><hr></div><h3><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/">Google launches Nano Banana 2 Lite and Gemini Omni Flash</a></h3><p><em>Google Blog &#8212; June 30, 2026</em></p><p>Two developer-facing models shipped the same day as Sonnet 5. <strong>Nano Banana 2 Lite</strong> is Google&#8217;s fastest and cheapest image model &#8212; 4-second generation times at $0.034 per 1K images, available immediately in Google AI Studio and the Gemini API. <strong>Gemini Omni Flash</strong> enters public preview as a natively multimodal video generation model supporting conversational editing from text, image, and video inputs at $0.10 per second of output. Both are accessible through the Gemini API, Google AI Studio, and the Enterprise Agent Platform starting today.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-07-02-copilot-cli-no-longer-needs-a-personal-access-token-in-github-actions/">Copilot CLI no longer needs a PAT in GitHub Actions</a></h3><p><em>GitHub Changelog &#8212; July 2, 2026</em></p><p>A simple change that removes significant operational overhead: <strong>GitHub Copilot CLI in Actions now authenticates with the built-in GITHUB_TOKEN</strong>, requiring only the <code>copilot-requests: write</code> permission. Personal access tokens &#8212; long-lived credentials that create security exposure and rotation headaches at scale &#8212; are no longer needed. AI credits bill directly to the organization. If your CI uses Copilot CLI today, this is a drop-in improvement with no behavior change.</p><div><hr></div><h3><a href="https://poolside.ai/blog/introducing-laguna-xs-2-1">Introducing Laguna XS 2.1</a></h3><p><em>Poolside &#8212; July 2, 2026</em></p><p>As frontier API costs become harder to justify at scale, the open-weight option keeps getting better. Poolside&#8217;s <strong>Laguna XS 2.1 is a 33B-parameter MoE model that runs on a single GPU</strong>, achieves 63.1% on SWE-bench Multilingual (+5.4 points over its predecessor), and is available free on Hugging Face and OpenRouter under the fully permissive OpenMDW-1.1 license. For teams that need local agentic coding &#8212; private codebases, air-gapped environments, or simply more predictable infrastructure costs &#8212; Laguna XS 2.1 is a credible alternative to frontier API tools, deployable via Ollama, llama.cpp, vLLM, or TensorRT-LLM, with a paid hosted tier at $0.10/$0.20 per Mtok if you prefer not to self-host.</p><div><hr></div><h2>&#128161; Others</h2><h3><a href="https://github.blog/changelog/2026-07-02-copilot-agent-session-streaming-is-now-in-public-preview/">Copilot agent session streaming is now in public preview</a></h3><p><em>GitHub Changelog &#8212; July 2, 2026</em></p><p>Enterprise visibility into agent activity just became possible. <strong>GitHub Enterprise Cloud customers can now stream Copilot agent session data</strong> &#8212; every prompt, response, and tool call &#8212; across all Copilot clients to a SIEM, event collector, or Microsoft Purview endpoint. Alternatively, pull up to 48 hours of records on-demand via REST API. As agents become standard in enterprise workflows, having an audit trail for what they did and why is increasingly non-negotiable &#8212; this is the first version of that answer from GitHub.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-07-02-upcoming-deprecation-of-gemini-2-5-pro-and-gemini-3-flash">Upcoming deprecation of Gemini 2.5 Pro and Gemini 3 Flash in GitHub Copilot</a></h3><p><em>GitHub Changelog &#8212; July 2, 2026</em></p><p><strong>Action required if your team uses Gemini models via Copilot:</strong> GitHub has announced the upcoming deprecation of both Gemini 2.5 Pro and Gemini 3 Flash from the Copilot model picker. Teams relying on either should plan a migration to Gemini 3 Pro or an alternative in the picker before the deprecation date. Given the simultaneous announcement that Kimi K2.7 Code is now available, this week&#8217;s GitHub Copilot model story is really about the picker becoming a more dynamic surface &#8212; older models out, new open-weight options in.</p><div><hr></div><p>This was a week where the word &#8220;default&#8221; did a lot of work. Sonnet 5 is now the default model in Claude Code. Background agents are now the default operating mode. Manual permissions are now the default safety posture. Anthropic shipped six releases in seven days, and each one quietly changed what it means to open a terminal. Meanwhile, Fable 5 came back &#8212; and what came with it was the first real attempt at a shared industry standard for measuring how dangerous a jailbreak is. A government suspended a model for three weeks and a cross-company safety framework came out the other side. And in the background, a Microsoft Research team published a paper showing you can cut your agent&#8217;s context token bill by 98% &#8212; because as your agents get more capable, the question of what they cost per task is no longer optional.</p><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know Fable 5 is back globally, why it was suspended, and what the CJS framework means for every lab that ships frontier models. You know Claude Sonnet 5 benchmarks near Opus 4.8 at less than half the price, that Claude Code is now opening your pull requests in the background, and that Microsoft Research just published a technique to cut your agent&#8217;s token bill by 98%. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #37]]></title><description><![CDATA[The week AI dev tooling started feeling like regulated infrastructure]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-37</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-37</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 29 Jun 2026 06:21:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ZFcV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Big Story:</strong> Anthropic launches Claude Tag &#8212; a persistent AI teammate inside Slack that 65% of Anthropic&#8217;s own product team uses to write code, with multiplayer collaboration, async task delegation, and admin-controlled tool access.</p></li><li><p><strong>The Release:</strong> Claude Code shipped five versions in one week, cutting streaming CPU usage by 37%, adding <code>/rewind</code> to undo a <code>/clear</code>, and routing all shell commands through auto-mode classification.</p></li><li><p><strong>The Efficiency Win:</strong> GitHub&#8217;s Copilot agentic harness, benchmarked against Claude, GPT-5.4, and GPT-5.5 on five industry benchmarks, matches model-vendor harnesses on task completion while consuming fewer tokens &#8212; the platform layer is quietly becoming the competitive moat.</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> OpenAI previewed GPT-5.6 &#8212; three models named Sol, Terra, and Luna &#8212; but almost nobody can use them yet. For the first time, a US executive order required OpenAI to coordinate with the federal government before releasing a frontier AI model to the public. Only ~20 trusted partner organizations have access. Sol handles complex coding and security research; Terra does the same work as GPT-5.5 at half the cost; Luna is the fastest and cheapest for everyday tasks. The restriction isn&#8217;t just a regulatory footnote &#8212; it&#8217;s the first time the US government has formally inserted itself into the frontier model release cycle. <a href="https://venturebeat.com/technology/openai-unveils-gpt-5-6-sol-terra-and-luna-models-but-only-accessible-to-limited-preview-partners-for-now-per-us-gov/">Read more &#8594;</a></p></blockquote><div><hr></div><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>Key Takeaways</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ZFcV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ZFcV!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!ZFcV!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!ZFcV!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ZFcV!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ZFcV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5847356,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/204070397?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ZFcV!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!ZFcV!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!ZFcV!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ZFcV!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb614280-23f6-476e-9ba5-b2b8134050d1_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><ul><li><p><strong>Microsoft&#8217;s first own-brand coding model is now GA on Copilot.</strong> MAI-Code-1-Flash, built entirely in-house by Microsoft AI, became generally available for Copilot Business and Enterprise on June 26. It&#8217;s optimized for high-volume iterative agentic workflows &#8212; fast and low-latency by design, billed at usage-based rates. Admins need to manually enable the policy. This marks Microsoft&#8217;s first step from model reseller to model maker inside its own developer platform. <a href="https://github.blog/changelog/2026-06-26-mai-code-1-flash-for-copilot-business-and-copilot-enterprise/">Read more &#8594;</a></p></li><li><p><strong>Codex Remote is generally available: AI coding you can approve from your phone.</strong> As of June 25, developers can start or continue Codex sessions on a connected Mac or Windows host and monitor and approve actions from an iOS or Android device. The release adds authenticated QR pairing between mobile device and host, and ships a new DigitalOcean plugin that provisions a cloud Droplet as a remote workspace from inside the Codex app. <a href="https://developers.openai.com/codex/changelog">Read more &#8594;</a></p></li><li><p><strong>GitHub Actions can now run steps in parallel.</strong> Four new keywords landed on June 25: <code>background: true</code> starts a step asynchronously; <code>wait</code>/<code>wait-all</code> blocks until specific background steps finish; <code>cancel</code> gracefully stops a background step; and <code>parallel</code> wraps multiple steps in concurrent execution with automatic waiting. Common use cases include running parallel builds, spinning up services while other steps proceed, and non-blocking telemetry uploads. <a href="https://github.blog/changelog/2026-06-25-actions-steps-can-now-be-run-in-parallel/">Read more &#8594;</a></p></li><li><p><strong>Free and Student Copilot users can no longer choose their AI model.</strong> Starting June 24, GitHub replaced manual model selection on Free and Student Copilot plans with automatic routing only. &#8220;Auto&#8221; dynamically picks the best available model per task from across the OpenAI, Anthropic, Google, and Microsoft model families &#8212; but which models are in the pool can change without notice. GitHub also retired the &#8220;(Preview)&#8221; label from all Microsoft-released models as part of the same update. <a href="https://github.blog/changelog/2026-06-24-changes-to-model-selection-for-free-and-student-plans/">Read more &#8594;</a></p></li><li><p><strong>5x more pull requests, but code review is collapsing.</strong> The Pragmatic Engineer&#8217;s Gergely Orosz published a June 23 analysis of what AI productivity data actually means at scale. Teams using agents ship 5x more PRs and generate 2.5x more code per developer than two years ago &#8212; but Cursor&#8217;s own data shows AI-generated code is increasingly being merged <strong>without any human review</strong>, a trend that accelerated sharply in February 2026. Meta reallocated engineers from security and quality to AI training tasks, then suffered a high-profile breach. Orosz&#8217;s conclusion: as output volume becomes cheap and automatic, &#8220;taste&#8221; &#8212; deliberate human judgment about what to ship &#8212; becomes the scarcest skill in the room. <a href="https://newsletter.pragmaticengineer.com/p/slow-down-to-speed-up">Read more &#8594;</a></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>&#128230; Releases &amp; News</h2><p><strong><a href="https://github.blog/changelog/2026-06-26-github-desktop-3-6-worktrees-and-deeper-copilot-integration/">GitHub Desktop 3.6: Worktrees and Deeper Copilot Integration</a></strong><br>GitHub Desktop 3.6, released June 26, adds native Git worktree support so developers can work across multiple branches simultaneously without stashing or re-cloning &#8212; a change that matters most when running parallel AI agent sessions on the same repository. Copilot now generates commit messages that respect custom instructions from <code>.github/copilot-instructions.md</code> and <code>AGENTS.md</code>, and provides AI-assisted merge conflict resolution. A <strong>model picker</strong> lets users choose from all Copilot-accessible models or bring their own key.</p><p><strong><a href="https://github.blog/changelog/2026-06-25-github-copilot-for-jira-is-now-generally-available/">GitHub Copilot for Jira Is Now Generally Available</a></strong><br>Copilot for Jira exited public preview on June 25 with three GA additions: real-time progress monitoring of the coding agent directly inside a Jira issue, the ability to give follow-up instructions in the Jira chat panel after the agent creates a draft PR (so it keeps refining the same PR instead of creating a new one), and streamlined setup requiring fewer configuration steps. Preview-phase features &#8212; model selection, Confluence context via MCP, custom agents, space-level guidance, and review request notifications &#8212; are all now generally available too.</p><p><strong><a href="https://github.blog/changelog/2026-06-22-new-features-and-claude-as-agent-provider-preview-in-jetbrains-ides/">Claude as Agent Provider in JetBrains IDEs &#8212; Public Preview</a></strong><br>As of June 22, JetBrains IDE users can select Claude as their Copilot agent provider by pointing the IDE to the Claude Code CLI path. Important caveat for the preview: <strong>Claude currently runs in bypass permissions mode</strong>, meaning all file edits and tool calls are auto-approved without confirmation. Configurable permissions are on the roadmap. The release also ships organization and enterprise agent support, a model picker, and a per-turn AI credits indicator.</p><p><strong><a href="https://cursor.com/changelog">Cursor 3.9: Customize Page, Plugin Marketplace, and Team Workspaces</a></strong><br>Cursor 3.9 landed on June 22 with a new Customize page that centralizes management of plugins, skills, MCPs, subagents, rules, commands, and hooks across user, team, and workspace scopes. A <strong>marketplace leaderboard</strong> surfaces the most-used extensions across your team for one-click installation. Plugin canvases ship prebuilt templates: Hex Canvas for data visualizations, Atlassian Canvas for real-time project visibility. Team Marketplaces now import plugin repositories from GitLab, Bitbucket, and Azure DevOps &#8212; not just GitHub.</p><p><strong><a href="https://github.com/sst/opencode/releases">OpenCode v1.17.11: Session Snapshots and Revert Controls</a></strong><br>OpenCode&#8217;s June 25 release adds the ability to roll a session back to any earlier message, including all associated file changes &#8212; a safety net that makes it easier to experiment aggressively and recover cleanly. Desktop improvements add <strong>Chrome-style tab cycling</strong> and draggable tabs. Bug fixes cover prompt draft attachment persistence and titlebar tab consistency. This brings session-level undo to what has been a rapid release cadence from the sst/opencode team.</p><p><strong><a href="https://github.com/sst/opencode/releases">OpenCode v1.17.10: MCP Server Instructions and Managed Provider Integration</a></strong><br>The June 24 OpenCode release injects <strong>MCP server instructions directly into the session context</strong>, giving the model visibility into what each connected MCP server can do without requiring a separate lookup. It also adds support for Opencode-managed provider integrations, a configurable keybind for the diff viewer, mobile bottom navigation in desktop mode, and collapsible server sections in the sidebar.</p><div><hr></div><p>The week&#8217;s theme is consolidation under pressure. GPT-5.6 lands behind a government gate. Claude Tag turns AI into a team member with a Slack handle. GitHub ships its own model. Harnesses are being benchmarked like infrastructure. And Orosz reminds us that the PRs-per-developer metric, taken alone, is a trap. The tools are genuinely getting better. The question of who&#8217;s watching the output isn&#8217;t going away.</p><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know GPT-5.6 exists but why almost nobody can use it yet, and you understand why a week with five Claude Code releases and a new Microsoft model actually matters for how your team works. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,</p><p>The Vibe Coding Team.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #36]]></title><description><![CDATA[Google killed Gemini CLI and 6,000 contributors' trust in one stroke. Plus Codex learns by watching you work, a single fake Sentry error hijacks every major coding agent.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-36</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-36</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Sun, 21 Jun 2026 08:07:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!y3L8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Big Story:</strong> Google shut down Gemini CLI on June 18 &#8212; 103K GitHub stars, 6,000 community contributors, and every CI/CD pipeline that depended on the <code>gemini</code> command stopped receiving responses overnight. The replacement, Antigravity CLI, is a closed-source Go binary with no 1:1 feature parity at launch.</p></li><li><p><strong>The Automation:</strong> Codex shipped Record &amp; Replay on macOS &#8212; show it a workflow once, and the language model converts the demonstration into a reusable, inspectable, editable skill. No scripting, no rule-based heuristics. Cursor matched the moment with /automate, Slack emoji triggers, and computer use for cloud agents &#8212; all on the same day.</p></li><li><p><strong>The Safety Net:</strong> Claude Code v2.1.183 now blocks destructive git commands in auto mode &#8212; <code>git reset --hard</code>, <code>git checkout -- .</code>, <code>git clean -fd</code>, and <code>git stash drop</code> are all rejected unless you explicitly ask for them. Same goes for <code>terraform destroy</code>, <code>pulumi destroy</code>, and <code>cdk destroy</code>.</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> A security team called Tenet published a new attack class called &#8220;Agentjacking&#8221; &#8212; and the results are uncomfortable. One fake Sentry error report, injected through a publicly accessible DSN, was enough to hijack AI coding agents into running attacker-controlled code on developer machines. Claude Code, Cursor, and Codex all fell for it with an <strong>85% success rate</strong>. The attack bypasses every traditional security control because every action in the chain is technically authorized. Tenet found <strong>2,388 organizations</strong> with exposed Sentry DSNs during their validation window. Sentry acknowledged the issue the same day it was disclosed &#8212; and declined to fix it, calling the attack class &#8220;technically not defensible&#8221; at the platform level. <a href="https://tenetsecurity.ai/blog/agentjacking-coding-agents-with-fake-sentry-errors/">Read more &#8594;</a></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><h2>Key Takeaways</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!y3L8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!y3L8!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!y3L8!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!y3L8!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y3L8!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!y3L8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5825813,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/202929797?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!y3L8!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!y3L8!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!y3L8!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y3L8!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e569193-30ef-4046-925c-ca44a92336a8_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><ul><li><p><strong>Gemini CLI&#8217;s shutdown is the open-source trust story of the year:</strong> Google accepted 6,000+ merged pull requests under Apache 2.0 with a CLA that granted perpetual, irrevocable rights &#8212; then used that code to build Antigravity CLI, a closed-source successor. Enterprise users with Gemini Code Assist Standard/Enterprise licenses keep access, but every free and Pro user lost their CLI on June 18 with no open-source alternative. <a href="https://www.techtimes.com/articles/318660/20260618/gemini-cli-shutdown-takes-effect-ci-cd-pipelines-break-go-based-antigravity-cli-arrives.htm">Read more &#8594;</a></p></li><li><p><strong>Claude Code Artifacts turn coding sessions into live, shareable web pages:</strong> Artifacts publish to a private URL on claude.ai, update in real-time as the session works, support version history, and enforce org-only visibility via strict CSP &#8212; 16 MiB per page, no external network requests. PR walkthroughs, dashboards, and incident pages that stay current without anyone manually updating them. <a href="https://claude.com/blog/artifacts-in-claude-code">Read more &#8594;</a></p></li><li><p><strong>Codex Record &amp; Replay teaches the agent by demonstration instead of description:</strong> Show Codex a workflow &#8212; uploading a YouTube video, filing an expense report &#8212; and it generalizes the steps into a reusable skill using the language model, not rule-based heuristics. The skill is inspectable and editable. Available on macOS for Plus, Pro, Business, Enterprise, and Edu tiers. <a href="https://developers.openai.com/codex/record-and-replay">Read more &#8594;</a></p></li><li><p><strong>Claude Design now imports your design system from GitHub and round-trips code with Claude Code:</strong> The /design and /design-sync commands create a bidirectional pipeline between design and code &#8212; the first time a major AI tool has closed the design-to-implementation loop without screenshots or manual rebuilds. Nine new export destinations (Adobe, Canva, Lovable, Replit, Vercel) and shared token limits across Claude products complete the overhaul. <a href="https://claude.com/blog/claude-design-stays-on-brand-for-daily-work">Read more &#8594;</a></p></li><li><p><strong>Cursor 3.8 makes automations triggerable from Slack and GitHub &#8212; with computer use for cloud agents:</strong> The /automate skill lets you describe an automation in plain language, but the real addition is five new GitHub triggers (issue comments, PR reviews, review thread resolution, Actions completion) and Slack emoji triggers that let teams kick off agent workflows without opening the IDE. <a href="https://cursor.com/changelog/06-18-26">Read more &#8594;</a></p></li></ul><div><hr></div><p>Growing at 20% new subscribers per week.</p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://docs.devin.ai/desktop/changelog">Devin Desktop v3.2.16: Plugin System for Devin Local</a></h3><p>Devin Local is becoming an extensible platform. The June 16 release introduces a <strong>plugin system</strong> (preview, opt-in for enterprises) that lets teams add custom tools, skills, and integrations to their local agent. Subagents can now call MCP tools directly &#8212; a meaningful expansion of what Devin Local can reach without routing through the cloud. Teams also gain CLI permission scopes to enforce terminal allow/deny lists, giving enterprises the governance layer they&#8217;ve been requesting since the Windsurf-to-Devin transition.</p><h3><a href="https://github.com/openai/codex/releases">Codex CLI v0.141.0: Encrypted Relay Channels and Plugin Marketplace</a></h3><p>The June 18 stable release ships <strong>87 changes</strong> with a clear security theme: remote executors now communicate through authenticated, end-to-end encrypted Noise relay channels, and cross-platform execution maintains native working directories across system boundaries. The plugin marketplace gets more informative automation &#8212; <code>codex plugin marketplace list --json</code> now includes each marketplace source, and plugin lists return from cached catalogs before refreshing in the background. Combined with tool search caching and reduced memory consumption, this is the most infrastructure-focused Codex release in months.</p><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code v2.1.183: Auto Mode Safety Blocks Destructive Commands</a></h3><p>The June 19 release is all about preventing the agent from doing things you didn&#8217;t ask for. Destructive git commands &#8212; <code>git reset --hard</code>, <code>git checkout -- .</code>, <code>git clean -fd</code>, <code>git stash drop</code> &#8212; are now <strong>blocked in auto mode</strong> unless you explicitly request them. The same goes for <code>git commit --amend</code> on commits the agent didn&#8217;t make, and <code>terraform destroy</code>, <code>pulumi destroy</code>, and <code>cdk destroy</code>. Also adds model deprecation warnings in print mode and a new <code>attribution.sessionUrl</code> setting to omit claude.ai session links from commits and PRs.</p><h3><a href="https://siliconangle.com/2026/06/17/ex-cisco-researchers-launch-tenet-security-lock-rogue-ai-agents/">Tenet Security Launches with $6M Seed Round and Agentjacking Disclosure</a></h3><p>The same day Tenet Security published its Agentjacking research, the company emerged from stealth with a <strong>$6M seed round</strong> led by The Westly Group (early SentinelOne investor). Founded by ex-Cisco AI Defense researchers, Tenet&#8217;s core technology &#8212; patent-pending &#8220;Agent-side Simulation&#8221; &#8212; predicts an AI agent&#8217;s likely next actions before they execute in production. The timing is deliberate: a new attack class and a funded startup to address it, launched on the same day.</p><h3><a href="https://opencode.ai/changelog">OpenCode v1.17.7 / v1.17.8 / v1.17.9: MCP Compatibility Push</a></h3><p>Three OpenCode releases in one week, all focused on making the MCP ecosystem work more reliably. v1.17.7 (June 14) fixes plugin client server reuse and adds ACP shell tool visibility. v1.17.8 (June 17) makes <strong>OpenAI-compatible providers accept MCP tool schemas</strong> that previously failed validation and fixes Cloudflare AI Gateway key handling. v1.17.9 (June 21) honors configured agent step limits and improves prompt caching. Small individually, but collectively they reflect OpenCode&#8217;s bet on being the most interoperable agent in the market.</p><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://kiro.dev/changelog/">Kiro CLI V3 Early Access: Spec-Driven Development Comes to the Terminal</a></h3><p>Kiro CLI V3 launches in early access on June 17, running alongside existing 2.x installations via <code>kiro-cli --v3</code>. The new version brings the unified agent harness used across IDE and Web to the terminal for the first time &#8212; with <strong>spec-driven development</strong>, a capability-based permissions model, enhanced hooks with a standalone file format, and tag-based agent configuration. Two days later (June 19), Kiro Web shipped Automations: schedule recurring work with GitHub/GitLab repos, up to 5 schedules per automation, each running autonomously in a sandbox and opening PRs when done.</p><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code v2.1.181: Inline Config, Bun 1.4, and Streaming Overhaul</a></h3><p>The June 17 release adds <code>/config key=value</code> syntax for changing settings without leaving the conversation &#8212; a small but meaningful friction reduction for power users who constantly toggle features. Under the hood, the <strong>bundled Bun runtime upgrades to 1.4</strong>, streaming now delivers long paragraphs line-by-line instead of waiting for the first break, and <code>sandbox.allowAppleEvents</code> opens macOS automation capabilities in sandboxed sessions. A new <code>CLAUDE_CLIENT_PRESENCE_FILE</code> environment variable lets server setups suppress mobile notifications.</p><div><hr></div><h2>&#128161; Others</h2><h3><a href="https://codingfleet.com/blog/terminal-bench-leaderboard-2026/">Fable 5 Tops Terminal-Bench 2.1 &#8212; But No One Can Use It</a></h3><p>The Terminal-Bench 2.1 leaderboard updated on June 17 with Claude Fable 5 entries &#8212; and the gap is striking. Fable 5 hit <strong>88.0%</strong>, the first model to break 85% on this benchmark, sitting <strong>4.6 points ahead</strong> of GPT-5.5 at 83.4%. Claude Code + Fable 5 scored 83.1%, Terminus 2 + Fable 5 at 80.4%. The irony: the model that topped the leaderboard remains suspended under a US export-control directive since June 12, meaning no one can actually use it.</p><div><hr></div><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know that Google killed Gemini CLI and 6,000 contributors&#8217; trust overnight, that a single fake Sentry error can hijack every major coding agent on the market, and that your Claude Code sessions can now become live web pages your team watches update in real time. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #35]]></title><description><![CDATA[The most capable coding model of the year shipped &#8212; and was pulled worldwide three days later. Plus Copilot's 50x billing shock and GitHub's all-in bet on becoming the platform for agents.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-35</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-35</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Sun, 14 Jun 2026 06:13:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nfGF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Big Story:</strong> GitHub Copilot&#8217;s switch to usage-based billing went live on June 1 &#8212; and developers are reporting 10x&#8211;50x cost increases overnight, with agentic coding sessions eating through a month of credits in a single sitting.</p></li><li><p><strong>The New Default:</strong> Windsurf retired after one year and shipped as Devin Desktop &#8212; Cascade replaced by a Rust-rewritten agent, Agent Command Center as the new home screen, and open ACP protocol letting any agent run alongside Devin in the same workspace.</p></li><li><p><strong>The Platform Bet:</strong> GitHub shipped Agentic Workflows in public preview, GitHub Copilot Sandboxes, and the Copilot SDK GA all in the same week &#8212; making GitHub the platform layer for running agents, not just the host for their output.</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> Claude Fable 5 launched on June 9 as the first publicly available Mythos-class model, and the benchmark gap is significant: 80.3% on SWE-bench Pro versus 69.2% for Opus 4.8 and 58.6% for GPT-5.5. The real number isn&#8217;t on a leaderboard, though &#8212; it&#8217;s in a Stripe engineering report. Fable 5 completed a codebase-wide migration of a 50-million-line Ruby codebase in one day. The same task had been estimated at two months of team work. At $10/$50 per million input/output tokens (2&#215; Opus 4.8), it was free on all paid Claude plans through June 22 &#8212; but the story took a sharp turn on June 12, when a US government export-control directive forced Anthropic to suspend access to Fable 5 and Mythos 5 worldwide, for every customer inside and outside the United States, citing national-security concerns over a potential jailbreak of Fable&#8217;s safeguards. Anthropic disputes the decision &#8212; calling the vulnerability a &#8220;narrow, non-universal jailbreak&#8221; no different from competing models &#8212; but is complying. All other Anthropic models remain available. The most capable coding model of the year was publicly accessible for three days. <a href="https://www.anthropic.com/news/fable-mythos-access">Read more &#8594;</a></p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Key Takeaways</h2><ul><li><p><strong>GitHub Copilot&#8217;s billing shock is the first crack in the &#8220;unlimited AI for a flat fee&#8221; model:</strong> Developers report projections jumping from $29/month to $750 and from $50/month to $3,000 &#8212; agentic coding sessions routinely hit $30&#8211;$40 each, blowing through a Pro user&#8217;s $10 monthly credit ceiling in one sitting. Code completions remain unlimited, but the economics of agentic workflows at scale have fundamentally changed. <a href="https://github.blog/changelog/2026-06-01-updates-to-github-copilot-billing-and-plans/">Read more &#8594;</a></p></li><li><p><strong>GitHub Agentic Workflows in public preview lets you define CI automation in plain Markdown:</strong> The system compiles natural-language workflow descriptions into standard GitHub Actions YAML, supports issue triage, CI failure analysis, and code simplification, and runs agents in sandboxed containers with an Agent Workflow Firewall &#8212; no personal access token required, GITHUB_TOKEN is enough. <a href="https://github.blog/changelog/2026-06-11-github-agentic-workflows-is-now-in-public-preview/">Read more &#8594;</a></p></li><li><p><strong>The Copilot SDK GA turns GitHub&#8217;s agentic engine into an embeddable building block:</strong> Available across six languages (Node.js, Python, Go, .NET, Rust, Java), it exposes custom tool registration, MCP server integration, lifecycle hook interception, and OpenTelemetry tracing &#8212; meaning any developer tool can now embed Copilot&#8217;s agentic reasoning layer, not just use it through GitHub&#8217;s surfaces. <a href="https://github.blog/changelog/2026-06-02-copilot-sdk-is-now-generally-available/">Read more &#8594;</a></p></li><li><p><strong>OpenCode sub-agents can now spawn sub-agents up to 5 levels deep:</strong> This mirrors Claude Code&#8217;s own nested sub-agent architecture (also shipping this week in v2.1.172) &#8212; a convergence across tools toward multi-agent pipeline depth as the default architectural pattern for complex coding tasks. <a href="https://github.com/anomalyco/opencode/releases">Read more &#8594;</a></p></li><li><p><strong>Cursor Bugbot is now 3x faster, finds 10% more bugs, and costs 22% less:</strong> The combination of speed, quality, and cost moving simultaneously in the same direction is unusual &#8212; Bugbot previously felt like a slow tax you paid for safety, and the <code>/review</code> command now makes it a natural pre-push habit. <a href="https://cursor.com/changelog">Read more &#8594;</a></p></li><li><p><strong>Kiro&#8217;s spec-driven workflow now works on GitLab and in the browser:</strong> Teams that live in GitLab can now plan work as reviewable requirements and task files in Kiro Web before the agent writes a line of code &#8212; extending a Spec-First approach that was previously limited to GitHub-connected repositories. <a href="https://releasebot.io/updates/kiro">Read more &#8594;</a></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!nfGF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!nfGF!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!nfGF!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!nfGF!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nfGF!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!nfGF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6732194,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/201951561?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!nfGF!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!nfGF!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!nfGF!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nfGF!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4f058eb-8f12-4957-91f4-25de60030741_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://devin.ai/blog/windsurf-is-now-devin-desktop/">Windsurf Is Now Devin Desktop</a></h3><p>Cognition retired the Windsurf brand on June 2 and shipped Devin Desktop as a standard over-the-air update. Three things materially changed: Cascade is replaced by Devin Local &#8212; a Rust rewrite that is <strong>up to 30% more token efficient</strong> and supports subagents; the Agent Command Center is now the default screen when you open the app instead of the code editor; and the product ships with open Agent Client Protocol support, letting Claude Agent, Codex, and OpenCode run alongside Devin in the same workspace. Existing Windsurf plans, pricing, and extensions carry over without changes. Cascade remains available until July 1, 2026.</p><h3><a href="https://techcrunch.com/2026/06/02/openai-launches-new-codex-tools-for-white-collar-work/">OpenAI Launches Codex for Every Role: Role-Specific Plugins and Sites</a></h3><p>OpenAI launched six role-specific Codex plugins (data analytics, creative production, sales, product design, equity investing, investment banking) alongside Codex Sites &#8212; the ability to create and host interactive web apps and dashboards that you can share via URL. <strong>Codex now has 5 million weekly active users, up 6x since the desktop app launched in February.</strong> Knowledge workers, who make up 20% of users but are growing 3x faster than developers, are the intended primary audience for this expansion. This marks a deliberate move from Codex as a coding tool to Codex as a cross-functional work platform.</p><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code v2.1.169: Safe Mode and /cd Command</a></h3><p>Claude Code ships two significant quality-of-life changes on June 8&#8211;9: <code>--safe-mode</code> disables all customizations while preserving authentication, giving you a clean diagnostic environment when a plugin or configuration breaks things; and <code>/cd</code> moves your session to a new working directory without breaking the prompt cache. The same release adds post-session hooks, fixes Up/Down arrow history navigation for long inputs, and tightens enterprise MCP policy enforcement on reconnect.</p><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code v2.1.172: Sub-Agents Spawn Sub-Agents</a></h3><p>The June 10 release enables sub-agents to spawn their own sub-agents <strong>up to 5 levels deep</strong>, unlocking recursive multi-agent pipeline architectures inside a single Claude Code session. Also in this release: a plugin marketplace search bar in <code>/plugin</code>, <code>model</code> attribute added to OTEL metrics for per-model cost attribution, and performance improvements that reduce idle CPU usage. Background agent directory isolation is also fixed, preventing cross-agent state leakage.</p><h3><a href="https://github.blog/changelog/2026-06-02-expanded-technical-preview-availability-for-the-github-copilot-app/">GitHub Copilot App Opens to All Subscribers with Canvases</a></h3><p>The GitHub Copilot app technical preview expands to all Pro, Pro+, Business, and Enterprise subscribers. The focal feature is <strong>Canvases</strong> &#8212; bidirectional work surfaces where agents write to actual work objects (plans, PRs, checklists) rather than chat messages, making progress visible and editable in real time. Voice input via on-device speech recognition is also included, and cloud-based agent sessions let you hand off from CLI into the app interface.</p><h3><a href="https://github.blog/changelog/2026-06-02-cloud-and-local-sandboxes-for-github-copilot-now-in-public-preview/">Cloud and Local Sandboxes for GitHub Copilot in Public Preview</a></h3><p>GitHub ships two sandbox types for Copilot agents: local sandboxes (built on Microsoft MXC technology, enabled via <code>/sandbox enable</code>) restrict Copilot&#8217;s access to filesystem, network, and system capabilities while running on your machine; cloud sandboxes (<code>copilot --cloud</code>) launch isolated GitHub-hosted Linux environments for stronger security boundaries. <strong>Both are included in standard Copilot seats</strong>, giving teams agentic workflow security without upgrading their plan.</p><h3><a href="https://github.blog/changelog/2026-06-02-copilot-cli-improved-ui-rubber-duck-prompt-scheduling-and-voice-input/">Copilot CLI: Rubber Duck Agent, Prompt Scheduling, and Voice Input</a></h3><p>The GitHub Copilot CLI gets four new capabilities at once: a Rubber Duck agent that critiques plans, designs, implementations, and tests before you ship; <code>/every</code> and <code>/after</code> slash commands for scheduling recurring or delayed prompts (e.g., <code>/every 30m run the frontend tests</code>); hold-space-bar voice dictation with on-device audio processing; and a redesigned experimental terminal UI with tabs for Sessions, Issues, PRs, and Gists. Rubber duck and voice are GA; scheduling and the new UI are via <code>/experimental</code>.</p><h3><a href="https://geminicli.com/docs/changelogs/latest/">Gemini CLI v0.45.0: Context Management Refactor and A2A Metadata</a></h3><p>The June 3 stable release completes a significant refactor of Gemini CLI&#8217;s context management system for improved architectural reliability, adds usage metadata exposure to the Agent-to-Agent (A2A) protocol for transparent resource monitoring across multi-agent flows, and resolves critical PTY resize errors and Termux relaunch loops that had been affecting terminal stability. This is likely one of the final major releases before Gemini CLI stops serving free and Google AI Pro/Ultra users on June 18.</p><h3><a href="https://github.com/anomalyco/opencode/releases">OpenCode v1.17.0: WSL Desktop and Faster File Search</a></h3><p>OpenCode v1.17.0 (June 10) adds WSL-backed Desktop support and Windows server management, making OpenCode a fully viable option for Windows-native development for the first time. File search is now faster using fff-backed tools, and a revamped sessions and servers UI improves navigation in long-running projects. The follow-on v1.17.4 (June 12) adds workspace-relative directory support for local MCP servers and connector-based authentication with provider credential storage.</p><h3><a href="https://releasebot.io/updates/kiro">Kiro Pro Max: $100/Month Tier and Self-Correcting Subagent Pipelines</a></h3><p>Kiro launches a Pro Max tier at $100/month offering 5,000 monthly credits &#8212; <strong>2.5x the Pro+ allocation</strong> &#8212; with premium model access and the full feature set. The same release ships a subagent self-correction pipeline: a reviewer stage can send tasks back to the implementer and loop until the work passes the bar, turning multi-agent pipelines into self-correcting workflows for use cases like code review and refactoring.</p><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://cursor.com/changelog/sdk-updates-jun-2026">Cursor SDK Updates: Custom Tools, Auto-Review Routing, and Nested Subagents</a></h3><p>Cursor&#8217;s June 4 SDK update adds significant programmability for teams building on the platform. <strong>Custom tools</strong> can be registered via <code>local.customTools</code> and exposed to the agent at runtime; <strong>auto-review routing</strong> lets teams define per-tool permission policies in <code>permissions.json</code> so approved tool calls run without interruption while higher-risk calls escalate; nested subagents are now supported; and JSONL and SQLite custom store options enable persistent agent memory across sessions. Worth reading for any team building internal tooling on top of Cursor&#8217;s agent runtime.</p><h3><a href="https://github.blog/changelog/2026-06-05-enterprise-managed-plugins-in-vs-code-in-public-preview/">Enterprise-Managed Plugins in VS Code: Distributing Agents at Scale</a></h3><p>Enterprises can now configure and auto-install Copilot CLI plugins across all licensed VS Code users via <code>settings.json</code>, removing manual setup from developer onboarding. Custom agents, shared MCP configurations, and enterprise-wide hook policies apply automatically &#8212; a meaningful governance lever for teams that have been managing plugin sprawl manually. The feature is in public preview and requires VS Code release 1.122 or later.</p><h3><a href="https://github.blog/changelog/2026-06-10-dedicated-security-review-command-now-available-in-copilot-cli/">GitHub Copilot CLI /security-review: On-Demand Security Scanning Before You Push</a></h3><p>The new <code>/security-review</code> slash command (experimental, public preview) analyzes local code changes and returns high-confidence findings scored by severity and confidence &#8212; covering injection flaws, XSS, insecure data handling, path traversal, and weak cryptography. It runs as an independent AI scan, designed to complement Dependabot and code scanning rather than replace them. The value proposition is speed: findings arrive without leaving the terminal, before code ever hits a PR.</p><div><hr></div><h2>&#128161; Others</h2><h3><a href="https://cursor.com/changelog/canvas-improvements">Cursor Design Mode: Select UI Elements, Talk to the Agent, Watch It Edit</a></h3><p>Cursor 3.7&#8217;s Design Mode inside Canvases lets users select and annotate UI elements directly in-browser &#8212; the agent sees the element, surrounding layout, and visual relationships, and can act on voice narration or typed annotations. It closes a workflow gap that has existed since canvas editing launched: the inability to communicate visually about specific components without describing them in text. The companion Context Usage Report breaks down exactly where your tokens go across system prompt, tool definitions, rules, and skills.</p><div><hr></div><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know that Claude Fable 5 compressed two months of Stripe engineering into a single day, that GitHub&#8217;s billing change made the agentic workflows your team relies on dramatically more expensive, and that sub-agents can now spawn sub-agents &#8212; five levels deep &#8212; across every major tool in your stack. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #34]]></title><description><![CDATA[More agentic coding launches in one week than any week this quarter &#8212; and the ROI questions are finally arriving.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-34</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-34</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 01 Jun 2026 05:16:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!57Oh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Race Heats Up:</strong> xAI enters agentic coding with Grok Build 0.1, Mistral rebrands entirely as Vibe with a unified work-and-code agent, and Codex brings Computer Use to Windows &#8212; more direct competition in the same week than any week this quarter.</p></li><li><p><strong>The Governance Shift:</strong> GitHub now lets enterprise admins assign specific AI models to specific organizations &#8212; model governance is no longer an afterthought, it&#8217;s a feature Copilot is actively selling to compliance-conscious CIOs.</p></li><li><p><strong>The Counter-Narrative:</strong> Engineering leaders are quietly capping per-engineer AI budgets as ROI questions surface &#8212; the era of unlimited token spend is hitting its first real headwinds.</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> Claude Opus 4.8 is the most honest model Anthropic has shipped. It&#8217;s approximately four times less likely than Opus 4.7 to let flawed code pass unremarked, scores 69.2% on SWE-bench Pro (up from 64.3%), and introduces Dynamic Workflows in research preview &#8212; the ability to orchestrate hundreds of parallel subagents in a single session to complete codebase-scale migrations across hundreds of thousands of lines of code. Fast mode is now three times cheaper than it was on Opus 4.7. The 41-day release cycle from 4.7 to 4.8 signals something: the competition right now is forcing releases faster than the calendar intended. <a href="https://www.anthropic.com/news/claude-opus-4-8">Read more &#8594;</a></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!57Oh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!57Oh!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!57Oh!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!57Oh!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!57Oh!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!57Oh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7253199,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/200011570?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!57Oh!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!57Oh!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!57Oh!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!57Oh!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a39c262-b8fc-47ec-a096-575081a17ae3_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><h2>Key Takeaways</h2><ul><li><p><strong>Cursor 3.6&#8217;s Auto-review Run Mode makes &#8220;run without babysitting&#8221; actually safe:</strong> A classifier subagent now evaluates every shell, MCP, and fetch call at runtime &#8212; allowlisted calls run immediately, sandboxable calls run in the sandbox, and everything else gets escalated to you. The result is fewer interruptions without the anxiety of full yolo mode. <a href="https://cursor.com/changelog/auto-review">Read more &#8594;</a></p></li><li><p><strong>xAI enters agentic coding with Grok Build 0.1 &#8212; 256K context, $1/$2 per 1M tokens, 100+ tokens/second:</strong> Unlike conversational Grok, this model is purpose-built for tool invocation and reasoning chains. It reads diagrams, UI mockups, and error screenshots, integrates with VS Code and JetBrains, and is already compatible with open-source agents like Hermes Agent and OpenClaw via OAuth. <a href="https://www.basenor.com/blogs/news/xai-launches-grok-build-0-1-agentic-coding-model-explained">Read more &#8594;</a></p></li><li><p><strong>Mistral unified its product under one brand called Vibe, and the AI assistant market is converging:</strong> Work Mode handles enterprise knowledge across Google Workspace, Slack, GitHub, and more; Code Mode runs remote coding agents and creates pull requests from a web surface or VS Code extension. Every Le Chat account carries over. The race is no longer between chat assistants &#8212; it&#8217;s between agent operating systems. <a href="https://mistral.ai/news/vibe-agent/">Read more &#8594;</a></p></li><li><p><strong>GitHub&#8217;s enterprise Copilot gets model governance at the org level:</strong> Admins can now assign specific models to specific organizations instead of applying a single enterprise-wide setting &#8212; a change that matters enormously for companies managing different compliance requirements across subsidiaries and teams. <a href="https://github.blog/changelog/2026-05-26-target-copilot-models-to-organizations-with-model-rules/">Read more &#8594;</a></p></li><li><p><strong>Engineering leaders are starting to cap AI token budgets:</strong> The Pragmatic Engineer reports that mid-sized and large companies are dampening AI agent spend through per-engineer monthly limits. After two years of uncritical adoption, ROI scrutiny is arriving &#8212; and teams that haven&#8217;t measured what they&#8217;re getting from their AI spend are about to be asked to justify it. <a href="https://newsletter.pragmaticengineer.com/p/the-pulse-a-trend-of-trying-to-cut">Read more &#8594;</a></p></li><li><p><strong>Karpathy said &#8220;agentic engineering&#8221; replaces vibe coding &#8212; but Jeff Gothelf noticed it&#8217;s just a new name for product management:</strong> Writing design specs, supervising agent plans, writing tests, managing permissions &#8212; that&#8217;s the PM job description. The judgment version of every role is now the job. The administrative version is automatable. Both engineers and PMs need to decide which side of that line they&#8217;re on. <a href="https://jeffgothelf.com/blog/karpathy-said-vibe-coding-is-obsolete-what-he-described-instead-is-product-management/">Read more &#8594;</a></p></li></ul><div><hr></div><p>Growing at 20% new subscribers per week.</p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://github.blog/changelog/2026-05-28-claude-opus-4-8-is-generally-available-for-github-copilot/">Claude Opus 4.8 Now Available in GitHub Copilot</a></h3><p>The same day Anthropic shipped Opus 4.8, it landed in GitHub Copilot for Pro+, Business, and Enterprise subscribers across VS Code, Visual Studio, JetBrains, Xcode, Eclipse, and GitHub&#8217;s web interface. Enterprise and Business admins must enable the policy in Copilot settings before their users can access it. <strong>A 15X premium request multiplier applies until Usage Based Billing launches June 1</strong> &#8212; worth noting for teams already managing Copilot token consumption.</p><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code v2.1.152&#8211;v2.1.158: Opus 4.8, Dynamic Workflows, and Auto-loading Plugins</a></h3><p>Five releases in four days. The headlining changes: Opus 4.8 is now the recommended model with <code>/effort xhigh</code> for difficult tasks (v2.1.154, May 28); Dynamic Workflows let you ask Claude to plan work and run tens to hundreds of agents in a single session (same release); <code>/code-review --fix</code> now applies review findings directly to the working tree (v2.1.152, May 27); <strong>plugins in </strong><code>.claude/skills/</code><strong> directories auto-load without needing marketplace publication</strong> (v2.1.157, May 29); <code>EnterWorktree</code> can switch between Claude-managed worktrees mid-session; and Auto mode arrives on Bedrock, Vertex, and Foundry for Opus 4.7 and 4.8 (v2.1.158, May 30).</p><h3><a href="https://github.blog/changelog/2026-05-26-copilot-memory-has-more-controls-for-deletion-scope-and-the-copilot-cli/">Copilot Memory Gets Deletion Controls, Repo-Level Toggle, and CLI Commands</a></h3><p>GitHub gave administrators and users more control over Copilot Memory: a new repository-level off switch lets admins disable memory entirely for a given repo from settings; deletion guidance now points users to the right place to remove a memory and downvotes entries; and three new CLI commands (<code>/memory on</code>, <code>/memory off</code>, <code>/memory show</code>) bring memory management into the terminal. Available to all paid Copilot plans in public preview.</p><h3><a href="https://developers.openai.com/codex/changelog">Codex CLI 0.134.0 and 0.135.0: History Search, Vim Mode, and Computer Use on Windows</a></h3><p>Two releases in rapid succession. <strong>0.134.0</strong> (May 26) adds local conversation history search with case-insensitive content matching, makes <code>--profile</code> the canonical selector across all CLI flows, and lets read-only MCP tools run concurrently via <code>readOnlyHint</code>. <strong>0.135.0</strong> (May 28) adds richer <code>codex doctor</code> diagnostics and named permission profiles. The May 29 app update brings the most notable change: <strong>Computer Use now works on Windows</strong> &#8212; Codex can see, click, and type in Windows desktop apps &#8212; plus remote control of Codex sessions from iOS, Android, or Mac.</p><h3><a href="https://geminicli.com/docs/changelogs/">Gemini CLI v0.44.0: Unified Auto Mode and Sublime/Emacs Support</a></h3><p>The May 27 release merges all specialized Auto modes into a single unified mode, simplifying configuration for teams running mixed workflows. <strong>Native support for Sublime Text and Emacs Client</strong> arrives as first-class integrations, and new <code>agent-tui</code> and <code>tui-tester</code> skills enable programmatic testing and automation of terminal UI applications &#8212; useful for anyone building or testing TUI-based developer tooling.</p><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://appwrite.io/blog/post/anthropic-just-launched-claude-opus-48-with-fast-mode-and-dynamic-workflows">Dynamic Workflows Explained &#8212; What They Are and How to Use Them</a></h3><p>Appwrite&#8217;s breakdown of Opus 4.8&#8217;s most consequential new feature: <strong>Dynamic Workflows</strong> let Claude plan and execute work across tens to hundreds of parallel subagents in a single session. The piece explains the practical use cases &#8212; codebase migrations across hundreds of thousands of lines, full-project refactors that would previously require manual task decomposition &#8212; and clarifies that this ships as a research preview available today in Claude Code. A useful orientation if you&#8217;re trying to understand when and why to reach for this versus a standard agentic flow.</p><div><hr></div><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know that Opus 4.8 can run hundreds of parallel subagents on your codebase in a single session, that engineering leaders are quietly starting to cap AI token spend, and that every AI assistant on the market is converging toward the same thing &#8212; an agent operating system &#8212; whether they admit it yet or not. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #33]]></title><description><![CDATA[The agent left the IDE. Your metrics didn't follow.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-33</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-33</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 25 May 2026 04:36:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4xzI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Big Story:</strong> Google I/O dropped Antigravity 2.0, Gemini 3.5 Flash, and Managed Agents &#8212; the most comprehensive agentic dev platform announcement since GitHub Copilot went GA.</p></li><li><p><strong>The Benchmark:</strong> SWE-bench Verified is now flagged as contaminated, and a new ranking study proves scaffolding quality matters as much as the model underneath.</p></li><li><p><strong>The Workflow Shift:</strong> Cursor embeds inside Jira, OpenAI puts Codex on your phone, and GitHub adds one-click CI fixes &#8212; the coding agent is no longer a separate tab, it lives inside every tool you already have open.</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> The Shai-Hulud/Megalodon supply chain attack is the most technically sophisticated assault on AI developer tooling to date. Wave 1 generated cryptographically valid SLSA Build Level 3 provenance attestations by hijacking legitimate build pipelines &#8212; meaning the signed artifacts were technically legitimate outputs from compromised pipelines. Wave 2 then deployed 5,718 malicious commits to 5,561 GitHub repositories in six hours. On May 12, the attackers open-sourced the worm itself, converting a targeted weapon into commodity attack infrastructure available to anyone. <a href="https://labs.cloudsecurityalliance.org/research/csa-research-note-shai-hulud-megalodon-supply-chain-cascade/">Read more &#8594;</a></p></blockquote><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!4xzI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!4xzI!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!4xzI!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!4xzI!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4xzI!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!4xzI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7259368,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/199146141?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!4xzI!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!4xzI!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!4xzI!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4xzI!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa36c05ad-aed0-42bf-99ef-10b2192f413b_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><h2>Key Takeaways</h2><ul><li><p><strong>The coding agent is leaving the IDE:</strong> Cursor is now inside Jira. Codex is now on your phone. GitHub Copilot fixes your failing CI jobs without you switching tabs. The agent isn&#8217;t a tool you open &#8212; it&#8217;s becoming part of every surface you already use. <a href="https://cursor.com/changelog/05-19-26">Read more &#8594;</a></p></li><li><p><strong>Cursor is becoming a model lab:</strong> Composer 2.5 matches Claude Opus 4.7 on SWE-Bench Multilingual (79.8% vs. 80.5%) at one-tenth the token cost &#8212; and it ships exclusively inside Cursor. The IDE that started as a thin wrapper around OpenAI is now training and shipping its own frontier-grade models. <a href="https://cursor.com/blog/composer-2-5">Read more &#8594;</a></p></li><li><p><strong>89% of engineering leaders say AI improved productivity &#8212; but 94% admit their metrics miss the actual costs:</strong> The Harness State of Engineering Excellence 2026 report finds 31% of developer time now goes to &#8220;invisible work&#8221; (reviewing AI code, fixing subtle AI bugs) that legacy productivity metrics don&#8217;t track. <a href="https://www.prnewswire.com/news-releases/harness-report-reveals-ai-has-outpaced-how-engineering-organizations-measure-developer-productivity-302770521.html">Read more &#8594;</a></p></li><li><p><strong>SWE-bench Verified is no longer a reliable benchmark:</strong> OpenAI found 59.4% of its tasks had flawed or unsolvable test cases, with contamination across all major frontier models. More importantly, the same model scored up to 17 points apart on identical benchmarks depending on the agent framework &#8212; scaffolding quality now matters as much as the model. <a href="https://www.marktechpost.com/2026/05/15/best-ai-agents-for-software-development-ranked-a-benchmark-driven-look-at-the-current-field/">Read more &#8594;</a></p></li><li><p><strong>Google&#8217;s I/O was the biggest agentic developer platform drop of the year:</strong> Antigravity 2.0 (desktop, CLI, SDK), Gemini 3.5 Flash (4x faster than Gemini 3.1 Pro at half the cost), and Managed Agents (one API call = full isolated Linux sandbox agent) all shipped the same week. Gemini CLI gets sunset for free users on June 18 &#8212; replaced by Antigravity CLI. <a href="https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2026-developer-highlights/">Read more &#8594;</a></p></li><li><p><strong>The MCP spec is getting its biggest overhaul since launch:</strong> The release candidate locked May 21 removes all session state, making MCP servers horizontally scalable behind plain load balancers. MCP Apps and a Tasks extension add UI rendering and long-running work as opt-in capabilities &#8212; the protocol is growing up for production. <a href="https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/">Read more &#8594;</a></p></li></ul><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://cursor.com/changelog">Cursor v3.4 + v3.5: Full-Screen Mode, Cloud Environments, and Automations</a></h3><p><strong>v3.4</strong> (May 13) introduced full-screen tab mode with a floating prompt bar, compact/balanced/detailed tool-call density controls, and multi-repo cloud development environments with Dockerfile-based configuration and <strong>70% faster builds</strong> via improved layer caching. <strong>v3.5</strong> (May 20) added Automations into the Agents Window, multi-repo reasoning for agents, no-repo automations for non-code tasks, and five new marketplace templates &#8212; with a 50% discount on agent runs for the first seven days of any new automation.</p><h3><a href="https://github.blog/changelog/2026-05-21-github-copilot-for-eclipse-is-open-source/">GitHub Copilot for Eclipse Goes Open Source Under MIT License</a></h3><p>GitHub published the full Eclipse plugin source code &#8212; including the implementation of chat, code completions, agent mode, Next Edit Suggestions, and MCP support &#8212; under the MIT license at github.com/microsoft/copilot-for-eclipse. The stated motivation is &#8220;community-driven innovation and increased transparency.&#8221; This does not open-source Copilot itself; it exposes one of its IDE front ends, including system prompts, context handling, and agentic workflow logic.</p><h3><a href="https://github.blog/changelog/month/05-2026/">GitHub Copilot Gets Auto Model Selection and One-Click CI Fixes</a></h3><p>Two notable Copilot updates landed in VS Code this week. <strong>Auto model selection</strong> (May 20) now routes tasks to the best model based on performance metrics, automatically matching model to task type across completions, chat, and agents. <strong>One-click CI fixes</strong> (May 18) let developers trigger the Copilot cloud agent directly from a failing GitHub Actions job to analyze the failure and propose a fix without leaving the Actions UI.</p><h3><a href="https://techcrunch.com/2026/05/14/openai-says-codex-is-coming-to-your-phone/">OpenAI Codex Comes to iPhone and Android</a></h3><p>Codex is now available in preview inside the ChatGPT mobile app on iOS and Android, across all plan tiers including Free. The mobile interface is a remote control surface for Codex sessions running on a Mac or devbox: developers can review outputs, approve commands, switch models, monitor terminal output, diffs, and test results from anywhere. More than <strong>4 million people now use Codex weekly.</strong> Windows desktop support is coming; no date announced.</p><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code v2.1.141&#8211;v2.1.149: /code-review, Agent Flags, and Background Session Hardening</a></h3><p>Eight releases in ten days. The most impactful changes: <code>/simplify</code> is now <code>/code-review</code> with an optional effort level parameter; the <code>claude agents</code> command accepts new flags to configure dispatched background sessions (<code>--add-dir</code>, <code>--settings</code>, <code>--mcp-config</code>, <code>--plugin-dir</code>, <code>--permission-mode</code>, <code>--model</code>, <code>--effort</code>); Fast Mode now defaults to Opus 4.7; <code>/usage</code> shows per-category cost breakdowns (skills, subagents, plugins, MCP); and background sessions now persist through macOS sleep/wake cycles with improved startup times (15s vs. 75s).</p><h3><a href="https://geminicli.com/docs/changelogs/">Gemini CLI v0.43.0: Surgical Edits and Session Portability</a></h3><p>The May 22 stable release improves code editing precision (the model now defaults to the <code>edit</code> tool for targeted changes), introduces session portability (export active sessions to files and import later via a CLI flag), and ships an adaptive token calculator for smarter context window management. <strong>Important note:</strong> Google has announced Gemini CLI will be replaced by Antigravity CLI for unpaid-tier and Google One users on June 18th.</p><h3><a href="https://windsurf.com/changelog">Windsurf: Claude Opus 4.7 Fast Mode Available</a></h3><p>Windsurf added Claude Opus 4.7 in fast mode on May 12, delivering <strong>~2.5x higher output speeds</strong> compared to standard Opus 4.7. The May 17 release (v2.3.9) fixed availability issues with the swe-check model, enhanced terminal processing performance, restored conversation sharing, and repaired Devin Local agent path resolution on WSL.</p><h3><a href="https://zed.dev/releases/stable">Zed 1.3.5: Terminal Threads, Parallel Agents, and Gemini 3.5 Flash</a></h3><p>Zed&#8217;s latest stable release adds Terminal Threads from the sidebar and Agent Panel, Git panel branch history views, inline image and Mermaid diagram rendering inside agents, and a new <code>subagent_model</code> setting enabling parallel agent execution across different parts of a codebase. The follow-up release (1.3.6, May 21) added native Google AI support, including Gemini 3.5 Flash with configurable thinking levels.</p><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://github.blog/changelog/2026-05-19-gemini-3-5-flash-is-generally-available-for-github-copilot/">How to Use Gemini 3.5 Flash in GitHub Copilot &#8212; Availability, Pricing, and Setup</a></h3><p>Gemini 3.5 Flash is now rolling out to GitHub Copilot Pro, Pro+, Business, and Enterprise subscribers across VS Code (1.115.0+), Visual Studio, JetBrains, Xcode, and Eclipse. <strong>Enterprise and Business admins must explicitly enable the Gemini 3.5 Flash policy</strong> in Copilot settings before users can access it. The model carries a 14x premium request multiplier (tentative). The rollout is gradual, so availability may vary &#8212; if you&#8217;re not seeing it yet, it&#8217;s coming.</p><div><hr></div><h2>&#128161; Others</h2><h3><a href="https://github.blog/changelog/month/05-2026/">GitHub Copilot Now Finds Issues with Natural Language &#8212; Semantic Search Goes GA</a></h3><p>Semantic issue search launched in Copilot Chat on GitHub.com, allowing developers to find, group, and analyze repository issues using natural language queries powered by a new semantic issues index. The feature is generally available on all Copilot plans. It&#8217;s a small but meaningful shift: instead of writing precise filter queries, you describe what you&#8217;re looking for and Copilot understands intent and context &#8212; the same pattern that is slowly displacing structured query interfaces across every developer workflow.</p><div><hr></div><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know that Cursor launched a proprietary model, that Google&#8217;s Antigravity 2.0 just made &#8220;managed agent&#8221; a one-API-call concept, and that the supply chain attack hitting AI dev tools right now is more sophisticated than anything we&#8217;ve seen before &#8212; and you didn&#8217;t have to spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #32]]></title><description><![CDATA[Anthropic held a developer conference, doubled your rate limits, and signed SpaceX to power it. The rest of the week kept up.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-32</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-32</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 11 May 2026 06:58:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_IZp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Agent Story:</strong> Anthropic shipped three new capabilities for Claude Managed Agents &#8212; Dreaming, Outcomes, and Multiagent Orchestration &#8212; turning a single-agent tool into a self-improving, parallelizable fleet. Netflix is already running orchestration in production.</p></li><li><p><strong>The Tool Story:</strong> OpenAI&#8217;s Codex moved into the browser with a Chrome extension that can test web apps, access DevTools, and use authenticated sessions for Gmail and internal tools &#8212; bringing Codex to where most real work actually lives.</p></li><li><p><strong>The Insight:</strong> GitHub&#8217;s engineering team published the data on how they cut agent token costs by 19&#8211;62% across five production workflows. The key finding was structural, not algorithmic: most agent turns are deterministic data-gathering steps that never needed an LLM in the first place.</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> On May 6, Anthropic held Code with Claude 2026 in San Francisco &#8212; its first developer conference &#8212; and rather than announcing new models, made its existing products dramatically more capable. Claude Code&#8217;s five-hour rate limits were doubled for all paid plans. Anthropic signed an agreement to use all compute capacity at SpaceX&#8217;s Colossus 1 data center: 300+ megawatts and 220,000+ NVIDIA GPUs available within a month. API volume is up 17x year-on-year. The conference framed the product direction clearly: Claude Code is no longer a terminal tool &#8212; it&#8217;s a platform, with CLI, IDE, desktop app, Remote Agents, Code Review, and now Routines as surfaces. <a href="https://www.anthropic.com/news/higher-limits-spacex">Read more &#8594;</a></p></blockquote><div><hr></div><p><strong>Growing at 20% new subscribers per week.</strong></p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: <strong>how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_IZp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_IZp!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!_IZp!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!_IZp!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_IZp!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_IZp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7723009,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/197182034?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!_IZp!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!_IZp!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!_IZp!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_IZp!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb4c43ab-cb09-4e57-a9e5-f93e8d3c98f3_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h2>Key Takeaways</h2><ul><li><p><strong>Claude Managed Agents can now self-improve between sessions:</strong> The Dreaming feature (research preview) reviews past agent sessions to extract recurring patterns, mistakes, and shared preferences across a team &#8212; updating memory stores automatically or with developer review. Outcomes adds rubric-based evaluation with up to 10% higher task success rates on the hardest tasks. Multiagent Orchestration lets a lead agent delegate subtasks to parallel specialists with their own models and tools, running on shared storage. <a href="https://9to5mac.com/2026/05/07/anthropic-updates-claude-managed-agents-with-three-new-features/">Read more &#8594;</a></p></li><li><p><strong>Cursor 3.3 closes the loop between coding and reviewing:</strong> The May 7 release ships a full PR review experience inside the IDE &#8212; Reviews, Commits, and Changes tabs &#8212; alongside &#8220;Build in Parallel,&#8221; which identifies independent parts of a plan and runs them simultaneously using async subagents. Teams can now go from initial plan to parallel execution to PR review without leaving Cursor. <a href="https://cursor.com/changelog/05-07-26">Read more &#8594;</a></p></li><li><p><strong>The cheapest token is the one you never send:</strong> GitHub&#8217;s May 7 analysis of five production agentic workflows found 19&#8211;62% token reductions came not from better prompting, but from removing LLM calls entirely for steps that didn&#8217;t need reasoning. Pruning unused MCP tools saved 8&#8211;12 KB of schema context per call; replacing GitHub MCP calls with direct CLI commands eliminated whole agent turns. <a href="https://github.blog/ai-and-ml/github-copilot/improving-token-efficiency-in-github-agentic-workflows/">Read more &#8594;</a></p></li><li><p><strong>GPT-5.5 Instant cuts hallucinations 52.5% and replaces GPT-5.3 as ChatGPT&#8217;s default:</strong> OpenAI&#8217;s May 5 release produces fewer hallucinations on high-stakes topics (law, medicine, finance), uses 30% fewer words, and scores 81.2 vs. 65.4 on AIME 2025 math benchmarks. Developers get it as <code>chat-latest</code> in the API, with GPT-5.3 remaining available for three months on paid plans. <a href="https://techcrunch.com/2026/05/05/openai-releases-gpt-5-5-instant-a-new-default-model-for-chatgpt/">Read more &#8594;</a></p></li><li><p><strong>Cursor Enterprise gets model-level spend controls before the GitHub Copilot billing shift:</strong> Four weeks before GitHub&#8217;s June 1 move to token-based billing, Cursor&#8217;s May 4 update gives enterprise admins granular model and provider blocklists, soft spending limits with 50/80/100% alerts, and per-surface usage breakdowns covering Cloud Agents, Automations, and Security Review. The timing is not coincidental &#8212; teams that didn&#8217;t track AI spend last quarter are now being told they need to. <a href="https://cursor.com/changelog/05-04-26">Read more &#8594;</a></p></li><li><p><strong>Security review is now table stakes for AI coding tools:</strong> Windsurf made Devin Review and Quick Review (10x faster bug detection via SWE-check) available to all subscribers; Snyk integrated Claude models into its AI Security Platform; and Opsera embedded DevSecOps agents &#8212; Architecture Analyzer, Security/SQL Scanner, Compliance Auditor &#8212; directly into Cursor. Three independent moves in the same week signal that security review is converging toward a default layer of the agentic coding stack, not an optional add-on. <a href="https://sdtimes.com/ai/may-8-2026-ai-updates-from-the-past-week-coder-agents-launch-snyk-claude-partnership-opsera-cursor-partnership-and-more/">Read more &#8594;</a></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://geminicli.com/docs/changelogs/">Gemini CLI v0.41.0 &#8212; Real-Time Voice Mode and Gemma 4 Support</a></h3><p>The May 5 stable release ships the real-time voice mode with cloud and local backends that had been in preview &#8212; the most significant UI expansion since the CLI launched. Workspace trust enforcement now secures <code>.env</code> loading in headless mode, and shell command validation is strengthened with a core tools allowlist. <strong>Gemma 4 models are now supported via the Gemini API</strong>, with the v0.42.0 preview release enabling them by default.</p><h3><a href="https://windsurf.com/changelog">Windsurf 2.2.17 &#8212; Devin Review and Quick Review for All Subscribers</a></h3><p>Devin Review &#8212; in-editor code review with bug detection and full codebase context &#8212; and Quick Review &#8212; <strong>10x faster bug detection powered by SWE-check</strong> &#8212; are now available to all Windsurf subscribers without a separate Cognition account. The release also improves the agent inbox with a list view, better session sidebar sorting, and fixes MCP server reliability and Devin Local agent stability.</p><h3><a href="https://developers.openai.com/codex/changelog">Codex CLI 0.130.0 &#8212; Remote Control, Plugin Sharing, Bedrock Auth</a></h3><p>The May 8 release adds <code>codex remote-control</code> as a simpler entrypoint for starting a headless, remotely controllable app-server &#8212; useful for CI/CD pipelines and background agent runners. Plugin details now expose bundled hooks with sharing and discoverability controls. Bedrock auth gains support for AWS console-login credentials, and <code>view_image</code> resolves files through the selected environment for multi-environment sessions.</p><h3><a href="https://opencode.ai/changelog">OpenCode v1.14.42 &#8212; Scout Agent for Repo Research</a></h3><p>The May 9 release introduces the <strong>Scout agent</strong> &#8212; a built-in tool for repository research, documentation lookup, and dependency-source inspection &#8212; giving agents a structured way to gather codebase context before making changes. Also adds HTTP API response compression for large non-streaming responses and an interactive split-footer mode for <code>opencode run</code>.</p><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code v2.1.136 &#8212; MCP Stability, Auto Mode Rules, WSL2 Image Paste</a></h3><p>The May 8 release brings 50+ fixes to Claude Code&#8217;s production deployment scenarios: <code>settings.autoMode.hard_deny</code> for classifier-based rules in auto mode; fixed MCP servers disappearing after <code>/clear</code>; fixed OAuth refresh token race conditions under concurrent server refreshes; and WSL2 image paste from Windows clipboard via PowerShell fallback. The week&#8217;s release cadence (8 releases from May 4&#8211;9) reflects the scale of active enterprise deployment across diverse environments.</p><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://simonwillison.net/2026/May/6/code-w-claude-2026/">Simon Willison: Live Blog of Code with Claude 2026</a></h3><p>Simon Willison&#8217;s real-time notes from Anthropic&#8217;s May 6 San Francisco developer conference capture every announcement as it happened &#8212; including the live coding demo with Boris Cherny and Jarred Sumner. The best single-source summary of what Claude Code&#8217;s new surfaces (Code Review, Remote Agents, Routines, desktop app) look like in practice, and what &#8220;Opus 4.7 has a real taste for visual design&#8221; actually means when demonstrated live.</p><div><hr></div><h2>&#128161; Others</h2><h3><a href="https://github.blog/changelog/2026-05-06-github-copilot-in-visual-studio-code-april-releases/">GitHub Copilot in Visual Studio Code &#8212; April Releases: Semantic Search, Inline Diffs, and Bring Your Own Keys</a></h3><p>The May 6 changelog covering VS Code releases v1.116&#8211;v1.119 introduces semantic indexing across all workspaces, meaning meaningful code search and grep-style queries across GitHub repos now work without per-project setup. Code changes from agents now appear as <strong>inline diffs in the chat thread</strong>, keeping review in context. Teams on Copilot Business and Enterprise can connect their own API keys from Anthropic, OpenAI, or Google &#8212; a notable shift that decouples Copilot from GitHub&#8217;s model choices for teams with existing provider commitments.</p><h3><a href="https://www.lennysnewsletter.com/p/code-with-claude-the-5-biggest-updates">Lenny&#8217;s Newsletter: Code with Claude 2026 &#8212; The 5 Biggest Updates Explained</a></h3><p>Claire Vo&#8217;s walkthrough of the five most significant announcements from Anthropic&#8217;s May 6 developer conference provides the clearest explanation of what Routines, Outcomes, and Dreaming mean for builders shipping AI products today. The episode reframes Routines as &#8220;higher-order prompts&#8221; &#8212; infrastructure for async developer workflows, not just scheduled scripts &#8212; and explains how Outcomes shifts agent evaluation from &#8220;did it run?&#8221; to &#8220;did it succeed against a rubric I defined?&#8221; Essential context for anyone building on top of Claude Managed Agents.</p><div><hr></div><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You know that Anthropic just doubled your Claude Code rate limits and signed SpaceX&#8217;s entire Colossus data center to back it up &#8212; and what that means for your team&#8217;s roadmap. You know that the cheapest token is the one you never send, and exactly which structural changes GitHub used to cut agent costs by up to 62%. You didn&#8217;t spend your weekend reading to know this. I did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #31]]></title><description><![CDATA[Karpathy called the paradigm shift. Four tools shipped to prove him right in the same week.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-31</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-31</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 04 May 2026 05:02:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BEcI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Infrastructure:</strong> Cursor released a TypeScript SDK that turns its agent runtime into programmable infrastructure &#8212; CI/CD pipelines and backend services can now invoke the full Cursor agent harness, with sandboxed cloud VMs that keep running after your laptop closes.</p></li><li><p><strong>The Model:</strong> Mistral shipped Medium 3.5 &#8212; a 128B dense model that merges reasoning, instruction-following, and code into one set of weights &#8212; alongside async cloud coding agents that run in parallel without blocking developers.</p></li><li><p><strong>The Shift:</strong> Lovable launched on iOS and Android with 100K+ Play Store downloads in its first days &#8212; and on iOS, Apple&#8217;s enforcement rules forced generated apps into the browser rather than running natively. Vibe coding just hit its first platform wall.</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> On April 30, Andrej Karpathy published his notes from Sequoia&#8217;s AI Ascent 2026, where he declared the era of vibe coding over and &#8220;agentic engineering&#8221; officially begun. He frames Software 3.0 as a paradigm where humans program LLMs through prompts, context, tools, and memory &#8212; and the context window is the program. The key inflection he identified: December 2025, when he went from writing 80% of his code manually to delegating 80% to agents. His sharpest insight isn&#8217;t about tools at all &#8212; it&#8217;s about taste: LLMs automate what you can verify; human judgment remains irreplaceable for deciding what&#8217;s worth building in the first place. <a href="https://karpathy.bearblog.dev/sequoia-ascent-2026/">Read more &#8594;</a></p></blockquote><div><hr></div><p><strong>Growing at 20% new subscribers per week.</strong></p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><h2>Key Takeaways</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!BEcI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!BEcI!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!BEcI!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!BEcI!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BEcI!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!BEcI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7899353,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/196299798?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!BEcI!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!BEcI!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!BEcI!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BEcI!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97c5a64a-a0a2-4073-928b-8088044b96c0_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><ul><li><p><strong>Cursor&#8217;s SDK makes its agent runtime a platform, not just an editor:</strong> The April 29 TypeScript SDK release gives developers programmatic access to the same harness powering Cursor&#8217;s desktop app, CLI, and web interface &#8212; including codebase indexing, MCP servers, hooks, and subagent delegation. Agents run on dedicated sandboxed VMs; early production adopters include Rippling, Notion, and Faire. The most common use case in early deployments: CI/CD pipelines that trigger agents to summarize changes, identify root causes for test failures, and open PRs autonomously. <a href="https://cursor.com/blog/typescript-sdk">Read more &#8594;</a></p></li><li><p><strong>Mistral Medium 3.5 sets a new bar for open-weight coding:</strong> Released April 29, the 128B dense model scores <strong>77.6% on SWE-Bench Verified</strong> &#8212; competitive with Claude Sonnet 4.6&#8217;s 79.6% &#8212; and runs self-hosted on as few as four GPUs at $1.50/M input tokens. The <code>reasoning_effort</code> toggle lets agents dial in heavier computation only when needed. Paired with Vibe cloud coding agents, teams can launch async sessions from the CLI, run jobs in parallel, and receive finished PRs as output &#8212; integrated with GitHub, Linear, Jira, and Slack. <a href="https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5">Read more &#8594;</a></p></li><li><p><strong>Lovable&#8217;s mobile launch is a leading indicator, not just a product news story:</strong> The iOS and Android app shipped April 28 with 100K+ Play Store downloads on day one &#8212; and on iOS, generated apps preview in browsers rather than run natively, a direct consequence of Apple&#8217;s enforcement against vibe-coding platforms. This constraint reveals the emerging regulatory boundary: Apple is actively shaping what AI-assisted app creation is allowed to look like on its platform, and every vibe-coding tool is now navigating that boundary. <a href="https://techcrunch.com/2026/04/28/lovable-launches-its-vibe-coding-app-on-ios-and-android/">Read more &#8594;</a></p></li><li><p><strong>Claude Code v2.1.126 closes the friction that limits real-world adoption:</strong> The May 1 release introduced <code>claude project purge</code> for full state cleanup, terminal-pasted OAuth for WSL2 and SSH environments where browser callbacks fail, and fixed CJK text rendering on Windows. These are not headline features &#8212; they are exactly the class of fixes that separate a tool developers trust for production work from one they use carefully in controlled conditions. <a href="https://github.com/anthropics/claude-code/releases">Read more &#8594;</a></p></li><li><p><strong>Cursor Security Review beta brings always-on security agents to Teams and Enterprise:</strong> The April 30 beta launch introduces two new agent types: a Security Reviewer that checks every PR for vulnerabilities, auth regressions, prompt injection risks, and agent tool auto-approvals (inline comments at the exact diff location); and a Vulnerability Scanner that runs scheduled codebase sweeps and posts findings to Slack. Both agents integrate with existing SAST, SCA, and secrets scanning tools via MCP. The timing is notable: security agents arriving in the same week that Cursor opened its runtime to external developers via the SDK. <a href="https://cursor.com/changelog/04-30-26">Read more &#8594;</a></p></li><li><p><strong>GitHub Copilot in Visual Studio now has a Debugger Agent that validates fixes against live runtime:</strong> The April 30 Visual Studio update goes beyond agentic code generation to introduce a Debugger Agent that reproduces bugs autonomously, forms hypotheses, instruments the code, and validates the fix through actual execution &#8212; not just static analysis. Cloud agent sessions also launch directly from the IDE, offloading remote work while developers continue locally. User-level custom agent definitions now travel across projects without repository commits. <a href="https://github.blog/changelog/2026-04-30-github-copilot-in-visual-studio-april-update/">Read more &#8594;</a></p></li><li><p><strong>GitHub Copilot&#8217;s shift to token-based billing signals the end of flat-rate AI tooling:</strong> Effective June 1, 2026, GitHub is replacing Premium Request Units with <strong>GitHub AI Credits</strong> ($0.01 each) &#8212; pricing based on token consumption rather than request count. Base prices stay the same (Pro at $10/month, Pro+ at $39/month), but the economics change fundamentally: Copilot Chat, CLI, cloud agents, and code review now draw from a monthly credit pool, with no rollover and no fallback to cheaper models when credits run out. GitHub&#8217;s CPO framed the core issue plainly: <em>&#8220;Today, a quick chat question and a multi-hour autonomous coding session can cost the user the same amount.&#8221;</em> Developer reaction was blunt &#8212; 707 downvotes vs. 15 upvotes in the community thread &#8212; with the sharpest concern being that agentic sessions using Opus 4.7 (now at a <strong>27x token multiplier</strong>) can burn through a monthly plan in a single afternoon. The broader implication: AI coding tools are converging toward cloud-style metered infrastructure, and developers will need to manage AI spend the same way they manage AWS bills. <a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/">Read more &#8594;</a></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://windsurf.com/changelog">Windsurf 2.1.29 &#8212; Devin for Terminal and Devin Local Agent</a></h3><p>April 28&#8217;s release brought Devin to the command line for all Windsurf subscribers &#8212; no separate Cognition account required. The terminal agent runs locally with full codebase access and can hand off to Devin&#8217;s cloud VM (with its own machine, test runner, and video recording) mid-session. A new Devin Local variant in the editor uses the same agent harness and is claimed to be <strong>30% more token-efficient</strong> than Cascade. Multi-model support includes Opus 4.7, GPT-5.5, and SWE-1.6, positioning it as a drop-in harness rather than a single-model tool.</p><h3><a href="https://cursor.com/changelog">Cursor Team Marketplace Now Requires No Repository to Set Up</a></h3><p>The May 1 update lets admins create team marketplaces without first connecting a repository, and adds direct management of first-party plugin install behavior &#8212; Default Off, Default On, or Required. Small infrastructure change, but it removes the last organizational barrier to teams standardizing on a shared plugin stack without a code review process.</p><h3><a href="https://opencode.ai/changelog">OpenCode v1.14.26&#8211;1.14.33 &#8212; Zed Support, Mistral Medium 3.5, Session Recovery</a></h3><p>OpenCode shipped seven releases in the April 26&#8211;May 2 window. Highlights: Zed editor selection support for editor context (Apr 26), configurable default shell for agent commands (Apr 27), Mistral Medium 3.5 integration with reasoning (Apr 29), desktop session recovery for path mismatch issues (Apr 29), Azure resource configuration persistence (May 1), and a fix for custom agents in plugins not loading (May 2). The project crossed <strong>150K GitHub stars</strong> and is used by over 6.5 million developers monthly.</p><h3><a href="https://github.com/anthropics/claude-code/releases">Claude Code v2.1.122 &#8212; Bedrock Service Tier Selection and PR Resume by URL</a></h3><p>The April 28 release adds the <code>ANTHROPIC_BEDROCK_SERVICE_TIER</code> environment variable (<code>default</code>, <code>flex</code>, or <code>priority</code>) for teams running Claude Code on AWS Bedrock who need predictable throughput. The <code>/resume</code> command can now locate sessions by PR URL &#8212; across GitHub, GitHub Enterprise, GitLab, and Bitbucket &#8212; rather than requiring a session ID lookup.</p><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://geminicli.com/docs/changelogs/latest/">Gemini CLI v0.40.0 &#8212; Offline Search, Four-Tier Memory, and Voice Mode Preview</a></h3><p>The April 28 stable release bundles ripgrep binaries into the single executable so <code>gemini</code> can search codebases without internet access &#8212; a critical improvement for air-gapped environments and CI runners. The memory system was fully overhauled: the legacy MemoryManagerAgent is replaced by a <strong>prompt-driven four-tier system</strong> that handles context more predictably across long sessions. A new <code>gemini gemma</code> command simplifies local Gemma model setup, and GitHub-style colorblind themes ship by default. The concurrent preview release (v0.41.0-preview.0) adds real-time voice mode with cloud and local backends &#8212; the first signal of where the Gemini CLI UI is heading.</p><div><hr></div><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You know that Cursor&#8217;s SDK just made its agent runtime programmable infrastructure &#8212; and that Mistral&#8217;s async cloud agents can run your refactors in parallel while you sleep. You know that GitHub Copilot is moving to token-based billing on June 1 and exactly what that means for your team&#8217;s spend before anyone else asks you. You didn&#8217;t spend your weekend reading to know this. I did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #30]]></title><description><![CDATA[GPT-5.5 doubled its API price on launch day. DeepSeek undercut it by 7x the same week. The model pricing war just stopped being theoretical.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-30</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-30</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 27 Apr 2026 04:37:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!afQI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Price War:</strong> OpenAI doubled GPT-5.5&#8217;s API price on launch day &#8212; $5/$30 per million tokens &#8212; while DeepSeek shipped V4-Pro at $1.74/$3.48, roughly one-seventh the cost. The frontier model market just bifurcated into two very different bets.</p></li><li><p><strong>The Money:</strong> In four days, Anthropic locked in $5 billion from Amazon and up to $40 billion from Google &#8212; two massive compute commitments from cloud giants that are also rivals, signaling the infrastructure race under AI dev tools is accelerating faster than the tools themselves.</p></li><li><p><strong>The Security Pivot:</strong> Replit shipped two AI-native security features in two days &#8212; an autonomous vulnerability auditor and a CVE auto-patcher &#8212; marking a shift from &#8220;AI helps you build&#8221; to &#8220;AI maintains what you built.&#8221;</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> GPT-5.5 launched on April 23 with a doubled API price and a claim of &#8220;a new class of intelligence&#8221; &#8212; but the benchmarks tell a more nuanced story. On Terminal-Bench 2.0 (autonomous agent tasks), GPT-5.5 scores 82.7% and leads the field. On SWE-Bench Pro (resolving real GitHub issues on production codebases), Opus 4.7 leads at 64.3% vs. GPT-5.5&#8217;s 58.6%. DeepSeek V4-Pro, released the same day at one-seventh the price, sits at 55.4% on SWE-bench &#8212; close enough to matter if you&#8217;re building at scale. The real story is not who won a benchmark. It&#8217;s that OpenAI just defined the high-price ceiling while DeepSeek defined the floor, and every developer team now has to make an explicit choice between them. <a href="https://the-decoder.com/openai-unveils-gpt-5-5-claims-a-new-class-of-intelligence-at-double-the-api-price/">Read more &#8594;</a></p></blockquote><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!afQI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!afQI!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png 424w, /__u/substackcdn.com/image/fetch/$s_!afQI!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png 848w, /__u/substackcdn.com/image/fetch/$s_!afQI!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png 1272w, /__u/substackcdn.com/image/fetch/$s_!afQI!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!afQI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png" width="1456" height="758" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:758,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5671659,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/195589954?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!afQI!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png 424w, /__u/substackcdn.com/image/fetch/$s_!afQI!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png 848w, /__u/substackcdn.com/image/fetch/$s_!afQI!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png 1272w, /__u/substackcdn.com/image/fetch/$s_!afQI!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6f8b36e-6461-4dde-98d5-8901a72e8447_2752x1432.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Key Takeaways</strong></h2><ul><li><p><strong>GitHub&#8217;s subscription model is cracking under agentic load:</strong> GitHub paused new Copilot Pro, Pro+, and Copilot Business signups after agentic workflows &#8212; parallelized, long-running sessions &#8212; started regularly incurring costs that exceed the plan price. Opus models were removed from Pro plans (now only in Pro+), and a move toward usage-based billing is being signaled. This is the first public admission that flat-rate AI coding subscriptions were designed for suggestions, not agents. <a href="https://github.blog/news-insights/company-news/changes-to-github-copilot-individual-plans/">Read more &#8594;</a></p></li><li><p><strong>DeepSeek V4 proves open-source can match closed-source on coding &#8212; at 1/7th the price:</strong> Released April 24 as a preview, DeepSeek V4-Pro (1.6T params, 49B active) matches GPT-5.4 on coding competition benchmarks, supports 1 million token context, and integrates with Claude Code, OpenCode, and OpenClaw. The architecture uses 27% of the compute of its predecessor for the same context length. MIT Technology Review&#8217;s analysis highlights a structural implication: V4 marks DeepSeek&#8217;s first model optimized for Huawei Ascend chips, advancing China&#8217;s push toward domestic AI infrastructure independence. <a href="https://techcrunch.com/2026/04/24/deepseek-previews-new-ai-model-that-closes-the-gap-with-frontier-models/">Read more &#8594;</a></p></li><li><p><strong>Anthropic secured $145 billion in infrastructure commitments in four days:</strong> Amazon invested $5 billion (April 20) and Google committed up to $40 billion (April 24), both structured around securing compute &#8212; AWS and Google Cloud capacity respectively, each at 5 GW. Anthropic&#8217;s annualized revenue hit $30 billion this month, up from $9 billion at end of 2025. The speed and scale of these two deals in the same week signals that the competition for AI model infrastructure is now operating on a timeline measured in days, not quarters. <a href="https://techcrunch.com/2026/04/24/google-to-invest-up-to-40b-in-anthropic-in-cash-and-compute/">Read more &#8594;</a></p></li><li><p><strong>Claude Code shipped three releases in 48 hours, each fixing real friction:</strong> v2.1.117 corrected a critical context window miscalculation for Opus 4.7 (was computing against 200K instead of 1M tokens), v2.1.118 added Vim visual mode and merged duplicate commands into <code>/usage</code>, and v2.1.119 made <code>/config</code> settings persist and extended <code>--from-pr</code> to GitLab, Bitbucket, and GitHub Enterprise. These aren&#8217;t headline features &#8212; they&#8217;re the kind of reliability fixes that signal a tool moving from early adopter to production-grade. <a href="https://github.com/anthropics/claude-code/releases">Read more &#8594;</a></p></li><li><p><strong>Agentic billing will replace flat subscriptions &#8212; and GitHub is the first domino:</strong> The GitHub Copilot pause is not an operational hiccup. It is the first major platform publicly acknowledging that the unit economics of agentic AI don&#8217;t fit inside a flat monthly fee. Expect token-based or outcome-based pricing to become the default across AI developer tools within the next 12 months. The developers building pipelines today are the ones who will be most directly affected when their vendors reprice. <a href="https://thenewstack.io/github-copilot-signups-paused/">Read more &#8594;</a></p></li></ul><div><hr></div><p>Growing at 20% new subscribers per week.</p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2><strong>&#128230; Releases &amp; News</strong></h2><h3><strong><a href="https://techcrunch.com/2026/04/24/deepseek-previews-new-ai-model-that-closes-the-gap-with-frontier-models/">DeepSeek V4 Preview: Near Frontier Performance at 1/7th the Cost</a></strong></h3><p>DeepSeek released V4-Pro and V4-Flash on April 24 &#8212; both supporting <strong>1 million token context windows</strong> and built on a mixture-of-experts architecture that runs at 27% of the compute of its predecessor for the same context length. V4-Pro (1.6T total params, 49B active) costs <strong>$1.74 input / $3.48 output per million tokens</strong>, compared to $5/$30 for GPT-5.5 and $5/$25 for Claude Opus 4.7. Both models integrate with Claude Code, OpenCode, and OpenClaw out of the box. Text-only for now &#8212; audio, video, and image capabilities are absent &#8212; and knowledge benchmarks trail leading models by roughly three to six months.</p><h3><strong><a href="https://geminicli.com/docs/changelogs/latest/">Gemini CLI v0.39.0 Released</a></strong></h3><p>Released April 23, Google&#8217;s open-source terminal agent shipped a <code>/memory</code> command that lets you review and patch skills extracted during previous agent sessions &#8212; essentially a continuous learning loop baked into the CLI. Plan Mode now requires user confirmation before activating any skill, adding a meaningful transparency checkpoint. The release also decouples the ContextManager and Sidecar architecture for better session resilience, the kind of infrastructure improvement that matters once you start running agents across multiple long-lived projects.</p><h3><strong><a href="https://www.anthropic.com/news/anthropic-nec">Anthropic and NEC Collaborate to Build Japan&#8217;s Largest AI Engineering Workforce</a></strong></h3><p>NEC Corporation becomes Anthropic&#8217;s first Japan-based global partner, deploying Claude to approximately 30,000 employees and building one of Japan&#8217;s largest AI-native engineering organizations using Claude Code. A new Center of Excellence will develop domain-specific AI products for finance, manufacturing, and local government &#8212; integrated into NEC&#8217;s BluStellar Scenario platform and Security Operations Center services. This is one of the first enterprise-scale deployments explicitly organized around Claude Code as a workforce tool rather than a productivity add-on.</p><div><hr></div><h2><strong>&#128218; Tutorials and Resources</strong></h2><h3><strong><a href="https://thenewstack.io/github-copilot-signups-paused/">GitHub Pauses Copilot Sign-ups as AI Coding Drives Up Compute Demand</a></strong></h3><p>The New Stack&#8217;s analysis contextualizes GitHub&#8217;s pause as a structural pricing failure: flat subscriptions were designed for autocomplete-style interactions, not for autonomous multi-agent sessions that can run for hours and spawn sub-agents in parallel. The article documents the shift happening in the background &#8212; GitHub is actively prototyping token-based billing tiers that reflect actual consumption. Worth reading not just for the GitHub news, but for the broader lesson about what happens when AI capability outpaces the business model that sells it.</p><div><hr></div><h2><strong>&#128161; Others</strong></h2><h3><strong><a href="https://techcrunch.com/2026/04/21/spacex-is-working-with-cursor-and-has-an-option-to-buy-the-startup-for-60-billion/">SpaceX Obtains Option to Acquire Cursor for $60 Billion</a></strong></h3><p>SpaceX obtained the right to acquire Cursor for <strong>$60 billion</strong>, preempting a $2 billion funding round that was in progress. The deal pairs Cursor&#8217;s distribution with SpaceX&#8217;s Colossus supercomputer. TechCrunch&#8217;s read is worth noting: neither Cursor nor xAI has models that match Anthropic or OpenAI, and the partnership is partly an attempt to paper over that gap. The deal is significant as a valuation data point and as a signal of how distribution assets are being priced in the current AI landscape &#8212; but whether it changes Cursor&#8217;s product trajectory depends entirely on whether the model gap gets closed.</p><h3><strong><a href="https://www.technologyreview.com/2026/04/24/1136422/why-deepseeks-v4-matters/">Three Reasons Why DeepSeek&#8217;s V4 Matters</a></strong></h3><p>MIT Technology Review&#8217;s analysis goes beyond the benchmark comparison to highlight three structural implications of DeepSeek V4. First: open-source frontier performance &#8212; V4-Pro matches Claude Opus 4.6 and GPT-5.4 on major benchmarks while remaining openly downloadable. Second: a memory efficiency breakthrough &#8212; V4 processes 1 million tokens using only 27% of the compute its predecessor required. Third, and most geopolitically significant: <strong>V4 is DeepSeek&#8217;s first model optimized for Huawei Ascend chips</strong>, marking a concrete step toward Chinese AI infrastructure independence from Nvidia. For developers, the first two points are immediately practical; the third shapes the landscape of model availability in the years ahead.</p><h3><strong><a href="https://www.technologyreview.com/2026/04/21/1135643/10-ai-artificial-intelligence-trends-technologies-research-2026/">MIT Technology Review: 10 Things That Matter in AI Right Now</a></strong></h3><p>MIT Technology Review&#8217;s first annual ranking, unveiled at their EmTech AI conference on April 21, names <strong>agent orchestration</strong> &#8212; multi-agent systems cooperating on complex goals &#8212; as one of the defining shifts in AI right now. Also on the list: AI co-scientists running autonomous research tasks, China&#8217;s open-source strategy distributing frontier models as downloadable packages, and a growing global resistance movement against AI development. For developers, the agent orchestration entry is the most actionable: the list characterizes it as moving from experimental to production-stage, the clearest signal yet that multi-agent architecture belongs in enterprise planning now, not next year.</p><h3><strong><a href="https://fortune.com/2026/04/21/services-are-the-new-software-sequoia-venture-capital-julien-bek-ai-native-eye-on-ai/">Services Are the New Software &#8212; Sequoia Partner on the Next $1 Trillion Company</a></strong></h3><p>Sequoia partner Julien Bek argues the world&#8217;s next $1 trillion company won&#8217;t sell software as a product &#8212; it will sell outcomes delivered by AI-powered software alongside human expertise. The thesis identifies a &#8220;sweet spot&#8221; for AI-native service firms: outsourced functions that are intelligence-heavy with minimal judgment requirements. Pricing shifts from per-seat to outcome-based; human experts monitor rather than execute &#8212; the aviation autopilot model. The implication for developers building AI tools: your customer&#8217;s mental model of what they&#8217;re buying is changing faster than most SaaS pricing pages reflect.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You know that GPT-5.5 leads on agent tasks but trails Opus 4.7 on real codebases &#8212; and that DeepSeek delivers 55% of the performance at 12% of the price. You know that GitHub&#8217;s pricing pause is not a glitch but the first crack in flat-rate AI subscriptions. You didn&#8217;t spend your weekend reading to know this. I did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes, Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #29]]></title><description><![CDATA[The week AI replaced a design team, fixed a security flaw by declaring it a feature, and made SWE-bench nearly obsolete.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-29</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-29</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 20 Apr 2026 04:24:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!912m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Security Story:</strong> OX Security disclosed an architectural RCE flaw in MCP affecting 200,000 servers and 150M+ downloads &#8212; Anthropic confirmed it&#8217;s working as designed and declined to patch it</p></li><li><p><strong>The Workforce Story:</strong> Snap cut 1,000 jobs (16% of headcount) because AI now generates 65% of its new code &#8212; the most explicit use of AI coding productivity as a direct justification for workforce reduction</p></li><li><p><strong>The Model Story:</strong> Claude Opus 4.7 ships with a new <code>xhigh</code> effort level, 10&#8211;14% better task success rates in enterprise testing, and higher-resolution vision for computer-use agents</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> Anthropic launched Claude Design on April 17 &#8212; a product that turns Opus 4.7 into a design collaborator capable of building interactive prototypes, pitch decks, and full design systems from your codebase. Designs hand off directly to Claude Code for implementation. Figma&#8217;s stock dropped the day it launched. This is the first product from a major AI lab that explicitly bridges the designer-developer gap through an agent rather than a tool. <a href="https://www.anthropic.com/news/claude-design-anthropic-labs">Read more &#8594;</a></p></blockquote><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!912m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!912m!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!912m!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!912m!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!912m!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!912m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7715930,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/194669813?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!912m!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!912m!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!912m!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!912m!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed1daa1a-6101-4156-9139-8c524e0c644e_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Key Takeaways</h2><ul><li><p><strong>Claude Opus 4.7 raises the floor on hard coding tasks:</strong> Anthropic&#8217;s new flagship model shows meaningful improvements specifically on the most difficult software engineering tasks, with enterprise testers reporting 10&#8211;14% better task success rates and fewer tool errors in complex multi-step workflows. The new <code>xhigh</code> effort level gives developers finer control over the reasoning-vs-latency tradeoff, and Claude Code defaults to it automatically. <a href="https://www.anthropic.com/news/claude-opus-4-7">Read more &#8594;</a></p></li><li><p><strong>Claude Code&#8217;s desktop redesign treats parallel agent orchestration as the default workflow:</strong> The new app adds a multi-session sidebar for managing concurrent tasks across multiple repos, drag-and-drop pane layout, an integrated terminal, and a built-in diff viewer &#8212; but the bigger shift is Routines: scheduled, API-triggered, and webhook-based automations that run on Anthropic&#8217;s infrastructure, not your laptop. Pro users get 5/day, Max 15, Team/Enterprise 25. This is Claude Code becoming a persistent infrastructure layer, not just a terminal tool. <a href="https://claude.com/blog/claude-code-desktop-redesign">Read more &#8594;</a></p></li><li><p><strong>The MCP security flaw is architectural and Anthropic won&#8217;t fix it:</strong> OX Security researchers found that MCP&#8217;s STDIO transport mechanism allows arbitrary OS command execution &#8212; not a bug, but a deliberate design decision affecting every MCP SDK across Python, TypeScript, Java, and Rust. Anthropic&#8217;s response: sanitization is the developer&#8217;s responsibility. With 200,000 vulnerable server instances and attack vectors through Cursor, Claude Code, and Windsurf, this places a significant undisclosed burden on every team shipping MCP-connected tools. <a href="https://www.theregister.com/2026/04/16/anthropic_mcp_design_flaw/">Read more &#8594;</a></p></li><li><p><strong>Snap&#8217;s announcement is the most explicit &#8220;AI replaced headcount&#8221; statement in enterprise tech so far:</strong> 1,000 jobs gone, 16% of workforce, $500M+ in annual savings &#8212; explicitly because AI generates 65% of new code, handles 1M+ monthly support questions, and flags 7,500+ bugs via a code-review agent. This is no longer a speculative future; it&#8217;s a current-quarter earnings call justification. <a href="https://techcrunch.com/2026/04/15/snap-is-cutting-1000-jobs-16-of-its-workforce/">Read more &#8594;</a></p></li><li><p><strong>Stanford&#8217;s AI Index 2026 puts a number on the SWE-bench leap: 60% to near 100% in one year:</strong> The 423-page report published April 14 documents the fastest single-year capability gain on any standardized software engineering benchmark on record. AI-driven developer productivity gains average 26% per MIT research cited in the report. Entry-level developer employment for ages 22&#8211;25 has declined ~20% since 2022. The benchmark improvement and the employment data are not unrelated trends. <a href="https://hai.stanford.edu/ai-index/2026-ai-index-report">Read more &#8594;</a></p></li><li><p><strong>Cloudflare Agents Week shipped the infrastructure layer that makes production agents actually feasible:</strong> Dynamic Workers (100x faster than containers), Sandboxes GA (persistent isolated Linux environments), Cloudflare Mesh (private agent networking), and AI Gateway (unified access to 70+ models) were all released during the April 13&#8211;17 event. The Think framework for long-running multi-step tasks in the Agents SDK is the piece most directly useful for developers building stateful coding agents. <a href="https://blog.cloudflare.com/welcome-to-agents-week/">Read more &#8594;</a></p></li></ul><div><hr></div><p><strong>Growing at 20% new subscribers per week.</strong></p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p><strong>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</strong></p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://github.blog/changelog/2026-04-16-manage-agent-skills-with-github-cli">GitHub Launches </a><code>gh skill</code><a href="https://github.blog/changelog/2026-04-16-manage-agent-skills-with-github-cli">: Agent Skill Management for the CLI</a></h3><p>GitHub launched <code>gh skill</code>, a new CLI command for discovering, installing, version-pinning, and publishing agent skills across Copilot, Claude Code, Cursor, Codex, and Gemini CLI. The tool includes <strong>content-addressed change detection</strong> that tracks git tree SHA values to identify genuine content changes &#8212; a direct supply chain safeguard designed for the era of agent-executed code. Skills can be locked to a specific tag or commit SHA with <code>--pin</code>. Requires GitHub CLI 2.90.0+; available now in public preview.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-04-17-github-copilot-cli-now-supports-copilot-auto-model-selection">GitHub Copilot CLI Gets Auto Model Selection (10% Discount Included)</a></h3><p>Copilot CLI&#8217;s auto model selection is now generally available for all Copilot plans, dynamically routing requests to GPT-5.4, GPT-5.3-Codex, Sonnet 4.6, and Haiku 4.5 based on the task and subscription tier. The model chosen is shown transparently in the CLI output so developers can verify routing. Paid subscribers receive a <strong>10% discount</strong> on model multipliers when using auto &#8212; a direct incentive to let Copilot decide rather than pinning to a specific model.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-04-14-model-selection-for-claude-and-codex-agents-on-github-com">GitHub Copilot: Model Selection Now Available for Claude and Codex Agents</a></h3><p>When assigning tasks to the Claude or Codex coding agents on github.com, you can now choose which underlying model powers the session. Claude options include Sonnet 4.6, Opus 4.6, Sonnet 4.5, and Opus 4.5; Codex options include GPT-5.2-Codex, GPT-5.3-Codex, and GPT-5.4. This gives teams the ability to route budget-sensitive tasks to faster models while reserving the most capable ones for complex multi-file refactors. Access is included with existing Copilot subscriptions.</p><div><hr></div><h3><a href="https://developers.openai.com/codex/changelog">OpenAI Codex Gets Computer Use, 90+ Plugins, and Persistent Memory</a></h3><p>OpenAI released a major update to Codex on April 16 that expands it well beyond a coding assistant. <strong>Computer use</strong> lets Codex operate macOS apps natively &#8212; seeing, clicking, and typing with its own cursor &#8212; for running tests and GUI-based workflows without leaving the coding context. The update also ships <strong>persistent memory</strong> (Codex now retains preferences and context across sessions), an in-app browser for previewing and requesting changes based on rendered output, and over 90 new plugins covering Atlassian, GitLab, CircleCI, and the Microsoft Suite. CLI v0.121.0, released the same day, adds a <code>codex marketplace</code> command, parallel MCP tool calls, and a bubblewrap-based secure devcontainer profile. Reaching more than 3 million weekly developers, this is the most direct feature-for-feature challenge OpenAI has mounted against Claude Code&#8217;s expanding desktop footprint.</p><div><hr></div><h3><a href="https://techcrunch.com/2026/04/15/openai-updates-its-agents-sdk-to-help-enterprises-build-safer-more-capable-agents/">OpenAI Agents SDK: Sandboxing and Durable Execution for Production Agents</a></h3><p>OpenAI&#8217;s updated Agents SDK adds two capabilities that move it meaningfully closer to production-grade infrastructure. <strong>Sandboxing</strong> isolates agent operations in controlled compute environments, keeping credentials out of the execution space where model-generated code runs. <strong>Snapshotting and rehydration</strong> enable durable execution &#8212; if a sandbox container fails or expires mid-task, the SDK restores agent state in a fresh container and continues from the last checkpoint. Launching first in Python; TypeScript support planned. Available to all API customers at standard pricing.</p><div><hr></div><h3><a href="https://github.com/sst/opencode/releases">OpenCode v1.4.7&#8211;v1.14.17: LLM Gateway, Azure Prompt Caching, and Opus 4.7 xhigh</a></h3><p>OpenCode shipped eight releases between April 15 and 19, with the substantive additions concentrated in v1.4.7&#8211;v1.4.9 (April 16-17). <strong>LLM Gateway</strong> is now a supported provider, with per-model usage reporting surfaced directly in the UI. <strong>Azure prompt caching</strong> ships with per-session cache keys by default &#8212; meaningful for teams running long sessions against Azure-hosted models. Claude Opus 4.7&#8217;s <code>xhigh</code> adaptive reasoning is enabled. The v1.4.5 release adds <strong>OTEL telemetry export</strong> to any OTLP trace backend, giving teams observability over agent runs without custom instrumentation. For teams evaluating open-source coding agents, OpenCode is moving at a notably faster cadence than either Claude Code or Codex on provider breadth and infrastructure integration.</p><div><hr></div><h3><a href="https://geminicli.com/docs/changelogs/latest/">Gemini CLI v0.38.1: Subagents, Chapters, and Context Compression</a></h3><p>Google shipped Gemini CLI v0.38.1 on April 15 with three structural improvements to how the agent handles long-running tasks. <strong>Subagents</strong> &#8212; previewed in a Google Developers Blog post the same day &#8212; now run with isolated context windows, custom system instructions, and their own tool sets; the primary agent delegates to a generalist, a cli_help specialist, or a codebase_investigator, and results collapse back into a single response to avoid context overflow. <strong>Chapters</strong> group tool calls into topic-labeled segments so sessions with dozens of steps remain navigable. A new <strong>context compression service</strong> distills conversation history automatically when approaching token limits. The preview release (v0.39.0-preview.0, April 14) adds a <code>/memory</code> inbox for reviewing agent-extracted skills and skill patching &#8212; the first step toward a persistent, session-to-session learning loop in Gemini CLI.</p><div><hr></div><h3><a href="https://huggingface.co/Qwen/Qwen3.6-35B-A3B">Qwen3.6-35B-A3B: Open-Weight Agentic Coding at 73.4% SWE-Bench</a></h3><p>Alibaba&#8217;s Qwen team released Qwen3.6-35B-A3B, an open-weight mixture-of-experts model (35B parameters, <strong>3B active at inference</strong>) optimized for agentic coding. It scores <strong>73.4% on SWE-bench Verified</strong> and 86.0% on GPQA, with a 262K native context window extensible to 1M tokens. The model includes always-on chain-of-thought reasoning, native tool-use via the Chat Completions API, and MCP support. It integrates with Qwen Code, the team&#8217;s open-source terminal-based coding agent &#8212; making this the most capable open-weight model currently available for self-hosted agentic coding workflows.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://www.infoq.com/news/2026/04/cursor-3-agent-first-interface/">Cursor 3&#8217;s Agent-First Interface: What Actually Changed for Developers (InfoQ Deep-Dive)</a></h3><p>InfoQ&#8217;s April 16 analysis synthesizes two weeks of developer reactions to Cursor 3&#8217;s Agents Window. The new <code>/worktree</code> command creates isolated git worktrees so agent changes happen separately from your main branch; <code>/best-of-n</code> runs the same task simultaneously across multiple models and compares outcomes. For teams evaluating whether to adopt the new interface: the piece concludes that Cursor 3 is less a productivity enhancement to the existing IDE and more an early prototype of a development model where the developer role is <strong>agent selection, prompting, and review</strong> &#8212; not code authoring.</p><div><hr></div><h3><a href="https://kblincoe.github.io/publications/2026_ICSE_SEIP_vibe-coding.pdf">Vibe Coding in Practice: First Academic Grey Literature Review (ICSE 2026)</a></h3><p>Presented April 15 at ICSE 2026 in Rio de Janeiro, this paper is the first systematic academic review of vibe coding practices drawn from 101 practitioner sources and <strong>518 firsthand behavioral accounts</strong>. The central finding: vibe coders are motivated by speed and accessibility, but <strong>QA practices are consistently skipped</strong> &#8212; many developers delegate testing back to the AI that generated the code, creating a closed loop with no external verification. The authors argue that the key missing tooling is not better code generation, but better <strong>automated guardrails that activate when QA steps are bypassed</strong>. Worth reading for team leads assessing vibe coding adoption risk.</p><div><hr></div><h2>&#128161; Others</h2><h3><a href="https://www.technologyreview.com/2026/04/13/1135675/want-to-understand-the-current-state-of-ai-check-out-these-charts/">MIT Technology Review: What the Stanford AI Index Actually Means for Developers</a></h3><p>MIT Technology Review&#8217;s April 13 summary of the Stanford AI Index 2026 focuses on the numbers most relevant to software teams. AI-related GitHub projects have surged to <strong>5.58 million</strong> &#8212; roughly a fivefold increase since 2020 &#8212; signaling grassroots developer adoption well beyond enterprise mandates. The benchmarks measuring AI progress are, in the report&#8217;s own words, &#8220;struggling to keep up&#8221; with actual capability gains. The article argues this benchmark lag is the real governance problem: the tools outpace the metrics we use to decide whether they&#8217;re ready.</p><div><hr></div><p>This week delivered what the past six months have been building toward: a model that can be trusted with your hardest coding tasks (Opus 4.7), an infrastructure layer that makes running those agents persistently practical (Cloudflare, Claude Code Routines), and the first tool from a major AI lab that targets the designer-developer boundary directly (Claude Design). At the same time, the MCP security flaw disclosure &#8212; and Anthropic&#8217;s response &#8212; revealed a governance model for agentic infrastructure that places the entire burden on developers. The capability curve and the accountability curve are not moving at the same speed.</p><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know that Claude Routines are not a feature &#8212; they&#8217;re a new execution model. You know why Anthropic&#8217;s &#8220;by design&#8221; response to the MCP flaw is actually the more alarming answer. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #28]]></title><description><![CDATA[Anthropic just decided it's in the infrastructure business.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-28</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-28</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 13 Apr 2026 06:01:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!B2hq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Big Story:</strong> Anthropic launched Project Glasswing &#8212; a $100M+ cybersecurity initiative using Claude Mythos Preview to hunt and fix zero-day vulnerabilities across every major OS and browser, in partnership with AWS, Apple, Cisco, Google, and Microsoft</p></li><li><p><strong>The Platform Story:</strong> Claude Cowork reached general availability with six enterprise governance features, while Claude Managed Agents launched in public beta &#8212; Anthropic is now directly competing in the agent infrastructure layer</p></li><li><p><strong>The Tooling Story:</strong> GitHub shipped Rubber Duck, nested subagents, and Dependabot-to-agent assignment in the same week; Cursor Bugbot learned to remember your feedback; and Windsurf launched a smart model router to stretch your quota further</p></li></ul><blockquote><p><strong>If you only read one thing this week:</strong> Project Glasswing is not a product &#8212; it&#8217;s a signal. Anthropic took an unreleased frontier model with the ability to find zero-days in every major OS and browser, partnered with AWS, Apple, Cisco, Google, and Microsoft, and deployed it defensively before anyone else could deploy it offensively. The infrastructure layer of AI is no longer neutral territory. <a href="https://www.anthropic.com/glasswing">Read more &#8594;</a></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!B2hq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!B2hq!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!B2hq!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!B2hq!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!B2hq!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!B2hq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8770904,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/193974854?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!B2hq!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!B2hq!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!B2hq!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!B2hq!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ecc0e7-66df-4a6a-a27d-0f4173fce545_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Key Takeaways</h2><ul><li><p><strong>Claude Managed Agents enter public beta, reducing prototype-to-production timelines by 10x:</strong> For $0.08 per session-hour on top of standard token rates, enterprises get fully managed infrastructure: sandboxed environments, long-running autonomous sessions, built-in tool execution (bash, files, web), and multi-agent coordination for parallelizing complex tasks. Notion, Rakuten, and Sentry are already live. This moves Anthropic into direct competition with startups building agent harnesses on top of the Claude API. <a href="https://the-decoder.com/anthropic-launches-managed-infrastructure-for-autonomous-ai-agents/">Read more &#8594;</a></p></li><li><p><strong>Claude Cowork is now generally available &#8212; and most of its enterprise users aren&#8217;t engineers:</strong> All paid plans gain access to Claude Cowork on macOS and Windows, with six new enterprise governance features including role-based access controls, group spend limits, and a Zoom MCP connector. The more striking data point: the vast majority of Cowork usage is coming from operations, marketing, finance, and legal &#8212; not engineering. The &#8220;developer tool&#8221; label may already be obsolete. <a href="https://claude.com/blog/cowork-for-enterprise">Read more &#8594;</a></p></li><li><p><strong>GitHub Copilot CLI&#8217;s Rubber Duck closes 74.7% of the Sonnet-to-Opus performance gap:</strong> An experimental second model from a different AI family independently reviews the primary agent&#8217;s plans after drafting, after complex implementation, and after writing tests. The result: near-Opus reasoning quality at Sonnet price points. This is a meaningful architectural signal &#8212; multi-model review is becoming a standard pattern for pushing coding agent quality. <a href="https://github.blog/ai-and-ml/github-copilot/github-copilot-cli-combines-model-families-for-a-second-opinion/">Read more &#8594;</a></p></li><li><p><strong>Dependabot can now hand security alerts directly to AI agents for remediation:</strong> When a vulnerability requires more than a version bump &#8212; API-breaking changes, complex cross-project updates &#8212; GitHub can now route the alert to Copilot, Claude, or Codex, which analyzes the dependency usage and opens a draft pull request with a proposed fix. The security-to-code pipeline just got a new autonomous step. <a href="https://github.blog/changelog/2026-04-07-dependabot-alerts-are-now-assignable-to-ai-agents-for-remediation/">Read more &#8594;</a></p></li><li><p><strong>Cursor Bugbot now learns from your feedback and hit a 78% resolution rate:</strong> Bugbot converts pull request review feedback into &#8220;learned rules&#8221; that improve future reviews &#8212; making it a system that gets better the more your team uses it. The same update adds MCP server support for Teams and Enterprise plans, a &#8220;Fix All&#8221; action, and a redesigned settings interface. <a href="https://cursor.com/changelog">Read more &#8594;</a></p></li></ul><div><hr></div><p>Growing at 20% new subscribers per week.</p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Subscribers also get <strong>Change Management in Agentic AI Adoption</strong> &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://github.blog/changelog/2026-04-08-github-copilot-in-visual-studio-code-march-releases/">GitHub Copilot in VS Code: Autopilot Mode, Nested Subagents, Video Support</a></h3><p>The VS Code Copilot March/April release (versions v1.111&#8211;v1.115) delivers three meaningful architectural changes. <strong>Autopilot</strong> &#8212; in public preview &#8212; allows agents to approve their own actions and automatically retry on errors, operating without manual intervention during multi-step tasks. <strong>Nested subagents</strong> allow one subagent to invoke another, enabling decomposition of complex workflows into coordinated specialist chains. And <strong>image and video attachment</strong> in chat means agents can now accept screenshots and screen recordings as inputs &#8212; and return images or recordings of their changes for review. Also included: integrated browser debugging with breakpoints, a unified customizations editor, and a <code>/troubleshoot</code> command for analyzing agent debug logs.</p><div><hr></div><h3><a href="https://windsurf.com/changelog">Windsurf Launches Adaptive Model Router for Quota Efficiency</a></h3><p>Windsurf&#8217;s new &#8220;Adaptive&#8221; option in the model picker automatically selects the optimal underlying model for each task &#8212; <strong>minimizing premium model consumption</strong> while maintaining consistent output quality. The model picker now shows exact token pricing and billing rates directly in the interface, alongside a prompt cache timer integrated into the context window indicator. Available for all self-serve Pro, Max, and Teams subscribers.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-04-10-copilot-cloud-agents-validation-tools-are-now-20-faster">Copilot Cloud Agent Validation Tools Are Now 20% Faster</a></h3><p>GitHub&#8217;s Copilot cloud agent now runs its validation suite &#8212; <strong>CodeQL, the GitHub Advisory Database, secret scanning, and Copilot code review</strong> &#8212; in parallel rather than sequentially, cutting total validation time by 20%. The change means the agent resolves detected issues faster without sacrificing the coverage or depth of security and quality checks. Configurable through Copilot settings in each repository.</p><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://github.blog/changelog/2026-04-08-github-mobile-research-and-code-with-copilot-cloud-agent-anywhere">GitHub Mobile: Research and Code with Copilot Cloud Agent from Your Phone</a></h3><p>GitHub extended Copilot cloud agent to the full mobile workflow, beyond just pull request management. From a phone, developers can now <strong>research a codebase</strong>, generate an <strong>implementation plan before writing any code</strong>, and push code changes to a branch &#8212; reviewing diffs and iterating before deciding whether to open a pull request. For teams with async or distributed workflows, this is a meaningful expansion: the delegation model works regardless of whether you&#8217;re at a desk.</p><div><hr></div><h2>&#128161; Others</h2><h3><a href="https://startupnews.fyi/2026/04/06/cursors-2-billion-bet-the-ide-is-now-a-fallback-not-the-default/">Cursor&#8217;s $2B Bet: The IDE Is Now a Fallback, Not the Default</a></h3><p>This piece cuts to the strategic implication of Cursor 3: the company with the fastest revenue growth in AI coding shipped a product that is <strong>not a code editor</strong>. The Cursor 3 interface is built around the Agents Window &#8212; parallel agents running across local, cloud, worktree, and SSH environments &#8212; with the traditional editor as a fallback when you need to touch code directly. The argument is that Cursor has bet its next growth phase on a world where agents orchestrate development and the IDE is the exception, not the default. Worth reading as a frame for where the entire tool category is heading.</p><div><hr></div><h3><a href="https://dev.to/alexmercedcoder/ai-tools-race-heats-up-week-of-april-3-9-2026-37fl">AI Tools Race: Week of April 3&#8211;9, 2026</a></h3><p>A useful aggregation of the week&#8217;s parallel storylines: <strong>Microsoft released Agent Framework 1.0</strong>, unifying Semantic Kernel and AutoGen with full MCP and A2A support into a single composable stack. <strong>MCP crossed 97 million monthly SDK downloads</strong>, with v2.1 adding Server Cards for capability discovery. And a JetBrains survey of 10,000+ professional developers found that <strong>Claude Code has reached 18% professional adoption</strong> &#8212; tying GitHub Copilot in active usage share despite not existing in measurable form 18 months ago. If you want a quick map of the week&#8217;s competitive movements, this is the fastest read.</p><div><hr></div><p>That&#8217;s a wrap for this week. Three Anthropic announcements in three days &#8212; Glasswing, Cowork GA, Managed Agents &#8212; add up to a company that has decided it is not just in the model business anymore. GitHub extended agent autonomy in six different directions inside a single changelog cycle. Cursor&#8217;s Bugbot now learns from your team&#8217;s feedback; Copilot&#8217;s Rubber Duck uses a second AI family to catch what the first one misses; Windsurf&#8217;s router picks your model so you don&#8217;t have to. The pattern across all of it: the agent is no longer a feature inside your tool &#8212; it&#8217;s becoming the infrastructure underneath your team.</p><p>Next week, the stack keeps moving. So does this newsletter. Fall behind one week, and you&#8217;ll spend the next three catching up.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You know why Anthropic&#8217;s three-day repositioning is a different category of event. You know that Rubber Duck isn&#8217;t a gimmick &#8212; it&#8217;s a signal that multi-model review is becoming a standard pattern. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item><item><title><![CDATA[Vibe Coding Weekly #27]]></title><description><![CDATA[The tools are moving faster than the trust that's supposed to contain them.]]></description><link>https://vibecodingweekly.substack.com/p/vibe-coding-weekly-27</link><guid isPermaLink="false">https://vibecodingweekly.substack.com/p/vibe-coding-weekly-27</guid><dc:creator><![CDATA[Angel Llosa]]></dc:creator><pubDate>Mon, 06 Apr 2026 06:02:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!wCnm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><strong>If you only read one thing this week:</strong> Cursor 3 is not an update &#8212; it&#8217;s a new product. The company rebuilt the entire interface around agents as first-class citizens: parallel execution, multi-repo support, local-cloud handoffs, and a Design Mode that lets agents target UI elements in-browser. The era of the AI IDE is over; this is the AI orchestration layer. <a href="https://cursor.com/blog/cursor-3">Read more &#8594;</a></p></blockquote><p><strong>This week in one satisfying refactor:</strong></p><ul><li><p><strong>The Infrastructure Story:</strong> Oracle cut up to 30,000 jobs in a single morning, replacing human capital with compute capital &#8212; $2.1B in restructuring charges funding a $156B AI data center buildout</p></li><li><p><strong>The Security Story:</strong> Sapphire Sleet compromised the axios npm package (100M+ weekly downloads) for three hours; Anthropic accidentally shipped Claude Code&#8217;s entire source code &#8212; 500,000 lines &#8212; to the public npm registry, then triggered accidental takedowns of 8,100 GitHub repos trying to clean it up</p></li><li><p><strong>The Open Model Story:</strong> Google released Gemma 4 &#8212; Apache 2.0, up to 256K context, runs on a Raspberry Pi 5 &#8212; while GitHub deprecated the entire GPT-5.1-Codex family and replaced it with GPT-5.3-Codex</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!wCnm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!wCnm!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!wCnm!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!wCnm!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wCnm!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_webp, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!wCnm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7280907,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://vibecodingweekly.substack.com/i/193268324?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!wCnm!, /__u/vibecodingweekly.substack.com/w_424, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!wCnm!, /__u/vibecodingweekly.substack.com/w_848, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!wCnm!, /__u/vibecodingweekly.substack.com/w_1272, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wCnm!, /__u/vibecodingweekly.substack.com/w_1456, /__u/vibecodingweekly.substack.com/c_limit, /__u/vibecodingweekly.substack.com/f_auto, /__u/vibecodingweekly.substack.com/q_auto:good, /__u/vibecodingweekly.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5da3558c-7629-4bb7-8f84-abd3633f42b2_2750x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Key Takeaways</h2><ul><li><p><strong>Cursor rebuilt its interface from scratch for the agent era:</strong> Cursor 3 replaces the VS Code&#8211;based editor with a unified workspace where humans orchestrate fleets of parallel agents across repos, environments, and platforms &#8212; locally, in the cloud, and via Slack or GitHub integrations. This is the clearest signal yet that the agent is the primary unit of development work, not the keystroke. <a href="https://cursor.com/blog/cursor-3">Read more &#8594;</a></p></li><li><p><strong>Oracle erased 30,000 jobs in a morning to fund AI infrastructure:</strong> On March 31, employees across the US, India, Canada, Mexico, and Uruguay received termination emails at 6am EST with no prior warning. The $2.1B restructuring charge is expected to generate $8&#8211;10B in annual savings that flow directly into data centers and GPUs &#8212; one of the most explicit examples of compute capital replacing human capital in the enterprise. <a href="https://tech-insider.org/oracle-30000-layoffs-ai-data-center-restructuring-2026/">Read more &#8594;</a></p></li><li><p><strong>The axios supply chain attack is a warning shot for every AI-assisted project:</strong> North Korean state actor Sapphire Sleet compromised axios &#8212; the HTTP client found in almost every JavaScript project &#8212; for three hours on March 31, injecting a cross-platform remote access trojan. Any project auto-updating axios during that window was exposed. AI agents that scaffold new projects with npm dependencies are now a fresh attack surface for exactly this kind of infiltration. <a href="https://snyk.io/blog/axios-npm-package-compromised-supply-chain-attack-delivers-cross-platform/">Read more &#8594;</a></p></li><li><p><strong>Anthropic accidentally leaked Claude Code&#8217;s entire source code &#8212; then accidentally took down 8,100 GitHub repos trying to contain it:</strong> On March 31, a misconfigured debug file bundled into a routine npm release exposed nearly <strong>500,000 lines of code</strong> &#8212; including feature flags for unshipped capabilities like background persistent assistants and remote phone control. Anthropic&#8217;s DMCA response then incorrectly targeted ~8,100 unrelated GitHub repositories before the company retracted the bulk of the notices. No customer data was exposed, but the incident is the sharpest reminder yet that the infrastructure underpinning AI coding tools runs on the same fragile, human-error-prone processes as everything else. <a href="https://www.theguardian.com/technology/2026/apr/01/anthropic-claudes-code-leaks-ai">Read more &#8594;</a></p></li><li><p><strong>Google&#8217;s Gemma 4 runs frontier-quality reasoning on edge hardware:</strong> Released April 2 under Apache 2.0, Gemma 4 comes in four sizes (2B to 31B parameters) and runs on mobile phones, Raspberry Pi 5, and in-browser via WebGPU &#8212; with 256K context support and 140+ language coverage. For developers building offline-capable or privacy-sensitive agentic applications, this removes a major constraint. <a href="https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/">Read more &#8594;</a></p></li><li><p><strong>GitHub Copilot&#8217;s cloud agent can now plan before it codes:</strong> A significant workflow shift: the Copilot cloud agent can now produce an implementation plan for your review before writing a single line of code, conduct deep research sessions grounded in your repository, and work on branches without immediately opening pull requests. The asynchronous, delegated development model just got a meaningful review layer. <a href="https://github.blog/changelog/2026-04-01-research-plan-and-code-with-copilot-cloud-agent/">Read more &#8594;</a></p></li><li><p><strong>Claude Code auto mode resolves the permissions dilemma:</strong> Anthropic&#8217;s auto mode is the middle ground developers have been looking for between constant manual approvals and the <code>--dangerously-skip-permissions</code> flag. Two AI classifiers evaluate each action before execution &#8212; one for prompt injection, one for intent alignment. False positive rate: 0.4% in production. The catch: ~17% of overeager actions still get through, so it&#8217;s designed for isolated environments, not production infrastructure. <em>(editorial inclusion &#8212; outside range, published March 25, 2026)</em> <a href="https://www.anthropic.com/engineering/claude-code-auto-mode">Read more &#8594;</a></p></li></ul><div><hr></div><p>Growing at 20% new subscribers per week.</p><p>The stories this week aren&#8217;t hard to find. What&#8217;s hard is knowing which ones actually matter before your team asks you on Monday.</p><p>That&#8217;s the only thing Vibe Coding Weekly does: cut through the volume so you arrive at the week with context, not anxiety.</p><p>Get the ebook <strong>Change Management in Agentic AI Adoption</strong> when subscribe &#8212; the framework for the conversation that always comes after &#8220;we should use AI more&#8221;: how to actually move an organization that didn&#8217;t ask to be moved. Included with every subscription.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/vibecodingweekly.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h2>&#128230; Releases &amp; News</h2><h3><a href="https://www.theguardian.com/technology/2026/apr/01/anthropic-claudes-code-leaks-ai">Anthropic Leaked Claude Code&#8217;s Source Code &#8212; Then Accidentally Took Down 8,100 GitHub Repos</a></h3><p>On March 31, a debug file accidentally bundled into a routine npm release exposed <strong>nearly 500,000 lines of Claude Code&#8217;s source code</strong> to the public registry &#8212; including feature flags for capabilities that are fully built but unshipped, like a background persistent assistant and remote control from a phone or browser. Anthropic confirmed no customer data or credentials were involved, attributing the incident to human error in the release packaging process. The response compounded the original story: Anthropic&#8217;s DMCA takedown request incorrectly targeted <strong>approximately 8,100 GitHub repositories</strong> &#8212; most unrelated to the leak &#8212; before the company retracted the bulk of the notices and described the overbroad enforcement as also accidental. For a company whose products are now running autonomously in developer workflows, two back-to-back unforced errors in a single day are an uncomfortable data point about operational maturity.</p><div><hr></div><h3><a href="https://www.macrumors.com/2026/03/31/ollama-now-runs-faster-apple-silicon-macs/">Ollama 0.19 Preview: Nearly 2x Faster on Apple Silicon via MLX</a></h3><p>Ollama&#8217;s 0.19 preview release switches to Apple&#8217;s MLX framework, delivering approximately <strong>1.6x faster prompt processing</strong> and <strong>nearly 2x faster response generation</strong> on Apple Silicon &#8212; with M5-series Macs seeing the largest gains. Initial support is limited to Qwen3.5, requires 32GB unified memory, and broader model support is planned. For developers running local coding agents or chat assistants on Mac, the responsiveness difference during extended sessions is described as noticeable.</p><div><hr></div><h3><a href="https://techcrunch.com/2026/03/31/salesforce-announces-an-ai-heavy-makeover-for-slack-with-30-new-features/">Salesforce Gives Slack 30 AI Features &#8212; Including MCP Client Support</a></h3><p>Salesforce announced 30 new AI features for Slack, the headline being that Slackbot now functions as a <strong>Model Context Protocol client</strong> &#8212; meaning it can make tool calls into external services across the 2,600+ Slack Marketplace apps and 6,000+ Salesforce AppExchange integrations. Slackbot can now coordinate with Agentforce agents, create Google Docs and Slides, and operate as an autonomous work assistant. The update positions Slack less as a messaging platform and more as an <strong>agentic operating system</strong> for enterprise workflows.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-04-02-copilot-organization-custom-instructions-are-generally-available/">Copilot Organization Custom Instructions Now Generally Available</a></h3><p>GitHub Copilot&#8217;s organization custom instructions &#8212; allowing admins to set default behavior guidelines that shape how Copilot behaves across all repositories in their organization &#8212; reached general availability. The feature applies across <strong>Copilot Chat on github.com, code review, and the cloud agent</strong>. Configuration lives in Organization Settings &#8594; Copilot &#8594; Custom Instructions. Available for Copilot Business and Enterprise admins.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-04-02-github-actions-early-april-2026-updates/">GitHub Actions: Early April 2026 Updates</a></h3><p>Three noteworthy additions to GitHub Actions: service containers now support <strong>entrypoint and command overrides</strong> via workflow YAML; OIDC tokens now carry <strong>repository custom properties as claims</strong>, enabling more granular cloud trust policies without referencing individual repo names; and Azure private networking for hosted runners gained <strong>failover network support</strong> in public preview for workflow continuity during subnet failures.</p><div><hr></div><h3><a href="https://github.blog/changelog/2026-04-03-gpt-5-1-codex-gpt-5-1-codex-max-and-gpt-5-1-codex-mini-deprecated/">GPT-5.1-Codex Family Deprecated in GitHub Copilot</a></h3><p>As of April 1, 2026, GPT-5.1-Codex, GPT-5.1-Codex-Max, and GPT-5.1-Codex-Mini are deprecated across all Copilot experiences &#8212; chat, inline edits, ask and agent modes, and code completions. The replacement is <strong>GPT-5.3-Codex</strong>. If your team has workflows or API integrations specifying these model names, update them now.</p><div><hr></div><h2>&#128218; Tutorials and Resources</h2><h3><a href="https://www.infoq.com/news/2026/04/pinterest-mcp-ecosystem/">Pinterest&#8217;s Production MCP Playbook: 66,000 Invocations/Month and 7,000 Hours Saved</a></h3><p>Pinterest&#8217;s engineering team published a detailed account of their production MCP deployment &#8212; the most concrete real-world data on MCP at scale published to date. Their architecture uses <strong>domain-specific cloud-hosted MCP servers</strong> (separate servers for Presto, Spark, and Airflow) connected through a central registry. The system processes 66,000 invocations per month across 844 active users, saving roughly <strong>7,000 hours per month</strong>. Two-layer authorization (JWT + service mesh) and mandatory human approval for sensitive operations kept governance intact. Worth reading for anyone planning an enterprise MCP rollout.</p><div><hr></div><h2>&#128161; Others</h2><h3><a href="https://news.harvard.edu/gazette/story/2026/04/vibe-coding-may-offer-insight-into-our-ai-future/">&#8216;Vibe Coding&#8217; May Offer Insight into Our AI Future</a></h3><p>Harvard researchers surfaced a tension that&#8217;s easy to miss in the excitement around AI-assisted development: vibe coding replaces technical knowledge with a different and harder-to-evaluate skill set &#8212; the ability to articulate ideas in natural language. The question isn&#8217;t whether you can build fast, it&#8217;s whether the resulting software is reliable, secure, and maintainable when you can&#8217;t read the code it contains. As Karen Brennan notes, the essential capability becomes &#8220;imagine possibilities, express clearly what we want to see in the world, review what we create, and iterate&#8221; &#8212; and not all developers are equally equipped to do that final step.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://vibecodingweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>Every week, a new model drops. A new agent framework ships. A new &#8220;this changes everything&#8221; thread goes viral. And you still have actual code to write.</p><p>Every Monday, you open your inbox and already know what matters. You&#8217;ve skipped three viral threads that turned out to be nothing. You didn&#8217;t spend your weekend reading to know this. We did.</p><p>That&#8217;s what Vibe Coding Weekly is. For developers, architects, tech leads, and everyone building or managing software in the age of AI.</p><p>Clean code and positive vibes,<br>Angel.</p>]]></content:encoded></item></channel></rss>