<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Kaitchup – AI on a Budget]]></title><description><![CDATA[Weekly tutorials and news on adapting large language models (LLMs) to your tasks and hardware using the most recent techniques and models. The Kaitchup proposes a collection of 180+ AI notebooks regularly updated.]]></description><link>https://kaitchup.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!xY7g!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png</url><title>The Kaitchup – AI on a Budget</title><link>https://kaitchup.substack.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 02 Sep 2026 10:05:48 GMT</lastBuildDate><atom:link href="/__u/kaitchup.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[The Kaitchup]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[kaitchup@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[kaitchup@substack.com]]></itunes:email><itunes:name><![CDATA[Benjamin Marie]]></itunes:name></itunes:owner><itunes:author><![CDATA[Benjamin Marie]]></itunes:author><googleplay:owner><![CDATA[kaitchup@substack.com]]></googleplay:owner><googleplay:email><![CDATA[kaitchup@substack.com]]></googleplay:email><googleplay:author><![CDATA[Benjamin Marie]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Qwen3.8 27B GGUF Benchmark: Q4 to Q1 Accuracy and Token Efficiency]]></title><description><![CDATA[15 GGUFs, 150M+ tokens]]></description><link>https://kaitchup.substack.com/p/qwen38-27b-gguf-benchmark-q4-to-q1</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-27b-gguf-benchmark-q4-to-q1</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Tue, 01 Sep 2026 16:21:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bSVx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Qwen3.8 27B is one of the best models for local AI right now. And for most local LLM users, local AI is experienced through the GGUF lens.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;1f3d7364-f3f0-42c2-b138-b29e033f9c11&quot;,&quot;caption&quot;:&quot;In a previous article, I benchmarked Qwen3.8 27B using its highest reasoning effort, xhigh. That shows the model at its best, but it does not tell us whether xhigh is the most efficient way to run it. Qwen3.8 27B also supports medium and low reasoning efforts, while thinking can be disabled completely.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwen3.8 27B Reasoning Benchmarks: Off vs Low vs Medium vs Xhigh&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-08-26T07:18:11.837Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!ULzf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwen38-27b-reasoning-benchmarks-off&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:212638885,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:17,&quot;comment_count&quot;:3,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>But how good are these GGUFs, actually?</p><p>We don&#8217;t have a clear answer. This is a problem I&#8217;ve discussed many times before.</p><p>Metrics such as KLD and PPL are useful for measuring how much a quantized model deviates from the original model, but they are not anchored in the accuracy space. A KLD or PPL value may indicate that one quantization is worse than another, but it cannot tell you when the degradation becomes large enough that the model stops being useful.</p><p>By design, these metrics cannot answer that question.</p><p>The only reliable way around this is the expensive one: run real benchmarks and generate millions of tokens.</p><p>So that&#8217;s what I did.</p><p>For this article, I generated more than 150 million tokens, requiring roughly 8 days of continuous benchmarking on an NVIDIA RTX Pro 6000.</p><p>I evaluated 15 GGUF versions of Qwen3.8 27B, ranging from Q4 down to Q1, produced by several different providers.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The main goal was to measure two things: accuracy and token efficiency. In other words, how low can you go on a 12 GB, 16 GB, or 24 GB GPU before the accuracy loss becomes significant, or before the model starts generating substantially more tokens to solve the same tasks?</p><blockquote><p><strong>Acknowledgments</strong></p><p><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=qwen38gguf">Verda</a><span> provided the RTX Pro 6000s to run the experiments described in this article.</span></p><p>Verda is a full-stack AI cloud, built for high-performance inference, training, and agentic workloads, with data privacy and sustainability at its core.</p><p><span>You can check them out </span><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=qwen38gguf">here</a><span>. There is a </span><strong>$50 coupon</strong><span> that you can redeem in your Verda account, after provisioning it with $5, to try their GPUs.</span></p><p><strong>Coupon code:</strong><span> </span><code>KAITCHUP-50</code></p><p><span>Follow </span><a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a><span>.</span></p><p><em><span>Note: I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon.</span></em></p></blockquote><h2>List of the evaluated Qwen3.8 GGUF models</h2><p><em>Note: The sizes listed below exclude the MTP and vision layers.</em></p><ul><li><p><a href="https://huggingface.co/unsloth/Qwen3.8-27B-GGUF">unsloth/Qwen3.8-27B-GGUF</a></p><ul><li><p>Qwen3.8-27B-UD-Q4_K_XL.gguf: 17.2 GB</p></li><li><p>Qwen3.8-27B-UD-Q3_K_XL.gguf: 12.8 GB</p></li><li><p>Qwen3.8-27B-UD-IQ3_XXS.gguf: 10.6 GB</p></li><li><p>Qwen3.8-27B-UD-Q2_K_XL.gguf: 9.5 GB</p></li><li><p>Qwen3.8-27B-UD-IQ2_XXS.gguf: 7.3 GB</p></li><li><p>Qwen3.8-27B-UD-IQ1_M.gguf: 6.7 GB</p></li></ul></li><li><p><a href="https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF">huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF</a></p><ul><li><p>Huihui-Qwen3.8-27B-abliterated-UD-Q4_K_XL.gguf: 17 GB</p></li></ul></li><li><p><a href="https://huggingface.co/AtomicChat/Qwen3.8-27B-GGUF">AtomicChat/Qwen3.8-27B-GGUF</a></p><ul><li><p>Qwen3.8-27B-AD-Q4_K_M.gguf: 16.8 GB</p></li><li><p>Qwen3.8-27B-AD-IQ3_S.gguf: 13.6 GB</p></li><li><p>Qwen3.8-27B-AD-IQ2_S.gguf: 10.8 GB</p></li></ul></li><li><p><a href="https://huggingface.co/bartowski/Qwen3.8-27B-GGUF">bartowski/Qwen3.8-27B-GGUF</a></p><ul><li><p>Qwen3.8-27B-IQ4_XS.gguf: 15.3 GB</p></li><li><p>Qwen3.8-27B-IQ3_XXS.gguf: 12.4 GB</p></li></ul></li><li><p><a href="https://huggingface.co/empero-ai/Qwen3.8-27B-Ridge-GGUF">empero-ai/Qwen3.8-27B-Ridge-GGUF</a></p><ul><li><p>Qwen3.8-27B-Ridge-3.7bpw.gguf: 12.3 GB</p></li></ul></li><li><p><a href="https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF">ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF</a></p><ul><li><p>Qwen3.8-27B-GSQ-RCO-IQ3_XXS.gguf: 10.1 GB</p></li><li><p>Qwen3.8-27B-GSQ-RCO-IQ2_S.gguf: 9.3 GB</p></li></ul></li></ul><h2>Qwen3.8 27B GGUF Benchmark Methodology and Evaluation Setup</h2><p>Each GGUF evaluation took roughly 10~15 hours on an NVIDIA RTX Pro 6000, so evaluating every available quantization was simply not practical.</p><p>All GGUFs were evaluated using the exact same set of 950 prompts, subsampled from MMLU-Pro, LiveCodeBench, and GPQA Diamond.</p><p>I used the inference hyperparameters recommended by Qwen rather than setting the temperature to 0. Greedy decoding would have made the results easier to compare, but it would also have produced significantly lower accuracy scores and substantially more generated tokens, increasing the overall evaluation cost. <em>Note: That said, with a much larger budget, I think temperature = 0 is a very good setting for evaluating quantization damage. At temperature 0, even the BF16 model performs worse on these tasks, but it remains relatively stable and does not collapse. Poorly quantized models, on the other hand, may degrade much more severely under the same conditions. In that sense, <strong>greedy decoding can act as a harsher stress test, making quantization-induced degradation easier to expose.</strong></em></p><p>The trade-off is that sampling introduces nondeterminism and fairly high score variance. To compensate for this, I ran <strong>each evaluation three times</strong> and averaged the results.</p><p>All GGUF models were served through a llama.cpp server. I set the maximum output length to 128K tokens and used low thinking mode. This still allows the model to generate sufficiently long reasoning traces to expose quantization damage, while keeping the total evaluation cost manageable.</p><h3>How to read these results</h3><p>The goal of this evaluation is <strong>not to produce a definitive ranking of GGUFs</strong>. Doing that properly across this many quantizations, with enough repetitions to confidently distinguish small differences, would easily cost $10,000+.</p><p>Instead, the goal is to identify which GGUFs show no clear signs of significant degradation compared with the original BF16 model.</p><p>For that purpose, I use a simple threshold: if a GGUF recovers more than 95% of the BF16 model&#8217;s accuracy, I consider it good enough.</p><p>In practice, a GGUF can occasionally score higher than the original BF16 model, but this <strong>should not be interpreted as the quantized model being better</strong>. It is simply a consequence of sampling and benchmark variance.</p><p>Below 95% accuracy recovery, I consider the degradation too large. At that point, you are probably better off using a smaller model at a healthier quantization level rather than forcing a heavily quantized 27B model to fit on your GPU.</p><p>There is also an important limitation to keep in mind.</p><p>All the benchmarks used here are single-turn, non-agentic tasks. Even though this evaluation is already expensive, it is still far too cheap to reliably tell us how well a given GGUF will perform on long-running agentic workloads.</p><p>For tasks such as agentic coding, even <strong>a seemingly small 1% accuracy loss can compound across many successive steps and tool calls</strong>. A model that looks nearly identical to BF16 on a single-turn benchmark may behave quite differently over a long agent trajectory.</p><p>If agentic workloads are your main use case, I would interpret these results asymmetrically: models near or below the 95% threshold are strong candidates to avoid, while models clearly above it should only be considered <em>potentially</em> good enough. How much quantization you can tolerate will ultimately depend on the length and complexity of your agentic tasks.</p><p><em>Note: I&#8217;m running more evaluations for UD-Q2_K_XL and UD-Q4_K_XL for agentic coding. It will take at least a week.</em></p><h2>Qwen3.8 27B GGUFs: Token Efficiency and Accuracy</h2><p>Below, I first present <strong>accuracy</strong> and then <strong>token efficiency</strong> before discussing the results together, because, as we will see, token efficiency is strongly correlated with accuracy degradation.</p><h3>Accuracy</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!bSVx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!bSVx!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg 424w, /__u/substackcdn.com/image/fetch/$s_!bSVx!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg 848w, /__u/substackcdn.com/image/fetch/$s_!bSVx!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg 1272w, /__u/substackcdn.com/image/fetch/$s_!bSVx!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!bSVx!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg" width="1694" height="1058.75" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:910,&quot;width&quot;:1456,&quot;resizeWidth&quot;:1694,&quot;bytes&quot;:16279,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/svg+xml&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/213309982?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!bSVx!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg 424w, /__u/substackcdn.com/image/fetch/$s_!bSVx!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg 848w, /__u/substackcdn.com/image/fetch/$s_!bSVx!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg 1272w, /__u/substackcdn.com/image/fetch/$s_!bSVx!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff359f989-7ae0-4386-a818-eb7284f6fd21_1600x1000.svg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Token Efficiency</h3>
      <p>
          <a href="/__u/kaitchup.substack.com/p/qwen38-27b-gguf-benchmark-q4-to-q1">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Qwen3.8 Needs Preserved Thinking, Q1 GGUFs Fall Apart, and the new GLM-5.3-Flash]]></title><description><![CDATA[The Weekly Kaitchup #157]]></description><link>https://kaitchup.substack.com/p/qwen38-needs-preserved-thinking-q1</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-needs-preserved-thinking-q1</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 29 Aug 2026 02:10:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of <strong>The Weekly Kaitchup</strong>, I mainly discuss the evaluations I&#8217;m currently running and what I plan to publish next. I also share a few thoughts on the new <strong>GLM-5.3 Flash</strong>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Qwen3.8 27B for Agentic Coding: Evaluations Are Still Running</h2><p>I&#8217;m spending a lot of compute testing different configurations, and I&#8217;m starting to see some very interesting results. <em>Note: Thanks to <a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=kaitchup157">Verda</a> for the compute sponsorship.</em></p><p>I&#8217;ll compile everything into a dedicated article, but since it will take some time before it&#8217;s ready, I thought I&#8217;d give you an early preview.</p><h3><code>preserve_thinking</code> is essential</h3><p>Qwen puts it this way in <a href="https://huggingface.co/Qwen/Qwen3.8-27B">the model card</a>:</p><blockquote><p>By default, Qwen3.8 retains thinking blocks from all historical messages, maintaining a complete reasoning trace across the conversation. This behavior, known as preserved thinking, ensures full context continuity and is especially beneficial for agent scenarios where decision consistency and reduced redundant reasoning are critical. It also improves KV cache utilization, optimizing inference efficiency in both thinking and non-thinking modes.</p></blockquote><p>I can confirm this, but I&#8217;d go even further: if you disable <code>preserve_thinking</code>, your agent will generate many more tokens and fail much more often.</p><p>It&#8217;s not just beneficial. It&#8217;s necessary.</p><p>For example, if you disable <code>preserve_thinking</code> with <code>xhigh</code>, you get <code>medium</code>-level performance (or less) while consuming significantly more tokens.</p><p>Also, double-check your agent setup. Claude Code wasn&#8217;t preserving thinking traces in my initial runs, and it took quite a few retries and Codex&#8217;s help before I got the configuration right. I haven&#8217;t had issues with mini-swe-agent or Pi.</p><h3>Pi Outperforms Qwen&#8217;s Published Score</h3><p>I made small modifications to Pi to better handle <code>NonZero</code> errors. With that change, my <code>xhigh</code> runs on DeepSWE 1.1 reached a score of <strong>46</strong>, four points higher than the result reported by Qwen with Claude Code.</p><p>I&#8217;ll publish the trajectories along with my modifications when I release the full article.</p><p>Maybe I should try OMP.</p><h3>What&#8217;s Still Running?</h3><p>I&#8217;m making a few final attempts to get a strong Claude Code run. So far, I haven&#8217;t managed to reproduce Qwen&#8217;s results.</p><p>My best DeepSWE 1.1 score with Claude Code is currently <strong>35.2</strong>, about seven points below the score Qwen reported with their Claude Code setup.</p><p>Claude Code was dropping the reasoning traces between turns, but I&#8217;ve now fixed the <code>preserve_thinking</code> issue. Hopefully, that will help close the accuracy gap.</p><p>I&#8217;m also running a few baselines using my best Pi configuration:</p><ul><li><p>Qwen3.6 27B with thinking enabled and <code>preserve_thinking</code> on</p></li><li><p>Qwen3.8 27B with <code>low</code> thinking</p></li><li><p>Qwen3.8 27B with thinking turned off</p></li></ul><p>These runs will probably be finished by Monday, and I expect to publish the article later in the week.</p><h2>Muse Glimmer Quantization Analysis</h2><p>All evaluations of the quantized Muse Glimmer models are now complete.</p><p>What remains is writing the analysis, but there&#8217;s so much going on right now that it&#8217;s been difficult to find the time. Still, I can already share a few results:</p><ul><li><p><a href="https://huggingface.co/cyankiwi/Muse-Glimmer-30B-AWQ-INT4">Cyankiwi&#8217;s 4-bit AWQ quantization</a> provides the best accuracy preservation.</p></li><li><p>On Blackwell, <a href="https://huggingface.co/RedHatAI/Muse-Glimmer-30B-NVFP4">NVFP4</a> is probably the better choice, especially for high-concurrency workloads.</p></li><li><p>My <a href="https://huggingface.co/kaitchup/Muse-Glimmer-30B-autoscheme-3.5bit">mixed-precision INT version</a> also performs well, but it slightly, and consistently, underperforms those two. That&#8217;s an acceptable cost to save 4 more GBs.</p></li></ul><p>And since Muse Glimmer is already kind of forgotten, I don&#8217;t yet have a clear schedule for when I&#8217;ll publish the full analysis.</p><p>Meta was also planning to release the weights for Muse Spark 1.2, but I think it has missed its window. Several newer models now look considerably stronger. Hopefully, Meta will skip directly to a newer version rather than releasing the current 1.2 weights.</p><h2>Ornith-1.5 Quantization</h2><p>I also created a mixed-precision quantization of Ornith-1.5 9B.</p><ul><li><p><a href="https://huggingface.co/collections/kaitchup/quantized-ornith-15">Quantized Ornith-1.5 9B</a></p></li></ul><p>The quantized Ornith-1.5 9B models are compatible with vLLM 0.26+ and should be particularly well suited to high-concurrency workloads.</p><p>With the KV cache quantized to FP8, the model should run comfortably on a 12 GB GPU. Without KV-cache quantization, you&#8217;ll want 16 GB to use the full context window.</p><p>If there&#8217;s enough demand, I&#8217;ll also quantize the 35B version.</p><h2>Qwen3.8 27B GGUF Evaluations</h2><p>I&#8217;ve now evaluated around a dozen Qwen3.8 27B GGUFs from several providers. The full evaluation has taken more than a week on a single RTX Pro 6000, with each individual GGUF typically requiring <strong>10&#8211;15 hours</strong>. The very low-bit, nearly broken quantizations can take even longer.</p><p>The main conclusion so far is fairly clear:</p><p><strong>Below roughly 9&#8211;9.5 GB, you&#8217;re probably giving up too much accuracy.</strong></p><p>As usual, the quantizations from Unsloth, LM Studio, AtomicChat, llama.cpp, and bartowski are very good. Down to <code>UD-Q2_K_XL</code> at 9.83 GB, they retain more than 95% of BF16 accuracy in my tests, including at context lengths of up to 128K tokens.</p><p>Below that, however, things deteriorate quickly.</p><p><code>UD-IQ1_M</code>, for instance, loses more than half of the BF16 accuracy on some coding evaluations. It&#8217;s impressive that a model compressed this aggressively still works at all, but that doesn&#8217;t necessarily make it useful.</p><p>On LiveCodeBench, <code>UD-IQ1_M</code> performs slightly worse than a <strong>4-bit Qwen3 4B</strong>, the first-generation Qwen3 model, while also requiring more memory and generating more tokens. On GPQA Diamond, its performance is not far from random.</p><p>So, in practical terms: <strong>I wouldn&#8217;t recommend Q1 quantizations of Qwen3.8 27B.</strong></p><p>If your memory budget forces you below roughly 9.5 GB, I think it makes more sense to use a smaller model at a higher precision, for example, a Q4 Qwen3.5 9B, or Ornith-1.5 9B for agentic workloads.</p><p><code>UD-Q2_K_L</code> looks more promising, although there&#8217;s an important limitation to my current results: I&#8217;ve only tested single-turn tasks, with context lengths up to 128K tokens.</p><p>For long-horizon agentic workloads, the picture could be different. Quantization errors can accumulate across many turns, tool calls, and reasoning steps, so a quantization that looks good on single-turn benchmarks may degrade more noticeably during an extended coding trajectory.</p><p>I&#8217;d like to test this properly with long-horizon agentic coding tasks. </p><p>For now, I expect the main Qwen3.8 27B GGUF evaluations to be completed this weekend. I&#8217;ll publish a dedicated article with the full <strong>accuracy and token-efficiency results</strong>, and I&#8217;m also evaluating several abliterated versions.</p><h2>GLM-5.3-Flash: Smaller, but Also with a New Efficient Architecture</h2><p>When the first rumors about <strong><a href="https://huggingface.co/zai-org/GLM-5.3-Flash">GLM-5.3-Flash</a></strong> appeared, I assumed we were getting another model in the same category as <a href="https://huggingface.co/zai-org/GLM-4.7-Flash">GLM-4.7-Flash</a>.</p><p>GLM-4.7-Flash was a lovely <strong>30B-A3B MoE</strong>: 30 billion parameters in total, with only around 3 billion activated per token. It was large enough to be capable, yet still small enough to be genuinely interesting for local inference. But it was quickly overshadowed. Gemma 4 and the Qwen3.5 MoEs arrived soon after and appeared significantly stronger, pushing GLM-4.7-Flash out of the spotlight almost immediately.</p><p>GLM-5.3-Flash is very, very far from that size.</p><p><strong>It has 320 billion parameters, with 18 billion active.</strong></p><p>At this point, &#8220;Flash&#8221; describes the inference economics rather than the size of the model.</p><p>GLM-5.3-Flash is a 45-layer Mixture-of-Experts model with 288 routed experts, eight of which are selected per token, plus a shared expert. </p><p>More interestingly, Z.ai has significantly reworked the attention architecture. GLM-5 already replaced conventional full attention with DeepSeek Sparse Attention, but GLM-5.3-Flash goes a step further: for the first time in the GLM family, it combines <strong>linear attention with periodic sparse-attention layers</strong>. Linear attention handles cheaper local and state-based processing, while sparse attention is used periodically to retrieve relevant information from the full context. The result is much lower compute and KV-cache requirements at very long context lengths.</p><p>The model also introduces <strong>Manifold-Constrained Hyper-Connections (mHC)</strong> to the GLM architecture, replacing conventional residual connections with a more flexible routing mechanism designed to improve training stability and scaling efficiency.</p><p>There is even an additional optimization called <strong>IndexPool</strong>, which compresses the index used by the sparse-attention layers. Z.ai says that compared with the full GLM-5.3, the Flash architecture reduces attention compute by about 3&#215; and average KV-cache size by about 4.4&#215;.</p><p>Another major change: this is the <strong>first natively multimodal model in the GLM-5 family</strong>. Text and vision are part of the model from pre-training rather than vision being bolted on afterwards. Z.ai says the new base model was trained on a 30-trillion-token multimodal corpus.</p><p>And the performance looks extremely strong.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!21aB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!21aB!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png 424w, /__u/substackcdn.com/image/fetch/$s_!21aB!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png 848w, /__u/substackcdn.com/image/fetch/$s_!21aB!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png 1272w, /__u/substackcdn.com/image/fetch/$s_!21aB!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!21aB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png" width="1456" height="908" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:908,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;bench_53&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="bench_53" title="bench_53" srcset="/__u/substackcdn.com/image/fetch/$s_!21aB!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png 424w, /__u/substackcdn.com/image/fetch/$s_!21aB!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png 848w, /__u/substackcdn.com/image/fetch/$s_!21aB!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png 1272w, /__u/substackcdn.com/image/fetch/$s_!21aB!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84e3bcc7-360c-4040-a614-db79a5d3faf8_4239x2643.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Z.ai also <strong>anonymously deployed</strong> the model before release under the name <strong>ox-alpha</strong> on OpenRouter and OpenCode. According to Z.ai, it became the most popular model of the week before anyone knew who had made it.</p><p>But the most interesting detail might be this one:</p><p><strong>Z.ai says all of that traffic was served using Chinese AI chips.</strong></p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisioning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/kaitchup.substack.com/subscribe"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p>]]></content:encoded></item><item><title><![CDATA[Qwen3.8 Flash Next Review: Benchmarks, Architecture, Memory Requirements, and Local Inference]]></title><description><![CDATA[Qwen&#8217;s 125B-A6B model, its 51B n-gram memory, and what it takes to run it locally]]></description><link>https://kaitchup.substack.com/p/qwen38-flash-next-review-benchmarks</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-flash-next-review-benchmarks</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Thu, 27 Aug 2026 05:54:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oEVE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffebf86b5-e288-4c79-83f6-f21b20c49507_2885x2930.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Qwen is moving very fast.</p><p>Only days after Qwen3.8 27B demonstrated just how much performance Qwen can now squeeze into a relatively conventional dense local model, Qwen is already showing us what Qwen4 will look like.</p><p>Qwen3.8-Flash-Next is a multimodal Mixture-of-Experts model with a 125B-parameter main model, only 6B parameters activated per token, another 51B parameters dedicated to a new n-gram embedding system, and a 4B MTP module for speculative decoding.</p><ul><li><p><a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next">Qwen/Qwen3.8-Flash-Next</a></p></li></ul><p>But the most interesting detail is in its name.</p><p>Flash-<strong>Next</strong> is not just another member of the Qwen3.8 family. Qwen explicitly describes it as an experimental preview of the architecture that will underpin Qwen4. In fact, its configuration file in Transformers already identifies the architecture as <code>qwen4_exp</code>.</p><p>This is similar to what Qwen did with Qwen3-Next before Qwen3.5. Qwen3-Next introduced the hybrid Gated DeltaNet and Gated Attention architecture that was then reused throughout Qwen3.5, Qwen3.6, Qwen3.7, and Qwen3.8.</p><p>Now Qwen is doing it again.</p><p>In this article, I&#8217;ll first look at how Qwen3.8-Flash-Next compares with Qwen3.8 27B. Then I&#8217;ll dig into the new architecture, especially Qwen Sparse Attention, Gated Residual, and the unusual 51B-parameter n-gram embedding table.</p><p>Finally, I&#8217;ll estimate its actual memory consumption at 256K context, look at NVFP4 and Unsloth&#8217;s GGUF quantizations, and see what kind of hardware may realistically run it.</p><h2>How good is Qwen3.8-Flash-Next?</h2>
      <p>
          <a href="/__u/kaitchup.substack.com/p/qwen38-flash-next-review-benchmarks">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Qwen3.8 27B Reasoning Benchmarks: Off vs Low vs Medium vs Xhigh]]></title><description><![CDATA[How much accuracy do extra reasoning tokens actually buy?]]></description><link>https://kaitchup.substack.com/p/qwen38-27b-reasoning-benchmarks-off</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-27b-reasoning-benchmarks-off</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Wed, 26 Aug 2026 07:18:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ULzf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ULzf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ULzf!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!ULzf!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!ULzf!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ULzf!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ULzf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png" width="644" height="483" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:644,&quot;bytes&quot;:2266012,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/212638885?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ULzf!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png 424w, /__u/substackcdn.com/image/fetch/$s_!ULzf!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png 848w, /__u/substackcdn.com/image/fetch/$s_!ULzf!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ULzf!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a8c4fcd-2b4d-4a5f-b625-0cfb9a84d9f7_1448x1086.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In a previous article, I benchmarked Qwen3.8 27B using its highest reasoning effort, <code>xhigh</code>. That shows the model at its best, but it does not tell us whether <code>xhigh</code> is the most efficient way to run it. Qwen3.8 27B also supports <code>medium</code> and <code>low</code> reasoning efforts, while thinking can be disabled completely.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;04c290d2-43b3-458e-893b-a5645d4adc2a&quot;,&quot;caption&quot;:&quot;Muse Glimmer was released only a few days ago as I write this, yet it has already been largely overshadowed by Qwen3.8 27B.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwen3.8 27B and Muse Glimmer Benchmarks: Accuracy, Token Efficiency and Memory Use&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-08-20T17:45:19.487Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!elrh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F219ba0ed-aa99-4654-a32a-affa7ad0c685_1520x574.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwen38-27b-and-muse-glimmer-benchmarks&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:211755468,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>In this article, I compare these four configurations to see how reasoning effort affects both accuracy and token usage. I ran a large number of tasks and generated hundreds of millions of tokens. For each mode, I measure how well Qwen3.8 performs and how many tokens it uses to get there.</p><p><strong>How much accuracy do we gain by letting Qwen3.8 think more, and how many extra tokens does that cost?</strong> </p><p><code>xhigh</code> should give the best results, but it may also generate much longer reasoning traces. If <code>medium</code> or <code>low</code> can retain most of the accuracy while using significantly fewer tokens, they may be better choices in practice.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>This is especially relevant for local inference. More reasoning tokens mean more generation time, more compute, and higher memory consumption. A small improvement in accuracy may not be worth it if the model requires two or three times as many tokens to achieve it.</p><p>I will first explain how Qwen3.8 reasoning effort works, how low, medium, and xhigh are implemented in the chat template, and how to set them correctly with vLLM and llama.cpp. I will then compare their accuracy and token usage across several benchmarks, before briefly putting the results in context against Qwen3.6 and Muse Glimmer.</p><blockquote><p><strong>Acknowledgments</strong></p><p><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=qwen38thinking">Verda</a><span> provided the RTX Pro 6000s to run the experiments described in this article.</span></p><p>Verda is a full-stack AI cloud, built for high-performance inference, training, and agentic workloads, with data privacy and sustainability at its core.</p><p><span>You can check them out </span><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=qwen38thinking">here</a><span>. There is a </span><strong>$50 coupon</strong><span> that you can redeem in your Verda account, after provisioning it with $5, to try their GPUs.</span></p><p><strong>Coupon code:</strong><span> </span><code>KAITCHUP-50</code></p><p><span>Follow </span><a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a><span>.</span></p><p><em><span>Note: I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon.</span></em></p></blockquote><p><em>Note: These experiments focus on non-agentic tasks, where each prompt is evaluated as a standalone problem. I am testing long-horizon agentic coding separately, and the runs are almost complete, but there is still a large amount of data to analyze and verify. I plan to cover those results in a separate article, probably next week.</em></p><h2>How Qwen3.8 Reasoning Effort Works</h2><p>Qwen3.8 supports three <code>reasoning_effort</code> values: <code>low</code>, <code>medium</code>, and <code>xhigh</code>. Thinking can also be disabled with <code>enable_thinking=false</code>. The goal of these settings is to control how much reasoning the model does before producing its final answer. <code>low</code> asks the model to keep its reasoning short and focused, while <code>xhigh</code> encourages more careful and thorough reasoning. <code>medium</code> sits between the two.</p>
      <p>
          <a href="/__u/kaitchup.substack.com/p/qwen38-27b-reasoning-benchmarks-off">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Ornith-1.5, LFM2.5 QAD, and a 10 GB Qwen3.8 27B]]></title><description><![CDATA[The Weekly Kaitchup #156]]></description><link>https://kaitchup.substack.com/p/ornith-15-lfm25-qad-and-a-10-gb-qwen38</link><guid isPermaLink="false">https://kaitchup.substack.com/p/ornith-15-lfm25-qad-and-a-10-gb-qwen38</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 22 Aug 2026 02:41:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>Last week was mostly about new models. This week is a little different.</p><p>The three releases I found most interesting are more about getting substantially more out of existing models through better post-training and quantization:</p><ul><li><p><strong>Ornith-1.5:</strong> a new generation of agentic models trained with a more complete self-improvement loop.</p></li><li><p><strong>LFM2.5 QAD:</strong> Liquid AI is using Quantization-Aware Distillation to make tiny Q4_0 GGUFs much more accurate.</p></li><li><p><strong>Escha-W2 for Qwen3.8 27B:</strong> a 10.15 GB version of Qwen3.8 27B that, according to Escha Labs, retains all of the FP8 model&#8217;s benchmark performance.</p></li></ul><p>There is a common theme here.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>We are getting increasingly sophisticated techniques for improving models <em>after</em> pre-training. Better RL can turn a relatively small base model into a much stronger agent. Training can compensate for quantization. And increasingly aggressive low-bit formats are starting to preserve much more accuracy than I would have expected a year ago.</p><blockquote><h3>Coming Soon in The Kaitchup</h3><p>I&#8217;m nearly done with my article on inference speed and latency, covering Muse Glimmer, Qwen3.8 27B, and Dflash-2. It&#8217;s a direct follow-up to the article I published yesterday:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;6ccb58ee-4f5b-47f5-b45a-d8d64acec46b&quot;,&quot;caption&quot;:&quot;Muse Glimmer was released only a few days ago as I write this, yet it has already been largely overshadowed by Qwen3.8 27B.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwen3.8 27B and Muse Glimmer Benchmarks: Accuracy, Token Efficiency and Memory Use&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-08-20T17:45:19.487Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!elrh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F219ba0ed-aa99-4654-a32a-affa7ad0c685_1520x574.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwen38-27b-and-muse-glimmer-benchmarks&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:211755468,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:10,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>I also now have all the results I need to compare the accuracy and token efficiency of Qwen3.8 27B across its off, low, medium, and xhigh thinking modes. One thing I can already say: <strong>xhigh is much better.</strong></p><p>As for my agentic coding experiments with Qwen3.8, comparing different agent harnesses across different thinking modes, I&#8217;m about halfway through.</p></blockquote><div><hr></div><h2>Ornith-1.5: Self-Scaffolding Becomes Self-Improvement</h2><p>I <a href="/__u/kaitchup.substack.com/p/this-week-in-open-models-tiny-lfm25">discussed Ornith-1.0 a few weeks ago</a>.</p><p>For agentic coding, the model operates inside a scaffold: the prompts, tools, memory strategy, retry logic, and general orchestration used to solve a task. Normally, humans design this scaffold and then keep it fixed during reinforcement learning.</p><p>Ornith-1.0 instead allowed the model to generate a task-specific scaffold and rewarded both the scaffold and the resulting solution. And since it works very well, <a href="https://huggingface.co/collections/ornith-ai/ornith-15">Ornith-1.5 </a>pushes this idea further. </p><p>The model no longer only creates the scaffold. It can now generate:</p><ul><li><p>the <strong>task</strong> it will learn from,</p></li><li><p>a <strong>scaffold</strong> for solving that task,</p></li><li><p>and the <strong>solution rollout</strong> itself.</p></li></ul><p>This effectively closes the loop.</p><p>Instead of humans providing a fixed collection of coding tasks and letting the model optimize how it solves them, the model can propose new tasks that are useful for training, build a strategy to solve them, execute that strategy, and then learn from the resulting reward.</p><p>I find this much more interesting than the benchmark tables.</p><p>Generating synthetic training tasks is obviously not new. What is interesting here is combining task generation with scaffold generation and reinforcement learning in one loop.</p><p>In principle, this allows the training distribution itself to evolve as the model improves.</p><p>If the model becomes good at one type of task, it can generate increasingly difficult variants rather than continuing to spend compute solving things it already knows how to do. What I&#8217;m curious about is the training cost. Because I assume finding new difficult tasks for the model that actually improve the reward is not that easy and could take many training steps</p><h3>Three Models: 9B, 35B-A3B, and 397B</h3><p>The Ornith-1.5 family currently contains:</p><ul><li><p>a <strong>9B dense model</strong></p></li><li><p>a <strong>35B-A3B MoE</strong></p></li><li><p>a <strong>397B MoE</strong></p></li></ul><p><em>Note: Did you notice? No mention of an &#8220;Ornith 1.5 31B&#8221;. The release of Ornith 1.0 was supposed to include a model based on Gemma 4 31B. This model was never published. My assumption: it&#8217;s extremely difficult to fine-tune Gemma 4 for agentic coding, and the Ornith 1.0 31B was never good enough. We need a Gemma 4.5 that's good at agentic tasks.</em></p><p>The 9B is particularly impressive:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!B0j6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!B0j6!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png 424w, /__u/substackcdn.com/image/fetch/$s_!B0j6!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png 848w, /__u/substackcdn.com/image/fetch/$s_!B0j6!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png 1272w, /__u/substackcdn.com/image/fetch/$s_!B0j6!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!B0j6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Ornith 1.5 9B Benchmark Results&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Ornith 1.5 9B Benchmark Results" title="Ornith 1.5 9B Benchmark Results" srcset="/__u/substackcdn.com/image/fetch/$s_!B0j6!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png 424w, /__u/substackcdn.com/image/fetch/$s_!B0j6!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png 848w, /__u/substackcdn.com/image/fetch/$s_!B0j6!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png 1272w, /__u/substackcdn.com/image/fetch/$s_!B0j6!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2fab2d-bd54-420a-ab07-a0dde3df6dd6_2787x1568.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I&#8217;ll publish quantized versions next week!</p><p>It&#8217;s safe to assume they are also preparing a version, Ornith 2.0?, based on Qwen3.8.</p><div><hr></div><h2>LFM2.5 QAD</h2><p>Liquid AI released new Q4_0 GGUFs for:</p><ul><li><p><a href="https://huggingface.co/LiquidAI/LFM2.5-230M-GGUF">LFM2.5-230M</a></p></li><li><p><a href="https://huggingface.co/LiquidAI/LFM2.5-350M-GGUF">LFM2.5-350M</a></p></li><li><p><a href="https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF">LFM2.5-1.2B-Instruct</a></p></li><li><p><a href="https://huggingface.co/LiquidAI/LFM2.5-2.6B-GGUF">LFM2.5-2.6B</a></p></li></ul><p>They were trained with <strong>Quantization-Aware Distillation</strong>, or QAD.</p><p>I briefly discussed <a href="/__u/kaitchup.substack.com/p/this-week-arcee-trinity-and-quantization">QAD earlier this year when NVIDIA used it for Nemotron 3 Nano</a>.</p><blockquote><p><strong>QAD</strong></p><p>With normal post-training quantization, we train the model at high precision and compress it afterward. If quantization changes the model&#8217;s outputs, we simply accept that loss.</p><p>QAD adds another training stage. You keep the high-precision model as the teacher and train the quantized model to reproduce its output distribution.</p><p>The loss can be viewed as:</p><p><code>KL(teacher || quantized student)</code></p><p>So instead of asking the quantized model to relearn the original tasks, we ask it to behave as similarly as possible to the unquantized model.</p><p>It only requires the original model, the quantized student, and suitable data for distillation.</p></blockquote><p>Quantization is generally less forgiving as models get smaller.</p><p>A 30B model often has enough redundancy that reducing weight precision to four bits causes little damage. With models containing only a few hundred million parameters, every parameter is doing more work.</p><p>Liquid reports that the new QAD Q4_0 checkpoints retain approximately:</p><ul><li><p>97.1% of BF16 performance for the 230M</p></li><li><p>96.5% for the 350M</p></li><li><p>97.4% for the 1.2B</p></li><li><p>96.6% for the 2.6B</p></li></ul><p>They significantly improve over Liquid&#8217;s previous PTQ Q4_0 checkpoints that I evaluated a few weeks ago:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;e750c673-63e9-4d5d-9dc4-4e027b3098f9&quot;,&quot;caption&quot;:&quot;Tiny language models that can run quickly with very limited memory are one of Liquid AI&#8217;s specialties. And while models with only a few hundred million parameters obviously cannot rival today&#8217;s multi-billion-parameter models, tiny models are now capable enough for many practical tasks. They can make tool calls, retain some basic world knowledge, and follow instructions surprisingly well.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;LFM2.5 230M and 350M: How Accurate Are the GGUF Versions?&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-07-08T03:20:47.548Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!2bDO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/lfm25-230m-and-350m-how-accurate&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:205072090,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:10,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>I ran the same evaluation settings on these QAD versions and independently confirmed LiquidAI&#8217;s conclusions: the new QAD Q4_0 versions are significantly better.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!yE63!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!yE63!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!yE63!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!yE63!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!yE63!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!yE63!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg" width="2361" height="1254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:2361,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:158462,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/212093788?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51ca8913-faf8-4c05-a8c5-da05ed5da7e1_2361x1254.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!yE63!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!yE63!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!yE63!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!yE63!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286bf8c3-7a9e-4dc1-a1fe-28993ef9af3a_2361x1254.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>But why focus on Q4_0?</h3><p>There are now many sophisticated GGUF quantization formats. Q4_0 is not one of them. It is old, simple, and widely supported. But that simplicity is also an advantage.</p><p>Hardware vendors and inference runtimes already know how to execute Q4_0 efficiently, including on CPUs and Arm devices.</p><p>The QAD versions match the quality of larger quantizations while decoding faster.</p><p>For the 230M and 350M models, the QAD Q4_0 checkpoints are between 4% and 33% faster, depending on the hardware.</p><div><hr></div><h2>Escha-W2: Qwen3.8 27B in 10.15 GB</h2><p>Two weeks ago, <a href="/__u/kaitchup.substack.com/i/209166534/escha-w2-qwen36-35b-compressed-to-123-gb">I wrote about Escha Labs&#8217; 2-bit version of Qwen3.6-35B-A3B</a>.</p><p>Now, Escha Labs (about which we don&#8217;t know much) has done the same thing with Qwen3.8 27B.</p><p>Escha reports:</p><ul><li><p><strong>10.15 GB</strong> for the entire model on disk</p></li><li><p><strong>82.6 tokens/s</strong> single-stream generation on an RTX 5090</p></li><li><p>approximately <strong>100% of FP8 benchmark performance</strong> averaged over the eight evaluations they have completed so far</p></li></ul><p>Like their Qwen3.6 release, the model runs through a custom SGLang-based runtime.</p><p>They also say the model and runtime are Apache-2.0 licensed and that GGUF versions are coming.</p><p>10.15 GB for a 27B dense model is extremely small.</p><h3>Almost No Benchmark Loss?</h3><p>This is the part I want to test carefully. Escha says the new W2 checkpoint averages around 100% of the FP8 model across the eight benchmarks they have evaluated.</p><p>There is even a surprising LiveCodeBench v6 result:</p><ul><li><p>Escha W2: <strong>86.81</strong></p></li><li><p>FP8: <strong>85.16</strong></p></li></ul><p>Is the low-bit model better than FP8? Obviously not. The problem is that their evaluation was too cheap: a 28K max output-token limit, and only ~17% of the benchmark was actually run. <em>Note: and most of the other benchmarks they used don&#8217;t even generate any tokens and rely on the model&#8217;s logits. </em></p><p>For Qwen3.8, evaluating quantization properly requires the full 262K-token context, because the model uses that context for reasoning. Quantization error compounds at every decoding step, so seeing little or no degradation early in the generation is encouraging, but it&#8217;s not enough to draw a conclusion.</p><p>LiveCodeBench and the other benchmarks they used, like GPQA, also have high variance. I run them four times before publishing my numbers. Yes, that&#8217;s not cheap, but it&#8217;s necessary.</p><p>My own evaluation of Escha-W2 is in the queue. I expect it to be significantly better than the other 2&#8211;3-bit versions of Qwen3.8, but certainly not as good as FP8. Hopefully, I&#8217;m wrong.</p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisioning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/kaitchup.substack.com/subscribe"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p>]]></content:encoded></item><item><title><![CDATA[Qwen3.8 27B and Muse Glimmer Benchmarks: Accuracy, Token Efficiency and Memory Use]]></title><description><![CDATA[Qwen3.8 is remarkably capable, while Muse Glimmer reveals a very different approach to reasoning and memory efficiency.]]></description><link>https://kaitchup.substack.com/p/qwen38-27b-and-muse-glimmer-benchmarks</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-27b-and-muse-glimmer-benchmarks</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Thu, 20 Aug 2026 17:45:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!elrh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F219ba0ed-aa99-4654-a32a-affa7ad0c685_1520x574.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Muse Glimmer was released only a few days ago as I write this, yet it has already been largely overshadowed by Qwen3.8 27B.</p><p>The difference in accuracy between the two models is striking, especially given their similar parameter counts and dense architectures. Muse Glimmer struggles in particular with long-horizon agentic coding, while Qwen3.8 27B performs remarkably well, even across different agents and with thinking effort set to medium.</p><p>Still, <a href="/__u/kaitchup.substack.com/p/muse-glimmer-metas-30b-model-built">as we saw in a previous article, Muse Glimmer may retain a few important advantages</a>:</p><ul><li><p>It has a relatively short maximum context length. That is a limitation in itself, but it also suggests that the model was not trained to produce the extremely long reasoning traces Qwen3.8 can generate.</p></li><li><p>Its KV-cache memory consumption is very low: it uses roughly <strong>4 times less memory</strong> than Qwen3.8.</p></li></ul><p>Taken together, these characteristics make Muse Glimmer considerably more memory-efficient than Qwen3.8 27B, despite being slightly larger.</p><p>In this article, I first present my own accuracy results for Muse Glimmer and Qwen3.8 27B across a broad range of tasks, with both models evaluated at <strong>xhigh thinking effort</strong>. The benchmarks cover coding, world knowledge, difficult mathematics, instruction following, and more.</p><p>I then look beyond raw accuracy to examine token efficiency: how many tokens each model typically needs to solve a problem, and how much memory each consumes once the length of its reasoning traces is taken into account. This gives a more complete picture of the trade-off between the two models than parameter count or benchmark accuracy alone.</p><blockquote><p><strong>Acknowledgments</strong></p><p><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=qwen38">Verda</a><span> provided the RTX Pro 6000s to run the experiments described in this article.</span></p><p>Verda is a full-stack AI cloud, built for high-performance inference, training, and agentic workloads, with data privacy and sustainability at its core.</p><p><span>You can check them out </span><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=qwen38">here</a><span>. There is a </span><strong>$50 coupon</strong><span> that you can redeem in your Verda account, after provisioning it with $5, to try their GPUs.</span></p><p><strong>Coupon code:</strong><span> </span><code>KAITCHUP-50</code></p><p><span>Follow </span><a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a><span>.</span></p><p><em><span>Note: I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon.</span></em></p></blockquote><p><em><strong>Note:</strong> For this article, I evaluated Qwen3.8 only in xhigh thinking mode. A follow-up will compare its low, medium, xhigh, and disabled thinking modes. I also plan to publish a separate article focused entirely on agentic coding with Qwen3.8, as there is a great deal to explore there. Results for quantized versions of both Muse Glimmer and Qwen3.8 will come later, as will a comparison of their inference speeds.</em></p><h2>Accuracy: Qwen3.8 Comes Out on Top, While Muse Glimmer Behaves Very Differently</h2><p>My results confirm that Qwen3.8 27B at xhigh thinking effort is a substantial improvement over Qwen3.6 27B:</p>
      <p>
          <a href="/__u/kaitchup.substack.com/p/qwen38-27b-and-muse-glimmer-benchmarks">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Qwen3.8 27B, Nemotron 3.5, Muse, DeepSeek V4 Pro: A Huge Week for Open-Weight AI]]></title><description><![CDATA[The Weekly Kaitchup #155]]></description><link>https://kaitchup.substack.com/p/qwen38-27b-nemotron-35-muse-deepseek</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-27b-nemotron-35-muse-deepseek</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Fri, 14 Aug 2026 20:13:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>This week was unusually rich in open-weight releases.</p><p>Meta released Muse Glimmer. Qwen finally released the weights of Qwen3.8 2.4T. NVIDIA released Nemotron 3.5 Lightning. DeepSeek just pushed a major V4 Pro update.</p><blockquote><p>And, just as I was finishing this article, <strong><a href="https://huggingface.co/Qwen/Qwen3.8-27B">Qwen3.8 27B</a> was released</strong>. It was worth the wait. According to Qwen&#8217;s evaluations, it improves very significantly over Qwen3.6 27B essentially everywhere, with particularly spectacular gains on long-horizon agentic coding: DeepSWE jumps from 13.3 to 42.2, while QwenSWEBench goes from 49.3 to 79.0. The 27B model also beats Qwen3.7-Plus on many coding and agentic benchmarks, and even beats Opus4.6 Max on QwenSWEBench, CoWorkBench, and LiveCodeBench. At this point, Qwen seems so far ahead that catching up is becoming difficult for everyone else.</p><p>There are also a few interesting changes compared with Qwen3.6. <code>preserve_thinking</code>, which Qwen3.6 introduced as an option for keeping reasoning traces across turns, is now <strong>enabled by default</strong>. Qwen3.8 also introduces <code>reasoning_effort</code>, with <code>low</code>, <code>medium</code>, and <code>xhigh</code> levels to trade reasoning depth against cost. Qwen now recommends <code>temperature=1.0</code> and <code>top_p=0.95</code> for thinking mode generally, and, for very long agentic runs, recommends allowing up to 262K tokens for reasoning and 131K for the final answer. I&#8217;ll publish a full analysis of Qwen3.8 27B within the next few days, with a separate look at its quantized versions later.</p><p>Note: <em>I think this &#8220;preserve_thinking&#8221; (which is a feature we can find in other models) should be carefully evaluated, especially its impact on inference cost. Qwen3.8 is a crazy thinker (they recommend allowing &#8220;up to 262K tokens for reasoning&#8221;). So if you preserve reasoning for each turn, the context may grow to millions of tokens. In practice though, for tool calls, reasoning is often much shorter and more/better reasoning may yield fewer turns, so I guess it can work. But I&#8217;m really interested to know how important it is, in terms of accuracy/efficiency, to preserve the reasoning traces. <strong>KV cache quantization will likely be very important</strong>.</em></p></blockquote><p>Finding enough time to cover all of this properly is hard. I don&#8217;t want to publish articles that simply repeat benchmark tables and model cards. A proper analysis means testing accuracy, token efficiency, memory, speed, and then doing it again for several quantized versions.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>For Glimmer, I already published an article covering the model and its architecture:</p><ul><li><p><a href="/__u/kaitchup.substack.com/p/muse-glimmer-metas-30b-model-built">Muse Glimmer: Meta&#8217;s 30B Model Built for Efficient Inference</a></p></li></ul><p>I also already have speed measurements for all the non-GGUF quantized versions I&#8217;m testing.</p><p>At the moment, essentially all my compute capacity, several RTX Pro 6000 GPUs provided by <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> is being used to evaluate the quantized versions that run through vLLM: NVFP4, INT4, and my own mixed-precision versions.</p><p>You can already find those here:</p><ul><li><p><a href="https://huggingface.co/collections/kaitchup/quantized-muse-glimmer">My quantized Muse Glimmer collection</a> (they all run, but consider all of them very bad until I publish some official evaluations)</p></li></ul><p>I&#8217;ll probably publish a short article focused only on speed first, since those results are already ready.</p><p>Then, expect perhaps two deeper articles comparing Glimmer with Gemma 4 and Qwen3.6/3.8, and another article focused specifically on the quantized versions. I may include GGUF versions in that one as well, but I&#8217;m not sure yet. There is only so much GPU time, and writing time, in a week.</p><p>So, for this Weekly Kaitchup, I&#8217;ll focus on the three releases that I won&#8217;t cover in more detail:</p><ul><li><p>Qwen3.8 2.4T: we finally know what is inside it, so I can check how well my predictions held up.</p></li><li><p>Nemotron 3.5 Lightning: essentially the new Nano-sized Nemotron.</p></li><li><p>DeepSeek V4 Pro 0813: a very large update that makes the word &#8220;Preview&#8221; much more meaningful.</p></li></ul><h2>Qwen3.8 2.4T: How Good Were My Predictions?</h2><p>A few weeks ago, before Alibaba disclosed the architecture, I tried to estimate what hardware would be required to run Qwen3.8 2.4T:</p><ul><li><p><a href="/__u/kaitchup.substack.com/p/qwen38-what-hardware-will-you-need">Qwen3.8: What Hardware Will You Need to Run Alibaba&#8217;s 2.4T Model?</a></p></li></ul><p>At the time, we knew only one particularly important number: <strong>2.4 trillion parameters</strong>.</p><p>Everything else had to be inferred.</p><p>Now we have the weights and architecture, so it&#8217;s a good opportunity to see what I got right and what should be corrected.</p><p>Qwen3.8-2.4T-A95B has 2.4T total parameters and 95B activated parameters. It has 92 layers and uses the hybrid architecture introduced with Qwen3.5, alternating Gated DeltaNet blocks with periodic full attention. Its MoE has 512 experts, with 10 routed experts plus one shared expert activated per token. Native context is 262K, extendable to just over one million tokens.</p><p>The downloadable 2.4T checkpoint is also more limited than the hosted Qwen3.8-Max: it is text-only and requires thinking mode, while the hosted Max version adds vision, non-thinking mode, built-in tools, and a 1M default context.</p><p>So, how did my predictions go?</p><h3><strong>&#9989; It is an extremely sparse MoE</strong></h3><p>This one was the most important and easiest assumption.</p><p>I wrote that a dense 2.4T model would make very little sense for inference and that Qwen3.8 was almost certainly going to be a highly sparse MoE.</p><p>That was correct.</p><p>Only 95B of 2.4T parameters are active for each token, or roughly 4% of the entire model.</p><h3><strong>&#9989; Nearly 99% of the parameters are in the routed experts</strong></h3><p>I assumed that approximately <strong>99% of Qwen3.8&#8217;s parameters would live inside routed experts</strong>, based largely on Kimi K2 and other very large MoEs.</p><p>This turned out to be close.</p><p>Using the released dimensions, 512 experts, 92 MoE layers, hidden size 8192, and expert intermediate size 2048, the routed expert matrices account for roughly <strong>2.371 trillion parameters</strong>, or about <strong>98.8% of the entire 2.4T model</strong>.</p><p>My estimate in the original article was 2.374T routed-expert parameters.</p><h3><strong>&#9989; The routing sparsity</strong></h3><p>I used Kimi K2 as a proxy.</p><p>Kimi routes each token through 8 of 384 experts, or about 2.08% of its routed experts.</p><p>Qwen3.8 uses 10 of 512, or about <strong>1.95%</strong>.</p><p>So, while the architectures are different, the routing sparsity I used for the estimate was close enough.</p><h3><strong>&#10060; I estimated ~75B active parameters</strong></h3><p>This is the main miss.</p><p>Using Kimi K2 as a proxy, I estimated that Qwen3.8 might require compute roughly comparable to a 75B dense model per token, before routing and communication overhead.</p><p>The actual number is <strong>95B activated parameters</strong>.</p><p>That is around 27% higher than my estimate.</p><p>So, the general prediction, a 2.4T model with only a small fraction active, was correct, but Qwen3.8 is somewhat more computationally expensive per token than I expected.</p><h3><strong>&#10060; Around 1.4&#8211;1.5 TB with experts-only NVFP4</strong></h3><p>Wrong but close.</p><p>I estimated <strong>1.387 TB</strong> for a version where the routed experts are stored in NVFP4 while the rest of the network remains at higher precision.</p><p>A community <a href="https://huggingface.co/RadixArk/Qwen3.8-2.4T-A95B-NVFP4">NVFP4 checkpoint from RadixArk</a> now does exactly that: only the routed-expert linear layers are quantized to NVFP4 while the other components retain their original precision.</p><p>Its repository is <strong>1.48 TB</strong>.</p><p>So my estimate was around 6&#8211;7% too low, but the overall hardware conclusion was correct.</p><p>A single 8xB300 node is enough to deploy this model.</p><h2>Nemotron 3.5 Lightning: Nano, Updated</h2><p>NVIDIA also released <strong><a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4">Nemotron 3.5 Lightning 30B-A3B</a></strong> this week.</p><p>If the naming feels slightly confusing, Nemotron 3.5 Lightning is effectively the updated model occupying Nemotron 3 Nano&#8217;s position in the family.</p><p>I previously tested the original Nano here:</p><ul><li><p><a href="/__u/kaitchup.substack.com/p/nemotron-3-nano-fast-cheap-and-surprisingly">Nemotron 3 Nano: A Very Fast Model That Doesn&#8217;t Think Too Much</a></p></li></ul><p>And I covered the much larger Super model here:</p><ul><li><p><a href="/__u/kaitchup.substack.com/p/nemotron-3-super-1m-tokens-small">Nemotron 3 Super: 1M Tokens, Small KV Cache</a></p></li></ul><p>Lightning still uses the same general idea of interleaving Mamba-2 processing with selected attention layers and sparse experts, and it supports context lengths up to 1M tokens.</p><p>So this is not a new size class. It is much closer to a refreshed Nano.</p><p>But NVIDIA has brought several ideas developed across the rest of the Nemotron 3 family back into this smaller model.</p><p>The biggest additions are around inference and agents.</p><p><strong>Multi-Token Prediction is now part of the model training</strong>, followed by an additional MTP-boosting phase. NVIDIA also releases DFlash and DSpark draft models for speculative decoding, so there are multiple ways to accelerate generation depending on concurrency and hardware.</p><p>This is interesting because MTP was already one of the important additions in Nemotron 3 Super.</p><p>NVIDIA also says that 3.5 Lightning received <strong>harness-optimized training</strong> for agent workloads. Large Nemotron models can do the expensive planning and reasoning, while Lightning is supposed to execute the many smaller calls produced by long-running agents.</p><p>There are BF16 and NVFP4 checkpoints, and NVIDIA says the model can reach up to 4x the output speed of similarly sized models. On its PinchBench test, NVIDIA reports roughly 86% accuracy while completing 10,000 tasks around 30% faster than Qwen3.6 35B at comparable accuracy. Those are NVIDIA&#8217;s numbers, so I would still like to reproduce the speed/accuracy trade-off independently.</p><h2>DeepSeek V4 Pro 0813: &#8220;Preview&#8221; Really Meant Preview</h2><p>Finally, DeepSeek (quietly) released <strong><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813">DeepSeek-V4-Pro-0813</a></strong>.</p><p>When DeepSeek introduced V4 in April, both V4 Pro and V4 Flash were explicitly called <strong>Preview</strong> models. The architecture was extremely innovative.</p><p>V4 Pro has around 1.6T parameters with 49B active and combines Compressed Sparse Attention with Heavily Compressed Attention. DeepSeek also introduced Manifold-Constrained Hyper-Connections and trained with the Muon optimizer. At a one-million-token context, DeepSeek reported that V4 Pro required only <strong>27% of the single-token inference FLOPs and 10% of the KV-cache memory</strong> of DeepSeek V3.2.</p><p>Accuracy was the problem.</p><p>The original V4 Pro was rather underwhelming for such a huge model, particularly on agentic tasks. Then DeepSeek released Flash-0731, and the Flash model was suddenly beating Pro Preview.</p><p>I wrote about that two weeks ago:</p><ul><li><p><a href="/__u/kaitchup.substack.com/p/deepseek-v4-flash-0731-and-inkling">DeepSeek-V4-Flash-0731 and Inkling Small: Smaller, but Better?</a></p></li></ul><p>At the time, I wrote that Flash-0731 was &#8220;a very promising sign for the next V4 Pro update.&#8221; That update is here now. The model now also exposes three reasoning-effort settings: low, high, and max.</p><p>The accuracy jump is much more important.</p><p>A few examples from DeepSeek&#8217;s own evaluation table:</p><ul><li><p>DeepSWE: 12.8 (it was <strong>lower than Qwen3.8 27B!</strong>) &#8594; 62.7<br>Current Flash-0731: 54.4</p></li><li><p>Terminal Bench 2.1: 72.1 &#8594; 87.9<br>Current Flash-0731: 82.7</p></li><li><p>NL2Repo: 38.5 &#8594; 61.5<br>Current Flash-0731: 54.2</p></li><li><p>Cybergym: 52.7 &#8594; 83.3<br>Current Flash-0731: 76.7</p></li><li><p>Toolathlon Verified: 55.9 &#8594; 74.1<br>Current Flash-0731: 70.3</p></li><li><p>AutomationBench: 12.8 &#8594; 31.8<br>Current Flash-0731: 25.1</p></li></ul><p><strong>Pro-0813 is now ahead of the current Flash-0731 model</strong>.</p><p>DeepSWE is the most spectacular improvement: <strong>12.8 to 62.7</strong>.</p><p>As I discussed this week, agentic evaluations can move significantly depending on the harness, configuration, allowed steps, tools, and other details. Independent testing remains necessary.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;497985b7-de19-49c9-b233-74e6f419582c&quot;,&quot;caption&quot;:&quot;Coding benchmarks increasingly test more than whether a model can write a correct function. They ask an agent to explore an unfamiliar repository, run commands, edit files, interpret failures, test its work, recover from mistakes, and decide when the job is actually finished.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Laguna S 2.1: How Agent Harnesses and Inference Budgets Shape Coding Performance&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-08-13T16:11:23.125Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!hTOO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/laguna-s-21-how-agent-harnesses-and&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:210853645,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:3,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisioning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/kaitchup.substack.com/subscribe"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I&#8217;ll probably add a B200 to cover Qwen3.8 27B better. Muse Glimmer has a short native max context length (131K tokens) and a small KV cache. An RTX Pro 6000 is enough. But Qwen3.8 27B&#8217;s KV cache is much larger, with a native context twice as large. 96 GB at 1.7 TB/sec is not enough with high concurrency.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Laguna S 2.1: How Agent Harnesses and Inference Budgets Shape Coding Performance]]></title><description><![CDATA[Testing long-horizon coding performance on DeepSWE and Terminal-Bench 2.1]]></description><link>https://kaitchup.substack.com/p/laguna-s-21-how-agent-harnesses-and</link><guid isPermaLink="false">https://kaitchup.substack.com/p/laguna-s-21-how-agent-harnesses-and</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Thu, 13 Aug 2026 16:11:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hTOO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hTOO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hTOO!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!hTOO!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!hTOO!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hTOO!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hTOO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png" width="640" height="360" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:640,&quot;bytes&quot;:2208006,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/210853645?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!hTOO!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!hTOO!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!hTOO!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hTOO!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Coding benchmarks increasingly test more than whether a model can write a correct function. They ask an agent to explore an unfamiliar repository, run commands, edit files, interpret failures, test its work, recover from mistakes, and decide when the job is actually finished.</p><p>That makes benchmark performance a product of at least three things: the underlying model, the agent harness driving it, and the amount of time and context the system is allowed to consume.</p><p>I evaluated Laguna S 2.1 on two long-horizon benchmarks, DeepSWE and Terminal-Bench 2.1, to understand how much these surrounding conditions affect both accuracy and cost. In particular, I wanted to see how the model behaves outside Poolside&#8217;s native agent harness, how much additional inference-time compute improves results, and what happens when an agent is given enough time to keep working after it gets stuck.</p><p>The results show a capable model whose accuracy can improve substantially when it is given more room to work. They also expose the other side of long-horizon agents: unsuccessful trajectories can consume enormous amounts of computation without getting any closer to a correct solution.</p><p>Perhaps most importantly, the experiments reinforce that model configuration is only part of the story. The agent wrapped around the model matters enormously.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In this article, I compare Laguna S 2.1 across DeepSWE and Terminal-Bench 2.1, examine how turn limits, timeouts, and token budgets affect accuracy and cost, analyze the main failure modes, and explain why the agent harness itself has become a critical part of coding-model evaluation.</p><blockquote><p><strong>Acknowledgments</strong></p><p><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=lagunas21">Verda</a><span> provided the B300 to run the experiments described in this article.</span></p><p>Verda is the full-stack frontier AI cloud, built for high-performance inference, training, and agentic workloads with data privacy and sustainability at its core.</p><p><span>You can check them out </span><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=lagunas21">here</a><span>. There is a </span><strong>$50 coupon</strong><span> that you can redeem in your Verda account, after provisioning it with $5, to try their GPUs.</span></p><p><strong>Coupon code:</strong><span> </span><code>KAITCHUP-50</code></p><p><span>Follow </span><a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a><span>.</span></p><p><em><span>Note: I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon.</span></em></p></blockquote><h2>Agentic coding evaluation setup</h2><p>I deliberately evaluated <a href="https://huggingface.co/poolside/Laguna-S-2.1">Laguna S 2.1</a> outside Poolside&#8217;s native pool harness.</p><p>For DeepSWE, I used mini-swe-agent.</p><p>For Terminal-Bench 2.1, I used the terminus-2 agent.</p><p>Poolside&#8217;s own Laguna S 2.1 evaluations use its pool agent harness, while the public DeepSWE leaderboard normally uses mini-swe-agent. My numbers should be read as measurements of Laguna operating inside these particular third-party agent setups, not as direct reproductions of Poolside&#8217;s published benchmark results.</p>
      <p>
          <a href="/__u/kaitchup.substack.com/p/laguna-s-21-how-agent-harnesses-and">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Muse Glimmer: Meta’s 30B Model Built for Efficient Inference]]></title><description><![CDATA[Inside Meta&#8217;s 30B local reasoning model and its tiny KV cache]]></description><link>https://kaitchup.substack.com/p/muse-glimmer-metas-30b-model-built</link><guid isPermaLink="false">https://kaitchup.substack.com/p/muse-glimmer-metas-30b-model-built</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Tue, 11 Aug 2026 02:01:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tgN2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!tgN2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!tgN2!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!tgN2!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!tgN2!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tgN2!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!tgN2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png" width="672" height="378" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:672,&quot;bytes&quot;:2313372,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/210643086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!tgN2!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!tgN2!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!tgN2!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tgN2!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">This isn&#8217;t quite as fun as generating images of llamas, but I&#8217;ll take it.</figcaption></figure></div><p>Meta is back, and it&#8217;s great.</p><p>For a while, the center of gravity in open-weight AI moved toward Qwen, DeepSeek, Google&#8217;s Gemma family, and a growing number of Chinese labs. Muse Glimmer gives Meta a genuinely interesting entry again.</p><p>Muse Glimmer is a ~30B-parameter multimodal reasoning model from Meta Superintelligence Labs, released under Apache 2.0.  </p><ul><li><p><a href="https://huggingface.co/collections/meta-models/muse-glimmer">Muse Glimmer</a></p></li></ul><p>Glimmer is distilled from Meta&#8217;s much larger Muse Spark model, supports tool use and multimodal inputs, has a 131K context window, ships with official 4-bit quantization, and targets consumer hardware. It&#8217;s clearly positioned as a competitor to Gemma 4 and Qwen3.6 open dense models.</p><p><em>How good is Glimmer compared with Gemma 4 and Qwen3.6?</em></p><p>The benchmark results put Glimmer surprisingly close to Qwen3.6. But benchmark scores alone tell us nothing about efficiency.</p><p>Glimmer uses an exceptionally lightweight attention design, and its relatively short 131K context window suggests it was not trained around the extremely long reasoning traces Qwen3.6 can produce, sometimes stretching beyond 100K tokens.</p><p>That opens up an interesting possibility: Glimmer may deliver performance close to Qwen3.6 while approaching Gemma 4&#8217;s token efficiency and using significantly less memory overall.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In this article, I&#8217;ll first look at what makes Muse Glimmer interesting as a local model, then break down its architecture in detail, including its hybrid attention pattern, extreme GQA, gated attention, positional encoding, multimodal stack, and speculative decoding. I&#8217;ll then dig into its KV-cache efficiency, with a direct 131K-context comparison against Gemma 4 31B and Qwen3.6-27B. </p><p>Finally, I&#8217;ll look at the benchmark results, where Glimmer gets surprisingly close to Qwen3.6, and discuss why token efficiency may turn out to be an even more important metric for agentic workloads.</p><p><em>Note: I&#8217;m also preparing a full analysis detailing the model accuracy, token efficiency, memory consumption, and inference speed for the original and quantized versions. It should take about a week to gather enough data.</em></p>
      <p>
          <a href="/__u/kaitchup.substack.com/p/muse-glimmer-metas-30b-model-built">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Qwen3.8 Is Almost Here — and Agent Benchmarks Are More Fragile Than They Look]]></title><description><![CDATA[The Weekly Kaitchup #154]]></description><link>https://kaitchup.substack.com/p/qwen38-is-almost-here-and-agent-benchmarks</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-is-almost-here-and-agent-benchmarks</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 08 Aug 2026 00:27:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>We have <a href="https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B">a countdown for the Qwen3.8 releases</a>:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!2A10!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!2A10!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 424w, /__u/substackcdn.com/image/fetch/$s_!2A10!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 848w, /__u/substackcdn.com/image/fetch/$s_!2A10!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2A10!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!2A10!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png" width="1299" height="642" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:642,&quot;width&quot;:1299,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:328112,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/210098494?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!2A10!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 424w, /__u/substackcdn.com/image/fetch/$s_!2A10!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 848w, /__u/substackcdn.com/image/fetch/$s_!2A10!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2A10!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Next week is going to be a busy one.</p><p>Here&#8217;s my plan:</p><ul><li><p>Start my own evaluation of Qwen3.8 27B as soon as the model is released. This time, I&#8217;ll also cover agentic coding. I plan to publish the full analysis the following week.</p></li><li><p>Run and report on evaluations of the GGUF versions by the end of the week.</p></li><li><p>Publish a deeper analysis of the quantized models within 14 days of release.</p></li></ul><p>This will be similar to what I have done for Qwen3.6.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;f45f8a39-0604-449f-b6b2-b8a7d5525311&quot;,&quot;caption&quot;:&quot;In a previous article, I found Gemma 4 31B to be superior or comparable to Qwen3.5 27B in most areas, with similar or better accuracy and lower latency.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwen3.6 27B vs Qwen3.5 27B vs Gemma 4 31B: Accuracy, Latency, Memory, and Token Efficiency Tested&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-05T11:39:02.874Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!vAQL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2862fcd9-12bf-4ca0-ba6a-23d6808c8806_1210x783.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwen36-27b-vs-qwen35-27b-vs-gemma&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:195830510,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:24,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>I&#8217;ll also release my own quantized versions using my new quantization pipeline that brings together AutoRound (for high-quality quantization), LLM Compressor (for packing with a vLLM-friendly format), and a custom repacker (to correct bugs with some layers). The goal is to finally make mixed-precision models, with different layers quantized anywhere from 2-bit to 8-bit, work well with vLLM.</p><p>Conceptually, it brings some of the flexibility that GGUF provides for llama.cpp to the vLLM ecosystem, while targeting much higher throughput under heavy concurrency, for example, when running multiple sub-agents in parallel, by leveraging the <a href="https://github.com/inclusionAI/humming">Humming kernel</a>.</p><p>I&#8217;ve already published Qwen3.6 27B variants built with this pipeline. </p><ul><li><p><a href="https://huggingface.co/collections/kaitchup/qwen36-wna16">Qwen3.6 WNA16</a></p></li></ul><p>The 3.8 and 4.3 bit versions work well. I&#8217;m still evaluating the 3.5 and 3.0 bit versions. I&#8217;ll update the model cards once I&#8217;m sure they all work well. I&#8217;ll also add versions with a quantized language modeling head later this week. I thought about quantizing the token embeddings, but this isn't well supported by vLLM for the Qwen3.5 architecture (it doesn&#8217;t crash, but only generates &#8220;!!!!!&#8221;; that&#8217;s a bug, not a quantization quality issue). </p><p>Now I&#8217;m ready to do the same for Qwen3.8.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The Harness Can Matter Almost as Much as the Model</h2><p>A very interesting test from Composio showed how much the agent harness can change the results, even when the model stays the same.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/composio/status/2085330847951970801?s=20&quot;,&quot;full_text&quot;:&quot;We ran DeepSeek V4 Flash through 4 agent harnesses (Claude Code, Codex, OpenCode, Oh My Pi) on 30 agentic tasks.\n\nA different harness won on each metric: success rate, cost, and speed. &#129525;&#129525;&#129525; &quot;,&quot;username&quot;:&quot;composio&quot;,&quot;name&quot;:&quot;Composio&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/2036542629220089856/zc7Eix-q_normal.jpg&quot;,&quot;date&quot;:&quot;2026-08-06T11:43:12.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HPCSzoPXEAAf-bw.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/WFy6QNgv8r&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:81,&quot;retweet_count&quot;:51,&quot;like_count&quot;:740,&quot;impression_count&quot;:75062,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Yes, it sounds obvious when stated like this, but I increasingly come across model evaluations claiming superiority over other models while relying on numbers produced with suboptimal agent harnesses for the evaluated models, or with harnesses configured in completely different ways (more turns, higher timeouts, &#8230;).</p><p>At best, these comparisons are useless. More often, they are actively misleading.</p><p>Composio ran DeepSeek V4 Flash through four different harnesses, Claude Code, Codex, OpenCode, and Oh My Pi, across 30 multi-step agentic tasks. The tasks involved live tools such as Gmail, Google Sheets, GitHub, Slack, Notion, Calendar, Airtable, and PagerDuty. A task was considered successful only if it passed every predefined check.</p><p><strong>Observations</strong></p><ul><li><p>Oh My Pi achieved the highest success rate, completing 17 out of 30 tasks. Claude Code and Codex each completed 16, while OpenCode completed 14.</p></li><li><p>OpenCode had the lowest estimated cost per successful task at $0.073, followed by Codex at $0.081. Oh My Pi came in at $0.103, while Claude Code was the most expensive at $0.195.</p></li><li><p>Speed produced yet another ranking. Claude Code had the lowest median completion time at 122.7 seconds, with OpenCode close behind at 129.7 seconds. Codex took 245 seconds, while Oh My Pi was the slowest at 272.4 seconds.</p></li></ul><p>So, <strong>it seems</strong> we have different trade-offs:</p><ul><li><p>Claude Code was the fastest, but also the most expensive.</p></li><li><p>Oh My Pi completed the most tasks, but was the slowest.</p></li><li><p>OpenCode was the cheapest, but also had the lowest success rate.</p></li></ul><p>But we also have another angle not tackled here: <strong>inference-time sampling</strong>. Rerun the same experiments, and you may get a very different picture. I agree with Composio&#8217;s general conclusion, but we need many more runs to conclude which is the cheapest, fastest, and most accurate harness.</p><p>This is also why evaluating models for agentic coding is extremely difficult, and extremely expensive to do properly.</p><p>I&#8217;m currently preparing an article focused on Laguna S2.1 and how I benchmarked it for agentic coding. I ran into far more issues than I would have liked to admit, but the process made one thing very clear to me: a large number of the agentic coding benchmark tables I see online should be treated with extreme caution.</p><p><strong>Without controlling for the harness, its configuration, tools, prompts, retry behavior, and execution environment, comparing model scores can quickly become meaningless.</strong></p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisionning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/kaitchup.substack.com/subscribe"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p>]]></content:encoded></item><item><title><![CDATA[ThinkingCap-Qwen3.6-27B Review: 2x Fewer Tokens, Same Accuracy?]]></title><description><![CDATA[A faster, more stable Qwen3.6 for local AI inference]]></description><link>https://kaitchup.substack.com/p/thinkingcap-qwen36-27b-review-2x</link><guid isPermaLink="false">https://kaitchup.substack.com/p/thinkingcap-qwen36-27b-review-2x</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Wed, 05 Aug 2026 01:11:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!MbCr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa76d07ac-f9de-490f-8a81-076f8b0bd919_912x522.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Qwen3.6 27B is one of the strongest models for local AI, but its extremely long reasoning traces can make inference painfully slow.</p><p>The community has tried to address this issue, but with limited success. One approach is to introduce a &#8220;thinking cap&#8221; that stops the reasoning process once the model reaches a predefined token budget. Another is to fine-tune the model on shorter, more efficient reasoning traces.</p><p>I examined both approaches in a previous article about Qwen3.5 and Qwen3.6. These approaches either performed poorly or required substantially more expensive fine-tuning.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;dd9a98dc-45de-40b9-b924-9befbadabef1&quot;,&quot;caption&quot;:&quot;LLMs now rely heavily on reasoning traces to improve accuracy. A reasoning trace is the intermediate text generated before the final answer, often delimited by tags such as <think>...</think>.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Reasoning Budgets vs. Structured CoT: Controlling LLM Thinking Tokens&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-25T16:10:24.390Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!5HDY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4419269c-1249-4fc1-8766-2f460c5bccb8_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/reasoning-budgets-vs-structured-cot&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:195230612,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:10,&quot;comment_count&quot;:6,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;f1e4fd13-172d-4608-b39f-d84286335242&quot;,&quot;caption&quot;:&quot;As we saw in previous articles, Qwen3.6 are very good LLMs for local AI.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwopus and REAP: Custom Qwen3.6 Models for Local Reasoning&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-06-17T20:22:48.702Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!DVuQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76668f-e137-416c-9826-d6d134dd6a60_1210x753.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwopus-and-reap-custom-qwen36-models&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:201494160,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>Hopefully, Qwen3.8 27B, which is expected to be released this week, will be a more efficient reasoner.</p><p>In the meantime, BottleCap AI has released a promising custom alternative that significantly reduces the reasoning cost of Qwen3.6: <strong><a href="https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B">ThinkingCap-Qwen3.6-27B</a></strong>.</p><p>ThinkingCap-Qwen3.6-27B is a post-trained version of Qwen3.6-27B. Its main improvement is behavioral: it has been trained to produce <strong>shorter reasoning traces</strong> before delivering an answer.</p><p>The model also tends to generate <strong>shorter final responses</strong>. According to BottleCap AI, this behavior emerged during training alongside the reduction in reasoning length.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In this article, I evaluate the model&#8217;s accuracy and token efficiency across several tasks. I compare it with the original Qwen3.6 and another model known for its efficient reasoning, Gemma 4 31B IT.</p><p>All three models are evaluated using exactly the same evaluation settings, with their recommended hyperparameters, making the results directly comparable. Every number reported here comes from my own testing.</p><p>I also examine the model&#8217;s reasoning stability. In particular, is ThinkingCap-Qwen3.6-27B less prone to excessively long or seemingly endless reasoning than the original Qwen3.6?</p>
      <p>
          <a href="/__u/kaitchup.substack.com/p/thinkingcap-qwen36-27b-review-2x">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[DeepSeek-V4-Flash-0731 and Inkling Small: Smaller, but Better?]]></title><description><![CDATA[The Weekly Kaitchup #153]]></description><link>https://kaitchup.substack.com/p/deepseek-v4-flash-0731-and-inkling</link><guid isPermaLink="false">https://kaitchup.substack.com/p/deepseek-v4-flash-0731-and-inkling</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Fri, 31 Jul 2026 19:33:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of <em>The Weekly Kaitchup</em>:</p><ul><li><p>DeepSeek V4 Flash: Now Better than DeepSeek V4 Pro (Preview)</p></li><li><p>Inkling-Small: Better than Inkling (?)</p></li><li><p>Escha-W2: Qwen3.6 35B Compressed to 12.3 GB and Faster</p></li></ul><div><hr></div><p>DeepSeek released an update for its V4 Flash model. As usual, DeepSeek released the weights for this update immediately, rather than imposing the waiting period we have seen from companies such as MiniMax and Qwen.</p><ul><li><p><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731">deepseek-ai/DeepSeek-V4-Flash-0731</a></p></li></ul><p>It now outperforms the much larger V4 Pro (Preview), with particularly strong gains in agentic coding:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!eyxR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!eyxR!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 424w, /__u/substackcdn.com/image/fetch/$s_!eyxR!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 848w, /__u/substackcdn.com/image/fetch/$s_!eyxR!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 1272w, /__u/substackcdn.com/image/fetch/$s_!eyxR!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!eyxR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png" width="561" height="383.1219512195122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1066,&quot;resizeWidth&quot;:561,&quot;bytes&quot;:82164,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/209166534?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!eyxR!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 424w, /__u/substackcdn.com/image/fetch/$s_!eyxR!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 848w, /__u/substackcdn.com/image/fetch/$s_!eyxR!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 1272w, /__u/substackcdn.com/image/fetch/$s_!eyxR!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>They did not publish results for other task categories. I checked <a href="https://artificialanalysis.ai/models?models=mimo-v2-5-pro%2Cgemini-3-5-flash-lite%2Cinkling%2Cclaude-sonnet-5%2Cminimax-m3%2Ccommand-a-plus%2Cgpt-5-6-luna%2Cnvidia-nemotron-3-ultra-550b-a55b%2Cqwen3-7-max%2Cgemini-3-6-flash%2Cgrok-4-5%2Cclaude-4-5-haiku-reasoning%2Cclaude-opus-5%2Cgpt-5-6-terra%2Cdeepseek-v4-pro%2Cgemma-4-31b%2Cclaude-fable-5%2Cmuse-spark-1-1%2Cgpt-5-6-sol%2Cmistral-medium-3-5%2Cgpt-5-5-pro%2Cgpt-oss-120b%2Cglm-5-2%2Ckimi-k3%2Cdeepseek-v4-flash%2Cdeepseek-v4-flash-0420">Artificial Analysis&#8217; additional results</a> and found that it also improved on other benchmarks, including GPQA Diamond, which is a very different task from agentic coding. In fact, none of the reported benchmarks appear to show a regression. This is not merely a more specialized model, it is a substantial improvement over the preview and a very promising sign for the next V4 Pro update.</p><p>On a related note, Qwen also released Qwen3.7 Flash this week. It is currently available only through APIs and is very inexpensive. The community has speculated that it may be a Qwen3.7 35B-A3B model and that Qwen will eventually release the weights. I am not certain this is correct, but it would make sense.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Inkling-Small: Smaller, but Better on Most Benchmarks</h2><p>Despite the name, <a href="https://huggingface.co/thinkingmachines/Inkling-Small">Inkling-Small</a> is not a lightweight model in the usual sense. It is still an open-weight, decoder-only multimodal model that accepts:</p><ul><li><p>Text</p></li><li><p>Images</p></li><li><p>Audio</p></li></ul><p>It follows the same general architecture as the larger Inkling. Both models use a sparse mixture-of-experts design, routing each token through six of 256 experts, along with two shared experts. Both also support context windows of up to one million tokens.</p><p>So Inkling-Small is mostly the same design, scaled down.</p><h3>How much smaller is it?</h3><p>The larger Inkling has:</p><ul><li><p>66 transformer layers</p></li><li><p>975 billion total parameters</p></li><li><p>About 41 billion active parameters per token</p></li></ul><p>Inkling-Small reduces that to:</p><ul><li><p>42 transformer layers</p></li><li><p>276 billion total parameters</p></li><li><p>About 12 billion active parameters per token</p></li></ul><p>That makes Inkling-Small a little over one-quarter the size of the larger model by total parameter count. Its active parameter count is reduced by a similar amount.</p><h3>The NVFP4 version is much easier to fit</h3><p>Inkling-Small requires around <strong>600 GB of aggregate VRAM in BF16</strong>.</p><p>The <a href="https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4">NVFP4 checkpoint</a> brings that down to roughly <strong>180 GB</strong>.</p><p>For comparison, the larger Inkling&#8217;s NVFP4 checkpoint is around <strong>592&#8211;600 GB</strong>. In other words, the quantized version of the large model takes roughly as much memory as Inkling-Small does in BF16.</p><p>The word &#8220;small&#8221; is doing some work here, but 180 GB is at least in the range of a compact multi-GPU server rather than a large GPU cluster.</p><h3>The smaller model is weirdly better at a lot of things</h3><p>The surprising part is that Inkling-Small beats the larger Inkling across most published head-to-head benchmarks.</p><p>A few examples:</p><ul><li><p><strong>SWE-bench Verified:</strong> 80.2% vs. 77.6%</p></li><li><p><strong>Humanity&#8217;s Last Exam:</strong> 31.6% vs. 29.7%</p></li><li><p><strong>IFBench:</strong> 82.2% vs. 79.8%</p></li></ul><p>The smaller model appears particularly strong in coding, tool use, instruction following, and several reasoning evaluations.</p><p>Thinking Machines attributes the improvement to a revised data mix, an updated training recipe, distillation from the larger Inkling model, and additional reinforcement learning focused on agentic coding. </p><p>The larger model still performs better in some areas, especially factual recall, broad knowledge coverage, and most audio evaluations.</p><p>Still, it is an unusual result: the smaller model uses far fewer parameters, needs much less memory, and yet comes out ahead on a majority of the reported benchmarks.</p><p>Another unexpected result is that it significantly outperforms Qwen3.5 397B despite being much smaller, especially on benchmarks where Qwen3.5 has always been exceptionally good, like GPQA Diamond.</p><p>Is it benchmaxxed?? I&#8217;m waiting for the community feedback!</p><div><hr></div><h2>Escha-W2: Qwen3.6 35B Compressed to 12.3 GB</h2><p>If you primarily use GGUF models locally, a 12.3 GB version of Qwen3.6 35B may not sound especially impressive. After all, that is roughly the size of a heavily quantized GGUF model, like a Q2_K_XL, and we already know those can perform well, as shown by the results I published a few months ago.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!6pUp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!6pUp!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 424w, /__u/substackcdn.com/image/fetch/$s_!6pUp!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 848w, /__u/substackcdn.com/image/fetch/$s_!6pUp!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6pUp!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!6pUp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png" width="1456" height="730" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b161b14e-e18d-425e-8437-318efff937b5_1680x842.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:730,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!6pUp!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 424w, /__u/substackcdn.com/image/fetch/$s_!6pUp!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 848w, /__u/substackcdn.com/image/fetch/$s_!6pUp!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6pUp!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The key difference with Escha-W2 is that it uses (presumably) a more advanced quantization method and comes with an optimized inference runtime. As a result, it should run faster than the typical GGUF model, while already supporting SGLang for significantly better throughput under concurrent workloads. The developers have also said that vLLM support is coming later.</p><h3>How They Did It?</h3><p>We don&#8217;t really know yet.</p><p>The original model is a mixture-of-experts model with:</p><ul><li><p>35 billion total parameters</p></li><li><p>Around 3 billion active parameters per token</p></li><li><p>256 experts</p></li></ul><p>Escha Labs keeps that underlying architecture but compresses most of the expert weights to roughly two bits. More precisely, it uses a mixture of two- and three-bit quantization for the expert projections, while keeping the dense layers in INT8.</p><h3>The whole model is 12.3 GB</h3><p>The resulting checkpoint takes up <strong>12.3 GB on disk</strong>.</p><p>It can run on:</p><ul><li><p>A single 24 GB GPU under the recommended configuration</p></li><li><p>A 16 GB GPU with reduced context length or fewer concurrent requests</p></li></ul><p>On an RTX 4090, Escha reports around <strong>225 tokens per second</strong> for a single stream. The same setup reaches about 1,321 tokens per second when serving 32 requests at once.</p><h3>Two bits, but roughly FP8-level benchmark scores</h3><p>The more interesting part is how little the compression appears to affect most of the reported evaluations.</p><p>Across six benchmark categories, Escha-W2 averages <strong>100.2% of the FP8 model&#8217;s score</strong>. That does not mean the quantized model is genuinely better, the small gains are mostly within normal evaluation variance, but it does suggest that the overall quality loss is limited.</p><p>A few examples:</p><ul><li><p><strong>MMLU-Pro:</strong> 80.9 versus 82.3 for FP8</p></li><li><p><strong><s>MATH-500:</s></strong><s> 93.8 versus 91.2</s> (ignore this one; this is too old and too easy for Qwen3.6)</p></li><li><p><strong>GPQA-Diamond:</strong> 77.8 versus 74.7</p></li><li><p><strong>BFCL tool use:</strong> 88.9 versus 88.2</p></li><li><p><strong>RULER long-context retrieval:</strong> 89.9 versus 89.4</p></li></ul><h3>Coding is the main place where it loses ground</h3><p>The clearest regression is on longer coding tasks, where quantization always does more damage.</p><p>On LiveCodeBench v6, <strong>Escha-W2 scores 62.6, compared with 67.0 for the FP8 baseline</strong>. <em>Note: I don&#8217;t know which framework they used to get a 67.0, but these scores seem very low. It should be around 85.0. This suggests that they ran it with a limited context length, like 32K max tokens. Not great. This means that the accuracy gap could be greater at longer context length, as quantization tends to be worse as sequence length increases.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!oa-Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!oa-Z!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 424w, /__u/substackcdn.com/image/fetch/$s_!oa-Z!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 848w, /__u/substackcdn.com/image/fetch/$s_!oa-Z!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oa-Z!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!oa-Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png" width="776" height="525" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:525,&quot;width&quot;:776,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:46780,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/209166534?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!oa-Z!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 424w, /__u/substackcdn.com/image/fetch/$s_!oa-Z!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 848w, /__u/substackcdn.com/image/fetch/$s_!oa-Z!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oa-Z!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>It is not a normal drop-in quantization</h3><p>There is one practical catch: Escha-W2 needs a custom runtime.</p><p>Its weights use a custom packed format, so the model does not simply load into standard inference software as a regular GPTQ, AWQ, or GGUF checkpoint. Escha currently provides two options:</p><ul><li><p>An SGLang-based runtime for concurrency, tool calling and structured output</p></li><li><p>A standalone ZML runtime aimed at single-user inference</p></li></ul><p>The problem with specialized runtimes is that it&#8217;s very hard work to maintain them. </p><p>The checkpoint is also text-only, even though the original Qwen architecture includes vision components.</p><p>When they release the vLLM runtime, I&#8217;ll double-check the accuracy on longer context and compare it with quantized versions of similar size.</p><p></p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisionning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/kaitchup.substack.com/subscribe"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p>]]></content:encoded></item><item><title><![CDATA[Bonsai 27B Review: Can a 3.9 GB 1-Bit Model Match Qwen3.6 27B?]]></title><description><![CDATA[An in-depth look at Bonsai 27B&#8217;s accuracy, token efficiency, reasoning stability, and production trade-offs.]]></description><link>https://kaitchup.substack.com/p/bonsai-27b-review-can-a-39-gb-1-bit</link><guid isPermaLink="false">https://kaitchup.substack.com/p/bonsai-27b-review-can-a-39-gb-1-bit</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Tue, 28 Jul 2026 16:06:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dlgR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!dlgR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!dlgR!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 424w, /__u/substackcdn.com/image/fetch/$s_!dlgR!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 848w, /__u/substackcdn.com/image/fetch/$s_!dlgR!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dlgR!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!dlgR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png" width="607" height="607" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:607,&quot;bytes&quot;:1283747,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/207459899?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!dlgR!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 424w, /__u/substackcdn.com/image/fetch/$s_!dlgR!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 848w, /__u/substackcdn.com/image/fetch/$s_!dlgR!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dlgR!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Prism ML has released impressive ternary and 1-bit variants of Qwen3.6 27B. The 1-bit version is only 3.9 GB, making it dramatically smaller than the original, which is roughly 55 GB. </p><ul><li><p><a href="https://huggingface.co/collections/prism-ml/bonsai-27b">Bonsai 27B Models</a> (Hugging Face)</p></li></ul><p>These Bonsai 27B models behave quite differently from the original Qwen3.6 model, particularly in terms of accuracy and token efficiency.</p><p><em>Can you get Qwen3.6 27B-level accuracy from a 3.9 GB model? </em></p><p>With reasoning enabled, Bonsai 27B can nearly match the accuracy of Qwen3.6 27B running without reasoning. Moreover, that dramatic reduction in memory comes at a substantial computational cost: Bonsai needs to generate far more tokens to achieve the same result, up to 14 times more on some coding tasks.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>When reasoning is enabled for both models, Qwen3.6 27B remains clearly ahead. It uses fewer tokens while achieving significantly higher accuracy. Still, what Prism ML has achieved is remarkable.</p><p>In this article, I take a deep dive into Bonsai 27B&#8217;s recipe, benchmark performance (my own numbers), and token efficiency. We will examine where the model struggles most, why its raw accuracy is lower, and why the results are nevertheless highly promising.</p>
      <p>
          <a href="/__u/kaitchup.substack.com/p/bonsai-27b-review-can-a-39-gb-1-bit">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Agentic AI at Two Different Scales: Nanbeige4.2-3B and Laguna S2.1 ]]></title><description><![CDATA[The Weekly Kaitchup #152]]></description><link>https://kaitchup.substack.com/p/agentic-ai-at-two-different-scales</link><guid isPermaLink="false">https://kaitchup.substack.com/p/agentic-ai-at-two-different-scales</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 25 Jul 2026 04:12:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of <em>The Weekly Kaitchup</em>, I&#8217;m looking at two of the week&#8217;s most interesting releases for agentic workloads: Nanbeige4.2-3B and Laguna S 2.1.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Agentic workloads include multi-step reasoning, tool use, interaction with external environments, and tasks that may require many actions before completion.</p><ul><li><p><a href="https://huggingface.co/Nanbeige/Nanbeige4.2-3B">Nanbeige/Nanbeige4.2-3B</a></p></li><li><p><a href="https://huggingface.co/poolside/Laguna-S-2.1">poolside/Laguna-S-2.1</a></p></li></ul><p>Despite this shared focus, they operate at very different scales.</p><p>Nanbeige4.2-3B is a compact dense model with approximately 4 billion total parameters and 3 billion non-embedding parameters. It is intended to make capable agentic behavior practical on consumer and workstation hardware.</p><p>Laguna S 2.1 is a 118-billion-parameter Mixture-of-Experts (MoE) model. It activates approximately 8 billion parameters for each token.</p><p>The two models represent different approaches to agentic AI. Nanbeige uses repeated computation over a relatively small set of weights, while Laguna uses sparse access to a much larger pool of learned parameters.</p><h2>Nanbeige4.2-3B</h2><p>While Nanbeige4.1 used a basic Llama architecture, Nanbeige4.2-3B uses a Looped Transformer.</p><p>Instead of containing a large number of unique transformer layers, the model has 22 physical decoder layers that are <strong>executed twice</strong>. The same weights are reused during the second pass, giving the model the computational depth of approximately 44 layer executions without storing 44 independent layers.</p><p>This design reduces weight memory, but it does not reduce inference computation to that of a normal 22-layer model. Each token must still pass through the transformer stack twice.</p><p>The model has 48 attention heads with a dimension of 128, and eight key/value heads. This is a lot. For instance, Qwen3.6 35B A3B has only two KV heads. </p><p>Looping introduces a further memory consideration. Unless the implementation explicitly shares KV-cache state between the two passes, each loop may require separate key/value entries. Nanbeige&#8217;s default configuration does not appear to enable loop-level KV sharing.</p><p>Moreover, I think models using the Looped Transformer should be better named. When you use a 3B model, you don&#8217;t expect it to be nearly as slow as a 6B model. small MoE models, like Qwen3.6 and Gemma 4, show the number of active parameters in their name, like 26B-A4B. For Looped Transformers, we should adopt a similar convention, like 3B-2P (for &#8220;2 passes&#8221;),  for instance.</p><h4>Nanbeige4.2-3B&#8217;s memory consumption</h4><p>A 16GB GPU should be able to run the model, without quantization, at moderate context lengths with careful settings. A 24GB GPU provides more room for longer prompts, greater concurrency, and runtime overhead.</p><p>Based on the model&#8217;s 22 physical layers, two loop passes, eight key/value heads, 128-dimensional heads, and BF16 KV values, a conservative estimate places KV-cache consumption at approximately 176 KiB per cached token for each active sequence.</p><p><em>I show in the following article how to estimate the KV cache memory consumption:</em></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;6e951bec-dcc3-4e7e-a5ff-d1b040652017&quot;,&quot;caption&quot;:&quot;Inference efficiency has become one of the main ways LLMs distinguish themselves. Two models can look similar on paper: roughly the same parameter count, similar context length, both marketed as fast, yet behave very differently once you actually try to serve them at scale. In this article, we focus on four recent open models built for efficient deployment using different architectures:&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The KV-Cache of Small MoEs: Qwen3, Qwen3.5/3.6, GLM 4.7 Flash, and Nemotron 3 Nano Compared&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-03-18T20:37:03.129Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!BR8P!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00c04d8e-bfbd-4069-b1a3-2be9d9e2e8ea_1751x790.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/the-kv-cache-of-small-moes-qwen3&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:191136984,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:27,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>At 16K tokens, this would require roughly 2.75 GiB of KV-cache memory. Using the complete 256K context could require roughly 44 GiB of KV-cache memory for a single active sequence.</p><p>This is very large. Almost twice as much as what Qwen3.6 27B would consume for the same context length.</p><h3>Using Nanbeige4.2-3B with vLLM</h3><p>At release, Nanbeige provided a dedicated vLLM branch for the model:</p><pre><code><code>git clone -b nanbeige42 https://github.com/Nanbeige/vllm.git
cd vllm
pip install -e .</code></code></pre><p>The model can then be served through vLLM&#8217;s OpenAI-compatible API:</p><pre><code><code>vllm serve Nanbeige/Nanbeige4.2-3B \
  --host 0.0.0.0 \
  --port 8000 \
  --enable-auto-tool-choice \
  --tool-call-parser nanbeige \
  --reasoning-parser nanbeige

 #use --host 127.0.0.1 if you are running it locally</code></code></pre><p>The reasoning and tool-call parsers are important for agent frameworks. They allow vLLM to separate reasoning content and structured tool calls from ordinary assistant text.</p><p>Nanbeige&#8217;s chat template exposes two controls.</p><p>The <code>enable_thinking</code> option determines whether the current response contains an explicit reasoning phase. The <code>preserve_thinking</code> option determines whether reasoning content from previous assistant turns remains in the conversation history.</p><p>For normal chat and question answering, preserved reasoning can usually be disabled. For multi-turn tool use, office workflows, and coding agents, the developers recommend retaining earlier reasoning content but I don&#8217;t understand how this can work well. Reasoning traces can be very long, like 50K+ tokens. Keeping them in the context, even if it&#8217;s just one, means that the model may reach its max context length at next turn.</p><h3>Performance</h3><p>According to the published evaluation, Nanbeige exceeded Qwen3.5-9B on most of the included agent, coding, and reasoning tasks. It also surpassed Gemma 4 12B on benchmarks including GDPval, SWE-Bench Verified, SWE-Bench Pro, and Terminal-Bench 2.0.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!z6QI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!z6QI!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 424w, /__u/substackcdn.com/image/fetch/$s_!z6QI!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 848w, /__u/substackcdn.com/image/fetch/$s_!z6QI!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 1272w, /__u/substackcdn.com/image/fetch/$s_!z6QI!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!z6QI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png" width="1456" height="993" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:993,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!z6QI!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 424w, /__u/substackcdn.com/image/fetch/$s_!z6QI!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 848w, /__u/substackcdn.com/image/fetch/$s_!z6QI!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 1272w, /__u/substackcdn.com/image/fetch/$s_!z6QI!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Agent results are influenced not only by the base model but also by the allowed reasoning budget, tool definitions, prompting strategy, conversation-state handling, retry policy, and maximum number of actions.</p><p><em>Reference: <a href="https://huggingface.co/Nanbeige/Nanbeige4.2-3B/blob/main/Nanbeige42_report.pdf">Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model</a></em></p><h2>Laguna S 2.1</h2><p>Laguna S 2.1 is a much larger coding-focused MoE model.</p><p>It contains 118 billion total parameters but activates approximately 8 billion parameters for each token. A routing mechanism selects a subset of experts during inference, allowing the model to access a large learned parameter pool without executing every parameter for every token.</p><p>Laguna contains 48 transformer layers. Twelve use global attention, while the remaining 36 use sliding-window attention with a window of 512 tokens. The global and local layers are interleaved in approximately a one-to-three ratio.</p><p>The global layers allow information to move across the complete input context. The sliding-window layers restrict attention to nearby tokens, reducing the cost of long-context inference and limiting KV-cache growth.</p><p>The model contains 256 routed experts and one shared expert. For each token, the router selects the top ten routed experts in addition to the shared computation.</p><p>Laguna also uses per-head softplus output gating.</p><blockquote><p><strong>Per-head softplus output gating</strong> gives each attention head its own learned volume control. For every token, the model can reduce, preserve, or amplify the contribution of individual heads before combining their outputs. The softplus function keeps these gates positive while allowing values above one, so useful attention heads can be strengthened rather than merely switched on or off.</p></blockquote><p>It uses interleaved reasoning. The model can reason, issue a tool call, receive the tool result, and resume its reasoning before selecting another action.</p><p>Poolside also provides a <a href="https://huggingface.co/poolside/Laguna-S-2.1-DFlash">DFlash draft model</a> for speculative decoding. These models generate candidate tokens that the main model can verify in parallel, potentially increasing output speed.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;d6c5a31a-2b8c-4985-a343-58e4d7652974&quot;,&quot;caption&quot;:&quot;Speculative decoding is becoming a popular way to accelerate LLM inference, with approaches such as MTP and DFlash. The idea is simple: a smaller or specialized draft model proposes future tokens, and the full target model verifies them. Matching tokens are accepted and when a mismatch occurs, only the valid prefix is kept and decoding falls back to the target model.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Train and Run DFlash Speculative Decoding&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-18T19:43:12.333Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sbvO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b274a34-f749-4a9f-b68b-0e0f501a9016_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/train-and-run-dflash-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196847181,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h3>Using Laguna S 2.1 with vLLM</h3><p>Laguna support is documented for vLLM 0.25.0 or newer:</p><pre><code><code>uv pip install -U "vllm&gt;=0.25.0"</code></code></pre><p>The model can run on a single B300 GPU:</p><pre><code><code>vllm serve poolside/Laguna-S-2.1 \
  --enable-auto-tool-choice \
  --tool-call-parser poolside_v1 \
  --reasoning-parser poolside_v1 \
  --default-chat-template-kwargs '{"enable_thinking": true}'</code></code></pre><p>The <a href="https://huggingface.co/poolside/Laguna-S-2.1-NVFP4">NVFP4</a> and <a href="https://huggingface.co/poolside/Laguna-S-2.1-INT4">INT4</a> versions consume fewer than 80 GB but you would need a 96 GB GPU, like an RTX Pro 6000, to exploit the context length of the model.</p><pre><code><code>vllm serve poolside/Laguna-S-2.1-INT4 \
  --enable-auto-tool-choice \
  --tool-call-parser poolside_v1 \
  --reasoning-parser poolside_v1 \
  --default-chat-template-kwargs '{"enable_thinking": true}'</code></code></pre><p>For coding-agent workloads, Poolside also recommends retaining prior <code>reasoning_content</code> in the conversation history.</p><p>As for the memory consumption of the KV cache, 256K tokens should consume around 24 GB.</p><h3>Target tasks</h3><p>Laguna S 2.1 is more specialized than Nanbeige.</p><p>Its primary target is long-horizon software engineering. This includes repository-level bug fixing, terminal interaction, shell-based workflows, multilingual code maintenance, codebase exploration, and answering questions that require understanding large software repositories.</p><h3>Performance</h3><p>Laguna&#8217;s sparse architecture and coding-focused training are particularly effective for terminal interaction, repository-level reasoning, and multilingual software engineering.</p><p>It looks like a very good model but so far, I have found community feedback to be mixed.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8N_3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8N_3!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 424w, /__u/substackcdn.com/image/fetch/$s_!8N_3!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 848w, /__u/substackcdn.com/image/fetch/$s_!8N_3!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 1272w, /__u/substackcdn.com/image/fetch/$s_!8N_3!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8N_3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg" width="1456" height="2295" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2295,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;benchmarks&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="benchmarks" title="benchmarks" srcset="/__u/substackcdn.com/image/fetch/$s_!8N_3!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 424w, /__u/substackcdn.com/image/fetch/$s_!8N_3!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 848w, /__u/substackcdn.com/image/fetch/$s_!8N_3!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 1272w, /__u/substackcdn.com/image/fetch/$s_!8N_3!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>reference: <a href="https://poolside.ai/blog/introducing-laguna-s-2-1">Introducing Laguna S 2.1</a></em></p><p>I&#8217;ll spend some time with Nanbeige and Laguna. If everything goes well, I&#8217;ll probably write an article using them for agentic coding.</p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisionning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/kaitchup.substack.com/subscribe"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p>]]></content:encoded></item><item><title><![CDATA[Qwen3.8: What Hardware Will You Need to Run Alibaba’s 2.4T Model?]]></title><description><![CDATA[Estimating the memory, storage, and GPU requirements for BF16, NVFP4, Q4, and TQ1 versions of Qwen3.8.]]></description><link>https://kaitchup.substack.com/p/qwen38-what-hardware-will-you-need</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-what-hardware-will-you-need</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Wed, 22 Jul 2026 16:29:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8fje!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8fje!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8fje!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 424w, /__u/substackcdn.com/image/fetch/$s_!8fje!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 848w, /__u/substackcdn.com/image/fetch/$s_!8fje!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8fje!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8fje!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png" width="507" height="253.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:507,&quot;bytes&quot;:804225,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/207990470?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8fje!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 424w, /__u/substackcdn.com/image/fetch/$s_!8fje!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 848w, /__u/substackcdn.com/image/fetch/$s_!8fje!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8fje!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Alibaba&#8217;s Qwen3.8 is a 2.4-trillion-parameter frontier model. The company has indicated that it plans to release the model&#8217;s weights, potentially making Qwen3.8 one of the largest openly downloadable AI models ever produced.</p><p>For comparison, Alibaba&#8217;s largest open-weight model to date, <a href="https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct">Qwen3-Coder-480B-A35B-Instruct</a>, is roughly five times smaller. Qwen3.8 would still have about 400 billion fewer parameters than Kimi K3.</p><p>With both K3 and Qwen3.8 performing nearly at the frontier, <em>was closing the accuracy gap with OpenAI and Anthropic largely a matter of scaling from hundreds of billions of parameters to more than two trillion?</em></p><p>An open-weight release for Qwen3.8 is excellent news. It could give researchers, companies, and the broader AI community unprecedented access to a model at this scale.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>But what could most people realistically do with a 2.4-trillion-parameter model?</p><p>Very little locally. Even with aggressive quantization, compression, or distributed inference, a model of this size would be far too large to run on a typical personal computer. Its release could still be highly significant for research institutions, infrastructure providers, and well-resourced community projects. For casual users, however, direct local deployment would remain largely impractical.</p><p>Let&#8217;s check, with a few assumptions, what you would need to run NVFP4, Q4, and Q1 GGUF versions.</p><h2>Qwen3.8: A Sparse MoE?</h2><p><em>As I write this, Alibaba has not disclosed when it plans to release the model&#8217;s weights. That could change quickly. They may even be released by the time this article is published.</em></p><p>Alibaba has not yet disclosed how many parameters are activated for each token, how many experts the model contains, or whether every layer uses mixture-of-experts routing.</p><p>Nevertheless, a dense 2.4-trillion-parameter transformer would be extraordinarily expensive to serve. It is thus reasonable to assume that Qwen3.8 uses a highly sparse mixture-of-experts, or MoE, architecture.</p><p>Until Alibaba publishes the architecture, Moonshot AI&#8217;s Kimi K2 provides a useful proxy for estimating how a trillion-parameter MoE model might distribute its weights.</p><p>Kimi K2 contains:</p><ul><li><p><strong>Total parameters:</strong> 1.04 trillion</p></li><li><p><strong>Activated parameters:</strong> 32.6 billion</p></li><li><p><strong>Routed experts:</strong> 384</p></li><li><p><strong>Experts selected per token:</strong> 8</p></li><li><p><strong>Routing sparsity:</strong> 8/384, or 1/48</p></li><li><p><strong>Shared experts:</strong> 1</p></li></ul><p><em>Note: These figures come from the <a href="https://arxiv.org/abs/2507.20534">Kimi K2 technical report</a>. </em></p><p><strong>98.9% of Kimi K2&#8217;s parameters are located inside routed experts</strong>, while only about 1.1% are always-active or otherwise non-routed weights. The trend is the same for all the MoE with 500B+ parameters: Nearly 99% of the parameters are in the routed experts.</p><h3>Applying the same ratio to Qwen3.8</h3><p>In other words, Qwen3.8 could store 2.4 trillion parameters while using only a small fraction of them for each token. If its routing sparsity resembled Kimi K2&#8217;s, the computation required per token might be comparable to that of a roughly 75-billion-parameter dense model, before accounting for routing and distributed-communication overhead.</p><p>BF16 stores each parameter using 16 bits, or two bytes.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;2.4\\text{T parameters}\\times2\\text{ bytes}\n=4.8\\text{ TB}&quot;,&quot;id&quot;:&quot;NRJXFDHYWW&quot;}" data-component-name="LatexBlockToDOM"></div><p>This does not include the KV cache, inference-engine workspace, CUDA graphs, temporary activations, multimodal components, tokenizer files or checkpoint metadata.</p><p>A practical deployment would consequently need more than 4.8 TB of combined accelerator memory. I won&#8217;t speculate on the KV cache size. There are too many variables that influence it: number of linear layers, number of KV heads, etc.</p><h2>Experts in NVFP4, remainder in BF16</h2><p>I hope Alibaba releases an official NVFP4 version. Producing a high-quality NVFP4 quantization of a model this large would be prohibitively expensive for most community projects. NVIDIA has also been relatively slow to publish its own NVFP4 conversions of open-weight models; for example, <a href="https://huggingface.co/nvidia/Qwen3.6-27B-NVFP4">its version of Qwen3.6 appeared only recently</a>.</p><p>NVFP4 stores each parameter as a 4-bit value, with a single FP8 scaling factor shared across all 16-value blocks. This results in an effective storage cost of approximately 4.5 bits per parameter, plus a negligible FP32 scaling value for each tensor.</p><p>Fortunately, routed experts tend to be relatively robust to quantization. In MoE models, the expert weights can often be stored at lower precision while the comparatively small set of shared, attention, embedding, and other always-active parameters remains at higher precision. This mixed-precision approach can substantially reduce storage requirements while limiting the effect on model quality.</p><p>So, with 99% of parameters quantized to NVFP4:</p><ul><li><p>2.374 trillion routed-expert parameters at 4.5 bits each</p></li><li><p>25.8 billion remaining parameters at 16 bits each</p></li></ul><p>The expert weights consume:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;2.374\\text{T}\\times\\frac{4.5}{8}\n\\approx1.335\\text{ TB}&quot;,&quot;id&quot;:&quot;JUTJBGJUHL&quot;}" data-component-name="LatexBlockToDOM"></div><p>The BF16 remainder consumes:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;25.8\\text{B}\\times2\n\\approx51.5\\text{ GB}&quot;,&quot;id&quot;:&quot;QAIKYJOHOD&quot;}" data-component-name="LatexBlockToDOM"></div><p>Total:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;1.335\\text{ TB}+0.0515\\text{ TB}\n\\approx1.387\\text{ TB}&quot;,&quot;id&quot;:&quot;ZKFYSZIKTS&quot;}" data-component-name="LatexBlockToDOM"></div><p>That&#8217;s an average of <strong>4.62 bits per parameter.</strong></p><p>Compared with a 4.8 TB BF16 checkpoint, that saves approximately:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;4.8-1.387=3.413\\text{ TB}&quot;,&quot;id&quot;:&quot;SMXBFEKRUQ&quot;}" data-component-name="LatexBlockToDOM"></div><p>or <strong>71.1% of the original weight memory</strong>.</p><h2>What about a typical Q4 GGUF?</h2><p>A common GGUF format such as Q4_K_M is approximately 4.8 bits per weight in representative models, while Q4_K_S is closer to 4.6 bits per weight. Actual ratios vary with architecture and which tensors remain at higher precision.</p><p>For a 2.4-trillion-parameter checkpoint:</p><ul><li><p><strong>Q4_K_S-like:</strong> 4.58 bits; approximately <strong>1.374 TB</strong></p></li><li><p><strong>Q4_K_M-like:</strong> 4.84 bits; approximately <strong>1.452 TB</strong></p></li><li><p><strong>Higher-overhead Q4:</strong> 5.0 bits; approximately <strong>1.50 TB</strong></p></li></ul><p>With runtime allocations and a useful KV cache, a deployment should budget at least 1.6&#8211;1.8 TB of usable memory, and more for long contexts or concurrent users.</p><h2>What would TQ1 weigh?</h2><p>I&#8217;m mentioning TQ1 since Unsloth has previously released very good TQ1 versions of Qwen3.5. We have no guarantee they can do/will do the same for Qwen3.8.</p><p>TQ1_0 uses a compact ternary representation at approximately 1.69 bits per weight.</p><p>Purely as storage arithmetic:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;2.4\\text{T}\\times\\frac{1.69}{8}\n=507\\text{ GB}&quot;,&quot;id&quot;:&quot;LKVCVHKHQG&quot;}" data-component-name="LatexBlockToDOM"></div><h2>Could Qwen3.8 run &#8220;locally&#8221;?</h2><p>It will not be a normal desktop model. Even the Q4 version would be roughly twenty times larger than a 70-billion-parameter Q4 model.</p><h3>Full BF16</h3><p>A BF16 deployment would need at least 4.8 TB for weights and roughly 5.5 TB after adding a modest 15% operational allowance.</p><p>That points toward configurations such as:</p><ul><li><p>Approximately <strong>24 B300 GPUs</strong>, assuming 288 GB each.</p></li><li><p>Approximately <strong>32 B200 GPUs</strong>, assuming 180 GB each.</p></li></ul><p>NVIDIA&#8217;s eight-GPU DGX B200 provides 1.44 TB of HBM and can be configured with 2&#8211;4 TB of system RAM. NVIDIA&#8217;s Blackwell Ultra B300 provides 288 GB of HBM per GPU, or approximately 2.3 TB in an eight-GPU system.</p><p>BF16 Qwen3.8 is firmly a multi-server deployment. </p><h3>NVFP4, Q4, and TQ1</h3><p>A 1.4&#8211;1.5 TB quantized model is more approachable, but &#8220;approachable&#8221; still means rack-scale hardware.</p><p>An <strong>eight-GPU B300 server with approximately 2.3 TB of HBM</strong> should have enough room for the quantized weights, runtime allocations, and a reasonable KV cache.</p><p>An eight-GPU B200 system has only 1.44 TB of HBM. It might barely load the 1.387 TB mixed-NVFP4 estimate, but it would leave almost no memory for the inference engine or KV cache. A two-node, 16-B200 configuration would be much more practical.</p><p>A 507 GB TQ1 checkpoint would still need roughly 600 GB or more after runtime overhead.</p><p>It could fit in:</p><ul><li><p>A 768 GB or 1 TB RAM server.</p></li><li><p>Four B200 GPUs with sufficient aggregate memory.</p></li><li><p>Three B300 GPUs in a custom configuration.</p></li><li><p>An eight-GPU node with considerable unused capacity.</p></li></ul><p>A current maximum-memory Mac Studio offers up to 512 GB of unified memory, which would be too tight once runtime overhead and the KV cache are included.</p><p>So, running even the most compressed version of Qwen3.8 won&#8217;t be cheap.</p><p>Hopefully, Alibaba will also release smaller versions of Qwen3.8.</p><p></p>]]></content:encoded></item><item><title><![CDATA[Inkling, Gemma 4 Updates, and 1-Bit Qwen3.6]]></title><description><![CDATA[The Weekly Kaitchup #151]]></description><link>https://kaitchup.substack.com/p/inkling-gemma-4-updates-and-1-bit</link><guid isPermaLink="false">https://kaitchup.substack.com/p/inkling-gemma-4-updates-and-1-bit</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 18 Jul 2026 00:29:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of <em>The Weekly Kaitchup</em>, we will discuss:</p><ul><li><p>Inkling and the Return of Global Competition in Open-Weight AI</p></li><li><p>Gemma 4: Small Updates with Significant Improvements</p></li><li><p>Bonsai Binary and Ternary Qwen3.6: Only 3.9 GB and It Still Works!</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Inkling and the Return of Global Competition in Open-Weight AI</h2><p>GLM, Kimi, MiniMax, and many of the largest recent open-weight language models have come from China. Outside China, NVIDIA has been one of the few companies to release a comparably large model recently, with Nemotron 3 Ultra.</p><p>Against this backdrop, the release of Inkling by Thinking Machines is a welcome development. Strong open-weight models from a wider range of organizations and countries are important for maintaining healthy global competition in frontier AI.</p><ul><li><p><a href="https://huggingface.co/thinkingmachines/Inkling">thinkingmachines/Inkling</a></p></li></ul><p>Inkling is a 975-billion-parameter multimodal model with approximately 41 billion parameters active per token. Its most interesting technical contributions are not its mixture-of-experts architecture or hybrid attention, both of which are now relatively common, but its approach to long-context modeling, optimization, and reinforcement learning.</p><h3>Long-context architecture</h3><p>Inkling uses learned relative positional representations instead of rotary positional embeddings. Thinking Machines reports that this approach performed better in its long-context extrapolation experiments.</p><p>The model combines these representations with an attention pattern consisting of five sliding-window attention layers for every global-attention layer. Most computation remains local, while periodic global-attention layers allow information to move across the full context. This idea of global-local layers is somewhat close to what Gemma 4 does.</p><p>Short convolutions are also applied after the attention key and value projections, as well as to the outputs of the attention and MLP branches. This introduces explicit local sequence mixing at several points within each transformer block, rather than relying exclusively on attention.</p><h3>Mixture-of-experts design</h3><p>Each mixture-of-experts layer contains:</p><ul><li><p>256 routed experts</p></li><li><p>Two shared experts (so, one more than most open-weight models)</p></li><li><p>Six routed experts selected for each token</p></li></ul><p>The outputs of the shared and routed experts are normalized together. This allows the routing mechanism to control the relative contribution of general-purpose shared computation and more specialized expert computation.</p><h3>Parameter-specific optimization</h3><p>Inkling uses Muon for large matrix parameters and Adam for other parameter types. Its weight decay is also adjusted alongside the learning-rate schedule instead of remaining fixed throughout training.</p><p>This represents a notably large-scale example of assigning different optimizers according to parameter structure.</p><h3>Learning to reason within a compute budget</h3><p>Thinking Machines reports running more than 30 million asynchronous rollouts while varying both the requested level of reasoning effort and the cost associated with generating additional tokens.</p><p>The objective is to teach the model to adapt its reasoning strategy to a given compute budget. This differs from simply imposing a maximum output length after training. Instead, the model learns to estimate when additional reasoning is likely to improve its answer and when the expected benefit is not worth the additional token cost.</p><p>Thinking Machines also randomized the available tools and their schemas during training. The goal was to prevent the model from becoming overly dependent on fixed function names, argument structures, or a single tool-calling format.</p><p>This could help the model generalize more reliably across unfamiliar tools and changing software environments.</p><h3>Competitive, but not yet the leader</h3><p>Inkling performs competitively with other leading open-weight models. However, on the reported benchmarks, GLM-5.2 remains slightly stronger across most evaluations despite being smaller.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ui-7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ui-7!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 424w, /__u/substackcdn.com/image/fetch/$s_!ui-7!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 848w, /__u/substackcdn.com/image/fetch/$s_!ui-7!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ui-7!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ui-7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png" width="956" height="1174" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1174,&quot;width&quot;:956,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:147740,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/207328806?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ui-7!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 424w, /__u/substackcdn.com/image/fetch/$s_!ui-7!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 848w, /__u/substackcdn.com/image/fetch/$s_!ui-7!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ui-7!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>They also released an NVFP4-quantized version that requires just 592 GB of memory, compared with roughly 1.9 TB for the original model.</p><ul><li><p><a href="https://huggingface.co/thinkingmachines/Inkling-NVFP4">thinkingmachines/Inkling-NVFP4</a></p></li></ul><p>The smallest GGUF made by Unsloth requires 270 GB + KV cache. I expect it to work very well despite the heavy compression since at that scale most of the parameters are routed experts, which are usually very robust to quantization.</p><ul><li><p><a href="https://huggingface.co/unsloth/inkling-GGUF">unsloth/inkling-GGUF</a></p></li></ul><div><hr></div><h2>Gemma 4: Small Updates with Significant Improvements</h2><p>The recent Gemma 4 changes are mainly runtime and interface updates rather than a new checkpoint.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;edc2f68d-ac52-4b6d-bcbe-71a79a88e727&quot;,&quot;caption&quot;:&quot;Last week, we compared Gemma 4 31B with Qwen3.5 27B and found that Gemma 4 31B outperformed it on most tasks while also being faster and more token-efficient.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Gemma 4 31B Quantization Comparison: Best FP8, NVFP4, and INT4 Models&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-04-20T18:00:18.907Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!rbnP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3c894d4-7925-49aa-991c-e71e4114b4e4_1210x843.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/gemma-4-31b-quantization-comparison&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:194307038,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:14,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>One update is broader FlashAttention 4 support on NVIDIA Hopper GPUs. Google reports higher prompt-processing throughput and lower time to first token. The main benefit is during prefill, when the model processes the complete prompt before generating its first output.</p><p>Gemma 4 also now exposes several visual-token budgets per image. Developers can use smaller allocations for classification, captioning, or video frames and larger allocations for OCR, documents, handwriting, or images containing small text.</p><p>This makes visual resolution a direct inference control. More visual tokens preserve more detail but also increase prompt length, prefill time, cache use, and the amount of context consumed by each image. The setting is useful for workloads in which some images require detailed analysis while others do not.</p><p>The other relevant change concerns chat and tool-use templates. These templates convert structured messages into the special-token format expected by the model.</p><p>Recent fixes address turn boundaries, tool-response handling, structured arguments, multimodal messages, and reasoning continuity across multiple tool calls. These are implementation changes, but they can directly affect agent reliability. An incorrect marker may cause the model to treat a tool result as a user message or to restart a reasoning sequence instead of continuing it.</p><p>Google reports accuracy improvements on benchmarks using tools thanks to these updates:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!zHo5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!zHo5!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!zHo5!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!zHo5!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!zHo5!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!zHo5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="/__u/substackcdn.com/image/fetch/$s_!zHo5!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!zHo5!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!zHo5!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!zHo5!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Important:</strong> The Gemma 4 repositories were updated in place. As a result, reproducing earlier evaluation scores requires loading the models with the corresponding revision explicitly specified.</p><div><hr></div><h2>Bonsai Binary and Ternary Qwen3.6: Only 3.9 GB and It Still Works!</h2><p>PrismML&#8217;s Bonsai applies binary and ternary weights to most of the language network of Qwen3.6-27B.</p><ul><li><p><a href="https://huggingface.co/collections/prism-ml/bonsai-27b">prism-ml/bonsai-27b</a></p></li></ul><p>The ternary version uses negative, zero, and positive weight states and occupies about 5.9GB. The binary version uses only negative and positive states and occupies about 3.9GB. Both use one half-precision scale for groups of 128 weights.</p><p>The low-bit format covers embeddings, attention projections, MLP layers, and the final language-model head. Many quantized models keep sensitive layers at higher precision, especially embeddings or output heads. Bonsai does not use those higher-precision exceptions in the language network.</p><p>The vision tower remains in four-bit precision. This suggests that the visual encoder is more sensitive to binary or ternary quantization than the language network. Vision performance still declines, likely because errors accumulate across the visual encoder, multimodal projection, and low-bit decoder.</p><p>Bonsai also requires custom CUDA and MLX kernels. A compressed weight file does not automatically produce fast inference if the runtime first expands the weights into a higher-precision matrix.</p><p>The custom kernels operate directly on the packed representation. Binary weights can be handled mainly as sign changes, while ternary weights add a zero state that can represent an absent connection. Accumulation and activations still require higher precision.</p><p>The reported phone deployment should be interpreted mainly as a weight-memory result. A 3.9GB model can fit within the memory range of a high-end mobile application, but runtime memory must also include the KV cache, activations, workspaces, the vision tower, and operating-system overhead. The full advertised context length is unlikely to be practical on a phone.</p><p>Despite the heavy compression, the model remains very capable:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!q4jg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!q4jg!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 424w, /__u/substackcdn.com/image/fetch/$s_!q4jg!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 848w, /__u/substackcdn.com/image/fetch/$s_!q4jg!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 1272w, /__u/substackcdn.com/image/fetch/$s_!q4jg!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!q4jg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png" width="1201" height="555" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:555,&quot;width&quot;:1201,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:95787,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/207328806?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!q4jg!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 424w, /__u/substackcdn.com/image/fetch/$s_!q4jg!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 848w, /__u/substackcdn.com/image/fetch/$s_!q4jg!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 1272w, /__u/substackcdn.com/image/fetch/$s_!q4jg!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The degradation compared with the original model is significant, but the accuracy is far better than all the other sub-4 GB models I have evaluated.</p><p>I&#8217;ll publish full analysis, accuracy, and token efficiency next week!</p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisionning it with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/kaitchup.substack.com/subscribe"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><p></p>]]></content:encoded></item><item><title><![CDATA[Qwen3.6-27B KV Cache Quantization in vLLM: Accuracy, Memory, and Speed]]></title><description><![CDATA[A smaller KV cache enables longer sequences and higher concurrency with virtually no loss in accuracy.]]></description><link>https://kaitchup.substack.com/p/qwen36-27b-kv-cache-quantization</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen36-27b-kv-cache-quantization</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Thu, 16 Jul 2026 16:24:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!alUj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dc43aa-10ea-4534-9c9e-0ba78a52affe_1210x843.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Long-context inference is often limited less by the model&#8217;s advertised context window than by the memory and bandwidth required for its KV cache.</p><p>During generation, attention layers store keys and values for every previous token. As context length and concurrency increase, this cache can consume tens of gigabytes and become a major memory-bandwidth bottleneck.</p><p>KV-cache quantization reduces that cost by storing keys and values at lower precision. Moving from 16-bit to approximately 4-bit storage can shrink the cache by nearly four times, enabling longer contexts and more concurrent requests.</p><p>In practice, however, a method that works well in a research implementation may perform poorly inside an optimized engine such as vLLM or llama.cpp. Quantization can add computational overhead, require specialized kernels, reduce hardware compatibility, and interfere with features such as speculative decoding.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In this article, I compare three 4-bit KV-cache formats in vLLM, <code>turboquant_4bit_nc</code>, <code>int4_per_token_head</code>, and <code>nvfp4</code>, using Qwen3.6-27B. I evaluate their memory usage, accuracy, token generation, and inference speed.</p><blockquote><p><strong>Acknowledgments</strong></p><p><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=kvquantization">Verda</a> provided the B200s and RTX Pro 6000s to run the experiments described in this article.</p><p>Verda is the full-stack frontier AI cloud, built for high-performance inference, training, and agentic workloads with data privacy and sustainability at its core.</p><p><span>You can check them out </span><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=kvquantization">here</a><span>. There is a </span><strong>$50 coupon</strong><span> that you can redeem in your Verda account, after provisionning it with $5, to try their GPUs. </span></p><p><strong>Coupon code:</strong><span> </span><code>KAITCHUP-50</code></p><p><span>Follow </span><a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a><span>.</span></p></blockquote><h3>Qwen3.6-27B KV Cache Size</h3><p>I&#8217;m going to compare KV-cache memory usage across different quantization methods. First, let&#8217;s establish the memory required for Qwen3.6-27B&#8217;s native 262,144-token context window.</p><p>Qwen3.6-27B has 64 language layers, but only 16 are Gated Attention layers that maintain a full KV cache. The remaining layers are Gated DeltaNet blocks, so they are not included when calculating the conventional attention KV cache.</p><p>Each Gated Attention layer has four KV heads, with a head dimension of 256. </p><p>In BF16, the KV cache for these full-attention layers consumes approximately 17.2 GB. This is substantially less than for models that use full attention in every layer, but it is still significant relative to the model weights themselves, which occupy roughly 54 GB.</p><p>On a device with 96 GB of memory, that leaves:</p><p>96 &#8722; 54 &#8722; 17.2 = 24.8 GB</p><p>This is not enough to run several full-context requests in parallel, for example, multiple sub-agents, especially once additional overhead is included. CUDA graphs, temporary buffers, runtime allocations, and other inference-engine components all consume extra memory.</p>
      <p>
          <a href="/__u/kaitchup.substack.com/p/qwen36-27b-kv-cache-quantization">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Efficient and Reasoning AI at the ACL 2026]]></title><description><![CDATA[The Weekly Kaitchup #150]]></description><link>https://kaitchup.substack.com/p/efficient-and-reasoning-ai-at-the</link><guid isPermaLink="false">https://kaitchup.substack.com/p/efficient-and-reasoning-ai-at-the</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 11 Jul 2026 00:46:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of <em>The Weekly Kaitchup</em>, I&#8217;ll highlight some of the most interesting work I saw at ACL 2026, which I attended this week in San Diego.</p><p>I&#8217;ll also discuss OpenAI&#8217;s recent take on benchmarking with SWE-Bench Pro, one of the most widely used coding benchmark. Once again, a SWE-Bench is being declared &#8220;bad.&#8221; I can&#8217;t say I&#8217;m too upset about having never spent compute on it.</p><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p><span>In collaboration with </span><a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a><span>, I&#8217;m sharing a </span><strong>$50 coupon</strong><span> that you can redeem in your Verda account, after provisionning it with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</span></p><p><strong>Coupon code:</strong><span> </span><code>KAITCHUP-50</code></p><p><span>Follow </span><a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a><span>.</span></p></blockquote><h2>ACL 2026: Why I Went, and Why I Skipped ICML 2026</h2><p>This year, two of the major annual research conferences shaping the development of AI, ACL and ICML, took place at the same time, on opposite sides of the Pacific: ACL in San Diego and ICML in Seoul.</p><p>While there is significant, and increasing, overlap between the two communities, they tend to emphasize different areas. ICML is often the place for deeper discussions about machine learning methods, algorithms, and theoretical foundations. ACL, by contrast, is where I usually find the most relevant work on evaluation, multilinguality, benchmarks, datasets, and language-centered applications.</p><p>That distinction is far from absolute. Many of the key ideas behind today&#8217;s Transformers and LLMs, as well as many of the benchmarks and evaluation tasks we still rely on, originated in papers published within the ACL community.</p><p>I chose to attend ACL this year primarily because it aligns more closely with my own research background. It is also the community where I know more people, having published several papers there during my Ph.D.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>ACL 2026 was capped at 3,500 onsite attendees. For reference, ACL 2025 in Vienna had more than 5,000 in-person attendees.</p><p>The smaller scale made ACL easier to navigate, but it was still impossible to see everything I had planned. As I did for NeurIPS 2025, I will mainly focus on the papers I actually saw and discussed at the conference rather than trying to summarize the full program.</p><p>I spent most of my time in the poster sessions, which were by far the best place to have useful conversations with authors. The talks were less crowded, and some large rooms were surprisingly almost empty, especially on the last day.</p><p>To keep this edition useful and focused, I will mostly discuss papers on LLM reasoning and efficiency, two of the major themes of ACL 2026 along with reinforcement learning, as it has been at most recent AI and machine learning conferences. </p><p>The program chairs even noted that papers on some of these topics seemed to have higher acceptance rates, which suggests that they were not only popular among authors, but also particularly appealing to reviewers and area chairs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!7KzG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!7KzG!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!7KzG!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!7KzG!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!7KzG!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!7KzG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg" width="588" height="783.8653846153846" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:588,&quot;bytes&quot;:4228692,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!7KzG!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!7KzG!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!7KzG!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!7KzG!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Note: For each paper discussed below, I also included a photo of the corresponding poster. The images are high-resolution, but I do not fully trust Substack to display them cleanly (depending on your device), so apologies if some details are difficult to read. ACL may also publish the posters directly on each paper&#8217;s page, or may have already done so by the time you read this. You can check by clicking the paper links.</em></p><div><hr></div><p><a href="https://aclanthology.org/2026.findings-acl.1717">Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models</a><br><em>Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei<br>Meta; New York University; Johns Hopkins University</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!prGW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!prGW!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!prGW!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!prGW!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!prGW!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!prGW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg" width="662" height="496.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:662,&quot;bytes&quot;:1662114,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!prGW!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!prGW!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!prGW!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!prGW!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This paper looks at a well-known problem in reinforcement learning for reasoning models: training that improves reasoning can also make the model forget broader capabilities. I often see that when running RL for recent models like Qwen3.5/Gemma 4. </p><p>The authors show that RLVR-style post-training may improve mathematical or multimodal reasoning while degrading skills such as perception, OCR, robustness, and general instruction following.</p><p>They propose RECAP, a replay-based method that keeps general-capability data in the training loop and dynamically reweights objectives. The main contribution is a clearer diagnosis of &#8220;reasoning gain&#8221; as a trade-off problem: a model can become better at narrow reasoning benchmarks while becoming less generally useful. RECAP is presented as a way to preserve the original model&#8217;s breadth while still gaining reasoning ability.</p><p>Very simple and intuitive.</p><div><hr></div><p><a href="https://aclanthology.org/2026.acl-long.1530/">SSSD: Simply-Scalable Speculative Decoding</a><br><em>Michele Marzollo, Jiawei Zhuang, Niklas Roemer, Niklas Zwingenberger, Lorenz K Muller, Lukas Cavigelli<br>Huawei; ETH Zurich</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!CCGF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!CCGF!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!CCGF!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!CCGF!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!CCGF!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!CCGF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg" width="1456" height="1941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1894802,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!CCGF!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!CCGF!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!CCGF!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!CCGF!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This paper proposes SSSD, a training-free speculative decoding method for making LLM inference faster.</p><p>Instead of relying on an additional trained draft model, like EAGLE3 or DFlash, SSSD retrieves likely n-gram continuations from the prompt, previously generated tokens, and a datastore, then verifies them with the main model.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;4d72fac1-b9b9-4e02-bc16-03acf1a974e9&quot;,&quot;caption&quot;:&quot;Speculative decoding is becoming a popular way to accelerate LLM inference, with approaches such as MTP and DFlash. The idea is simple: a smaller or specialized draft model proposes future tokens, and the full target model verifies them. Matching tokens are accepted and when a mismatch occurs, only the valid prefix is kept and decoding falls back to the target model.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Train and Run DFlash Speculative Decoding&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-18T19:43:12.333Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sbvO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b274a34-f749-4a9f-b68b-0e0f501a9016_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/train-and-run-dflash-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196847181,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>It avoids extra model training and is easier to deploy across changing workloads, domains, and languages. </p><p>The paper reports latency reductions of up to 2.9x over standard autoregressive decoding, making it relevant for production serving where simplicity and robustness matter as much as peak speed.</p><div><hr></div><p><a href="https://aclanthology.org/2026.findings-acl.677/">Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction</a><br><em>Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, Shafiq Joty<br>Salesforce AI Research; University of Wisconsin-Madison; The University of Hong Kong</em></p><p>Early layers of an LLM can already identify many of the tokens that will matter for answering a query. The authors turn this into GemFilter, a training-free method that uses early layers as filters to select a much smaller subset of the input before later computation.</p><p>The method is aimed at long-context inference, where latency and GPU memory grow quickly with input length. By compressing the input token set, GemFilter speeds up inference while keeping the important evidence visible to the model.</p><p>A nice aspect of the work is interpretability: because the method explicitly selects input tokens, humans can inspect what the model kept. The paper reports a 2.4x speedup and 30% GPU memory reduction compared with strong baselines, while performing well on Needle-in-a-Haystack and competitively on LongBench. </p><div><hr></div><p><a href="https://doi.org/10.1162/TACL.a.692">A Survey on Memory-Efficient Fine-Tuning for Large Language Models</a><br><em>Yeachan Kim, Mingyu Lee, SangKeun Lee<br>Hankuk University of Foreign Studies; Korea University</em><br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!pVhK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!pVhK!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!pVhK!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!pVhK!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!pVhK!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!pVhK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg" width="1456" height="1941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1750655,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!pVhK!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!pVhK!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!pVhK!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!pVhK!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This survey organizes the landscape of memory-efficient fine-tuning for LLMs. I recommend reading it if you are lost in the recent progress made in this area. </p><p>One point I strongly agree with, although it is not very well developed in the paper, is that these methods are often poorly evaluated.</p><p>The paper mentions GLUE-style tasks. I did not double-check how widely GLUE is still used in this line of work, but even replacing it with old and relatively easy benchmarks such as QNLI, PIQA, or CommonsenseQA, like they did in this paper, would not solve the problem. People use these benchmarks because they are cheap to run.</p><p>As I often argue here, if your benchmarks do not require the model to generate millions of tokens, they are unlikely to tell you much about how a generative language model actually performs, unless of course you thouroughly study what the benchmarks has generated instead of just looking at the scores.</p><div><hr></div><h2>KV Compression</h2><p><a href="https://aclanthology.org/2026.acl-long.1683/">LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning</a><br><em>Haoyue Zhang, Hualei Zhang, Xiaosong Ma, Jie Zhang, Song Guo<br>The Hong Kong University of Science and Technology; The Hong Kong Polytechnic University</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!OPck!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!OPck!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!OPck!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!OPck!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!OPck!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!OPck!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg" width="656" height="874.5164835164835" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:656,&quot;bytes&quot;:1905771,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!OPck!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!OPck!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!OPck!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!OPck!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This paper studies KV-cache compression for long chain-of-thought reasoning. The key finding is &#8220;Token Importance Recurrence&#8221;: tokens that seem unimportant at one decoding step can become important again later. This makes simple greedy KV eviction risky, because it may remove tokens that the model will need in future reasoning steps.</p><p>LazyEviction addresses this by delaying eviction and observing attention patterns over a window before deciding what to discard. The method is designed for long reasoning sequences, where memory cost grows with every generated token, and it aims to reduce memory while avoiding sudden reasoning failures caused by prematurely evicting useful context.</p><p>This is another interesting piece of work on KV cache compression, one of many presented at the conference.</p><p>As usual, though, the real test will be implementation. Unless a method lands in frameworks such as vLLM or llama.cpp, it is hard to know whether it actually improves inference efficiency in practice, or whether it interferes with other optimizations already implemented in these systems.</p><div><hr></div><p><a href="https://aclanthology.org/2026.acl-long.1542/">Anchoring the Cache: Mitigating Contextual Hallucination in KV-Compressed Long-Context Summarization</a><br><em>Yu Fu, Chen Luo, Josef Valvoda, Xin Zhang, Xuejing Lei, Xiao Pan, Hui Liu, Yue Dong<br>University of California, Riverside; Amazon</em><br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Yn1p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Yn1p!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Yn1p!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Yn1p!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Yn1p!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Yn1p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg" width="1456" height="1941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1848955,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Yn1p!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Yn1p!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Yn1p!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Yn1p!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This paper studies an important side effect of KV-cache compression in long-context summarization: it can make models hallucinate more. </p><p>The authors show that aggressive compression may preserve surface-level summarization metrics while weakening factual grounding, because retrieval heads drift away from the source context during generation.</p><p>Their method, HalluKV, anchors selected retrieval heads to the source by selectively removing generated KV pairs from those heads during decoding. The idea let the model keep the efficiency benefits of cache compression, but prevent the heads responsible for source retrieval from over-attending to the model&#8217;s own generated text.</p><p>The result is a decoding-time intervention that reduces hallucination while maintaining long-context efficiency. The paper reports that KV compression can increase hallucination scores by up to 3.36&#215;, and that HalluKV reduces hallucination across multiple models and datasets while preserving competitive summary quality. </p><div><hr></div><p><a href="https://aclanthology.org/2026.findings-acl.1314/">Quantize What Counts: More for Keys, Less for Values</a><br><em>Mohsen Hariri, Alan Luo, Weicong Chen, Tianyi Zhang, Qifan Wang, Xiaotian Han, Vipin Chaudhary<br>Case Western Reserve University; Rice University; Meta</em><br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!AyUv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AyUv!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!AyUv!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!AyUv!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!AyUv!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AyUv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg" width="1456" height="1941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2121396,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AyUv!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!AyUv!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!AyUv!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!AyUv!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Keys and values of the KV cache should not necessarily receive the same number of bits: key projections tend to have larger spectral and Frobenius norms than value projections, meaning keys carry higher-magnitude, more quantization-sensitive information. </p><p>The authors use this observation to justify a simple rule: allocate more precision to keys and compress values more aggressively.</p><p>The paper turns this into a geometry-driven design principle for mixed-precision KV-cache quantization. Instead of tuning key/value bit splits heuristically for each model or task, it argues that model weight geometry can predict which cache component needs more protection. Empirically, key-favored allocations such as 4-bit keys and 2-bit values preserve up to 98.3% of the accuracy of uniform 4-bit KV quantization while saving memory, making the result useful for long-context inference where KV cache memory dominates.</p><p>A very informative paper to better understand why keys are more sensitive to quantization.</p><div><hr></div><h2>More Comments on the ACL 2026</h2><p>The social events at the ACL are always great. This year, the organizers reserved the <a href="https://uss-midway-museum.sandiegotourismtickets.com/">USS Midway</a> for an evening dinner. The USS Midway is an aircraft carrier turned museum, and the venue was genuinely impressive.</p><p>These social events are often among the most useful parts of a conference. They make it much easier to meet people outside your usual circle, and this one was no exception.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ko6e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ko6e!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Ko6e!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Ko6e!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Ko6e!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ko6e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:696600,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ko6e!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!Ko6e!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!Ko6e!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!Ko6e!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The next ACL will be in Kyoto, so I will probably end up playing local guide: I lived and worked there for six years.</p><p>Hopefully, the conference will not coincide with the worst of Kyoto summer: 35&#176;C or more, 70%+ humidity, and the occasional typhoon. Otherwise, the most important survival tips may have nothing to do with AI.</p><div><hr></div><h2>SWE-Bench Is Not a Good Benchmark, Again</h2><p><em>Unrelated to the ACL, but very interesting findings published by OpenAI this week.</em></p><p>OpenAI has used SWE-Bench Pro to benchmark all its models since GPT 5.2.</p><p>And yet, according to <a href="https://openai.com/index/separating-signal-from-noise-coding-evaluations/">OpenAI&#8217;s recent audit</a>, many SWE-bench Pro tasks fail for reasons unrelated to model capability.</p><p>Examples include:</p><ul><li><p>hidden requirements that never appear in the issue description</p></li><li><p>contradictory specifications</p></li><li><p>unit tests that reject perfectly reasonable implementations</p></li><li><p>grading logic that expects one specific implementation instead of checking whether the issue is actually fixed</p></li><li><p>mismatches between the GitHub issue, the merged pull request, and the evaluation tests</p></li></ul><p>Because SWE-bench Pro is built automatically from real repositories, OpenAI argues that these inconsistencies are common enough to substantially affect reported scores.</p><p>OpenAI&#8217;s main conclusion is:</p><blockquote><p><strong>Approximately 30% of SWE-bench Pro tasks are broken.</strong></p></blockquote><p>Their recommendation is therefore to &#8220;<strong>carefully inspect benchmark results instead of treating leaderboard rankings as ground truth.&#8221; </strong>Yes, of course, but who actually does this? Most AI labs want to report the strongest possible numbers. If you publish a score of 80 and your competitor publishes 90, most people will not ask whether that 10-point gap reflects a real performance difference or a benchmarking artifact. They will simply see the lower score, and it will look bad.</p><p>Ironically, OpenAI itself helped create <strong>SWE-bench Verified</strong> in 2024 because it believed the original SWE-bench contained many impossible or ambiguous tasks. Now, after auditing SWE-bench Pro, it argues that even this newer benchmark still contains enough flawed tasks that raw leaderboard scores should not be over-interpreted.</p><p>My take is that these benchmarks are becoming increasingly complex: automatically generated, frequently refreshed, available in multiple versions, and often difficult to interpret consistently. As a result, I think people will develop more and more trust issues around published scores and may eventually start disregarding them altogether.</p><p>However, this is less of a problem when evaluating different versions of the same model. For example, when comparing a quantized model to the original, the goal is not necessarily to establish the model&#8217;s absolute performance. What matters is whether the quantized version performs as closely as possible to the non-quantized one.</p><p>In that setting, the original score matters less. The main objective is to minimize the accuracy delta between the two versions. Even a flawed benchmark like SWE-Bench can still be useful for this purpose, as long as it is used consistently across both model variants.</p><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/kaitchup.substack.com/subscribe"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><p></p>]]></content:encoded></item><item><title><![CDATA[LFM2.5 230M and 350M: How Accurate Are the GGUF Versions?]]></title><description><![CDATA[230M or 350M GGUFs?]]></description><link>https://kaitchup.substack.com/p/lfm25-230m-and-350m-how-accurate</link><guid isPermaLink="false">https://kaitchup.substack.com/p/lfm25-230m-and-350m-how-accurate</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Wed, 08 Jul 2026 03:20:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2bDO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Tiny language models that can run quickly with very limited memory are one of Liquid AI&#8217;s specialties. And while models with only a few hundred million parameters obviously cannot rival today&#8217;s multi-billion-parameter models, tiny models are now capable enough for many practical tasks. They can make tool calls, retain some basic world knowledge, and follow instructions surprisingly well.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!2bDO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!2bDO!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 424w, /__u/substackcdn.com/image/fetch/$s_!2bDO!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 848w, /__u/substackcdn.com/image/fetch/$s_!2bDO!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2bDO!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!2bDO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png" width="1456" height="978" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:978,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!2bDO!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 424w, /__u/substackcdn.com/image/fetch/$s_!2bDO!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 848w, /__u/substackcdn.com/image/fetch/$s_!2bDO!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2bDO!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">https://huggingface.co/LiquidAI/LFM2.5-230M</figcaption></figure></div><p>LFM2.5 is a family of small language models. In this article, I will focus specifically on the tiny 230M and 350M versions. I ran both models on a large set of benchmarks with two goals:</p><ol><li><p>Reproduce some of the results published by Liquid AI.</p></li><li><p>Better highlight what these models can and cannot do.</p></li></ol><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>At first glance, the size difference between the 350M and 230M models may seem negligible: only 120M parameters. But that gap feels much larger when you realize that 120M parameters represent roughly 34% of the 350M model&#8217;s total size. On memory-constrained devices, especially when using longer context lengths, this difference can matter a lot. In 16-bit precision, it reduces weight memory consumption from about 0.71 GB for the 350M model to about 0.46 GB for the 230M model.</p><p>And given how accurate is quantization now, it could be possible to make a quantized version of the 350M that is both smaller and better than the 230M.</p><p>This raises another question that I will address in this article:</p><p><em>Should you use a low-bit quantized version of the 350M model, or a higher-precision version of the 230M model?</em></p><p>To answer that, we will compare accuracy versus model size for both models across different GGUF quantization formats released by Liquid AI and Unsloth.</p><h2>LFM2.5 230M and 350M: What They Can and Cannot Do</h2>
      <p>
          <a href="/__u/kaitchup.substack.com/p/lfm25-230m-and-350m-how-accurate">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[DSpark and NVIDIA's Qwen3.6 NVFP4 Models]]></title><description><![CDATA[The Weekly Kaitchup #149]]></description><link>https://kaitchup.substack.com/p/dspark-and-nvidias-qwen36-nvfp4-models</link><guid isPermaLink="false">https://kaitchup.substack.com/p/dspark-and-nvidias-qwen36-nvfp4-models</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 04 Jul 2026 03:21:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c7_f!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of The Weekly Kaitchup:</p><ul><li><p>Which NVFP4 Version of Qwen3.6 27B Should You Use?</p></li><li><p>DSpark: DeepSeek&#8217;s New Speculative Decoding Method</p></li></ul><p>I&#8217;m at the ACL 2026 and stopped by the Qwen corner to ask the important question:</p><blockquote><p><strong>Should we expect Qwen3.7 27B soon?</strong></p><p>Answer:<br><em>We haven&#8217;t made a final decision yet, but we&#8217;re committed to open-sourcing more models and have some exciting releases coming soon.</em></p></blockquote><p>Not very informative&#8230; It sounds like a well-prepared answer. And they told me they already got this question a lot during just a single morning at the conference.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!D0d3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!D0d3!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!D0d3!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!D0d3!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!D0d3!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!D0d3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg" width="550" height="412.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:550,&quot;bytes&quot;:4715993,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/204695456?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!D0d3!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!D0d3!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!D0d3!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!D0d3!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Which NVFP4 Version of Qwen3.6 27B Should You Use?</h2><p>NVIDIA released an NVFP4 version of Qwen3.6 27B. It is a good occasion to revisit the other NVFP4 variants that have been released and to put NVIDIA&#8217;s checkpoint into context.</p><p>When I published my evaluations and analysis of quantized Qwen3.6 27B, only a few NVFP4 versions were available. There were so few that I had to make one myself with <a href="https://github.com/intel/auto-round">AutoRound</a>.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;bcf583e6-8211-4164-bd6d-c6ee65102265&quot;,&quot;caption&quot;:&quot;Like I did for Qwen3.5 and Gemma 4, let&#8217;s see how well Qwen3.6 27B holds up when quantized into different formats: FP8, INT4, and NVFP4.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwen3.6 27B Quantization: FP8 vs INT4 vs NVFP4&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-12T06:45:55.916Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!5Qvj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b412df9-eef1-4bfe-aa6b-b90105838569_1210x903.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwen36-27b-quantization-fp8-vs-int4&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196100250,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:15,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p><strong><a href="https://huggingface.co/kaitchup/Qwen3.6-27B-autoround-nvfp4-linearattn-BF16">kaitchup/Qwen3.6-27B-autoround-nvfp4-linearattn-mtp-BF16</a></strong><br>This is my version made with AutoRound. It uses NVFP4 while keeping the linear-attention and MTP layers in 16-bit. The model size is <strong>28.6 GB</strong>.</p><p><strong><a href="https://huggingface.co/Peutlefaire/Qwen3.6-27B-NVFP4">Peutlefaire/Qwen3.6-27B-NVFP4</a></strong><br>This version goes further and quantizes more <code>linear_attn</code> layers too. The model size is <strong>20.6 GB</strong>, with a 19.7 GB main model file and an 849 MB MTP file. Compared with my AutoRound MTP BF16 version, that is roughly 8.0 GB smaller.</p><p><a href="/__u/kaitchup.substack.com/p/qwen36-27b-quantization-fp8-vs-int4">In my analysis</a>, I showed that quantizing the attention path to NVFP4 significantly degraded the model. It did not break the model, but it made it more prone to endless thinking, and accuracy went down on nearly all benchmarks. Check the full analysis for details.</p><p><strong><a href="https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4">unsloth/Qwen3.6-27B-NVFP4</a></strong><br>Unsloth also released an NVFP4 checkpoint. Its config applies NVFP4 quantization to most Linear layers, while leaving some of the linear-attention projections (linear_attn.out_proj) in BF16. You can see it as a version between mine and a full NVFP4 quantization. The model size is <strong>26.4 GB</strong>.</p><p><strong><a href="https://huggingface.co/nvidia/Qwen3.6-27B-NVFP4">nvidia/Qwen3.6-27B-NVFP4</a></strong><br>The NVIDIA version is interesting because it is mixed precision: attention and linear-attention projections are FP8 W8A8, while the MLP layers and LM head are marked as <strong>W4A16_NVFP4</strong>. The checkpoint is <strong>21.9 GB</strong>, which makes it <strong>6.7 GB smaller</strong> than my 28.6 GB AutoRound MTP BF16 version and <strong>4.5 GB smaller</strong> than Unsloth&#8217;s 26.4 GB version.</p><p>This also means that the NVFP4 parts of NVIDIA&#8217;s model do not use NVFP4 activations. They are W4A16: 4-bit NVFP4 weights, but 16-bit activations. The attention path, meanwhile, is FP8 rather than NVFP4. With Blackwell GPUs, this likely leaves speed on the table compared with W4A4 NVFP4. In my experience, NVFP4 activations can be almost harmless with good calibration.</p><p>Another important detail: NVIDIA&#8217;s <code>config.json</code> includes an FP8 KV-cache scheme. That matters a lot for speed and memory use, especially on devices where memory bandwidth is the bottleneck. I already saw some third-party speed comparisons on X comparing this checkpoint against other NVFP4 models without matching the KV-cache format, which is not an apples-to-apples comparison. At minimum, KV-cache dtype should be controlled explicitly across all models; vLLM, for example, can be run with an explicit FP8 KV cache.</p><p>NVIDIA also reports its own evaluation numbers against an FP8 baseline. In their table, the NVFP4 checkpoint is very close to the FP8 checkpoint across MMLU-Pro, GPQA Diamond, HLE, &#964;&#178;-Bench, MMMU Pro, SciCode, AIME 2025, AA-LCR, and IFBench. This is useful, but it is still not a community-wide comparison against the other NVFP4 checkpoints under identical settings.</p><p>It is also worth mentioning the PrismaQuant family.</p><p><strong><a href="https://huggingface.co/rdtand/Qwen3.6-27B-PrismaQuant-5.5bit-vllm">rdtand/Qwen3.6-27B-PrismaQuant-5.5bit-vllm</a></strong><br>This version uses a sensitivity-driven per-Linear allocation. PrismaQuant chooses between NVFP4, MXFP8, and BF16 under a 5.5-bit-per-parameter target. According to the model card, bulk dense MLP layers are usually NVFP4 W4/A4, higher-sensitivity dense linears can be MXFP8 W8/A8, and the most sensitive tensors, norms, biases, embeddings, and LM head stay BF16.</p><p>The model size is 22.7 GB, which slighly larger to the one released by NVIDIA, probably due to the LM head staying at higher precision.</p><p>There are newer PrismaQuant-family variants too. <strong><a href="https://huggingface.co/rdtand/Qwen3.6-27B-PrismaAURA-5.5bit-vllm">PrismaAURA 5.5-bit</a></strong><a href="https://huggingface.co/rdtand/Qwen3.6-27B-PrismaAURA-5.5bit-vllm"> </a>uses an AURA/KL-Fisher allocation over NVFP4, FP8, and BF16, while <strong><a href="https://huggingface.co/rdtand/Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm">PrismaSCOUT 5.31-bit</a></strong> is described as a Blackwell-oriented successor to the earlier 5.5-bit PrismaQuant artifact. PrismaSCOUT uses NVFP4 plus selected BF16, includes MTP tensors, and its repository is <strong>20.2 GB</strong>.</p><h2>How to Choose?</h2><p>We do not yet have a good comparison. Quantization is cheap; evaluation is expensive.</p><p>What we need is a benchmark suite run on all these checkpoints with exactly the same settings: same inference engine, same chat template, same thinking mode, same KV-cache dtype, same context length, same sampling settings, same MTP/speculative-decoding settings, and same hardware.</p><p>Until then, I would assume the following:</p><ul><li><p>There is no strong reason to keep all activations in 16-bit on Blackwell if NVFP4 activation quantization is well calibrated.</p></li><li><p>Keeping the attention path out of NVFP4 is safer.</p></li><li><p>Unsloth&#8217;s version and my AutoRound version are safer choices if your priority is a more conservative quantization strategy.</p></li><li><p>PrismaSCOUT and the other PrismaQuant-family models are worth testing, but I would not rank them without independent, apples-to-apples accuracy numbers. This could be the best version. We simply don&#8217;t know.</p></li></ul><p>So, for now, I would use either Unsloth&#8217;s version or my AutoRound MTP BF16 version if accuracy safety is the main concern. They are also better documented and easier to reason about than many one-off community checkpoints. If memory footprint matters more, NVIDIA&#8217;s version and PrismaSCOUT are especially interesting.</p><div><hr></div><h2>DSpark: DeepSeek&#8217;s New Speculative Decoding Method</h2><p>DSpark is a speculative decoding method for making LLMs generate faster without changing the final model that verifies the text. </p><p>Instead of asking the full model to produce one token at a time, DSpark, like any other speculative decoding method, lets a smaller draft module propose several future tokens, then asks the full model to verify those tokens in a batch. When the draft is right, the server accepts multiple tokens from one target-model pass. When the draft is wrong, the server keeps the longest valid prefix and continues from there.</p><p>The target model we want to accelerate remains the source of truth, while the drafter tries to guess what the target model is likely to accept next.</p><h2>What DSpark Changes</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!VX38!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!VX38!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!VX38!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!VX38!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!VX38!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!VX38!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;DSpark architecture: a parallel draft backbone, a lightweight serial head linking adjacent tokens, a confidence head scoring each token, and a hardware-aware scheduler choosing how much of the block to verify.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="DSpark architecture: a parallel draft backbone, a lightweight serial head linking adjacent tokens, a confidence head scoring each token, and a hardware-aware scheduler choosing how much of the block to verify." title="DSpark architecture: a parallel draft backbone, a lightweight serial head linking adjacent tokens, a confidence head scoring each token, and a hardware-aware scheduler choosing how much of the block to verify." srcset="/__u/substackcdn.com/image/fetch/$s_!VX38!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!VX38!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!VX38!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!VX38!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">https://deepseek.ai/blog/deepseek-dspark-speculative-decoding</figcaption></figure></div><p>The weakness of many speculative decoders is suffix decay. Drafting the first token is easy but drafting the fifth, sixth, or seventh token is much harder because every position depends on earlier guesses. Fully parallel drafters are fast, but later draft positions often become noisy. Autoregressive drafters are more faithful, but they give back too much of the latency advantage because they draft sequentially.</p><p>DSpark tries to sit between those two extremes. It keeps a DFlash-style parallel draft backbone, then adds a lightweight Markov logit-bias head. <em>Note: I explained how DFlash works here:</em></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;4fc4f756-15cc-4d6a-a2b0-a64c213eaa20&quot;,&quot;caption&quot;:&quot;Speculative decoding is becoming a popular way to accelerate LLM inference, with approaches such as MTP and DFlash. The idea is simple: a smaller or specialized draft model proposes future tokens, and the full target model verifies them. Matching tokens are accepted and when a mismatch occurs, only the valid prefix is kept and decoding falls back to the target model.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Train and Run DFlash Speculative Decoding&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-18T19:43:12.334Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sbvO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b274a34-f749-4a9f-b68b-0e0f501a9016_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/train-and-run-dflash-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196847181,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>That head lets each draft position condition on previously sampled tokens inside the draft block, which improves the later positions without turning the whole drafter into a slow autoregressive model.</p><p>The other important piece is confidence. DSpark adds a confidence head that predicts the acceptance probability of each draft position. That gives the serving system a practical scheduling signal. When the drafter is confident, the runtime can verify more proposed tokens. When confidence drops, or when the server is under load, it can avoid wasting verification work on tail tokens that are likely to be rejected. In other words, DSpark is not just a better drafter; it is a more serving-aware drafter.</p><h3>How it works in the decoding loop</h3><p>At each generation step, DSpark first prepares a block of draft tokens. The target model then scores that block in one verification pass. The runtime accepts the longest prefix that agrees with the target model&#8217;s sampling path and discards the rest. If several drafted tokens survive, the user sees several tokens produced for roughly one target-model step. If only the first token survives, the system behaves closer to normal decoding for that step.</p><p>The Markov head just helps the drafter keep later draft positions coherent. The confidence head helps the runtime decide how much speculation is worth attempting. Together, they attack the two practical bottlenecks of speculative decoding: low acceptance at later positions and wasted verification work under real serving load.</p><h3>Speed results</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!qSuE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!qSuE!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 424w, /__u/substackcdn.com/image/fetch/$s_!qSuE!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 848w, /__u/substackcdn.com/image/fetch/$s_!qSuE!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qSuE!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_webp, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!qSuE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png" width="1297" height="696" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:696,&quot;width&quot;:1297,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:302141,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/204695456?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!qSuE!, /__u/kaitchup.substack.com/w_424, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 424w, /__u/substackcdn.com/image/fetch/$s_!qSuE!, /__u/kaitchup.substack.com/w_848, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 848w, /__u/substackcdn.com/image/fetch/$s_!qSuE!, /__u/kaitchup.substack.com/w_1272, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qSuE!, /__u/kaitchup.substack.com/w_1456, /__u/kaitchup.substack.com/c_limit, /__u/kaitchup.substack.com/f_auto, /__u/kaitchup.substack.com/q_auto:good, /__u/kaitchup.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>DSpark improves per-user generation speed by about 60&#8211;85% over the previous MTP-1 baseline on DeepSeek-V4-Flash, and by about 57&#8211;78% on DeepSeek-V4-Pro at matched throughput. Offline accepted-length results also reportedly improve over Eagle3 and DFlash. As usual with speculative decoding, the exact gain depends on the prompt mix, sampling settings, hardware, batch pressure, and how often the drafter&#8217;s later tokens are accepted. Once widely available for other open-weight models, like Qwen3.6 and the larger Gemma 4, it will be interesting to see how fast is it compared with high MTP, like MTP-4 or 5 which are much faster than MTP-1 for these two models.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;35cea535-aba5-400d-96ae-db8ee9bc7d80&quot;,&quot;caption&quot;:&quot;Qwen3.6 inference is faster when MTP layers are enabled to draft tokens. Gemma 4 now supports this option as well. Both model families also have public DFlash speculator checkpoints, which can draft blocks of tokens in a single forward pass.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;DFlash vs MTP: Qwen3.6 Speculative Decoding Benchmarks with vLLM and llama.cpp&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-06-02T18:08:30.973Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!79OP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a729fb2-7c13-44ac-bad1-08602638a0f9_1530x742.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/dflash-vs-mtp-qwen36-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:198352019,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:8,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>You can find more results in the technical report:</p><p><a href="https://github.com/deepseek-ai/DeepSpec/blob/main/DSpark_paper.pdf">DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation</a></p><h3>What DeepSeek released</h3><p>The <a href="https://github.com/deepseek-ai/DeepSpec">DeepSpec</a> repository currently includes three draft-model algorithms: DSpark, DFlash, and Eagle3. </p><p>DeepSeek also released trained draft checkpoints for four target models: <a href="https://huggingface.co/deepseek-ai/dspark_qwen3_4b_block7">Qwen3-4B</a>, <a href="https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7">Qwen3-8B</a>, <a href="https://huggingface.co/deepseek-ai/dspark_qwen3_14b_block7">Qwen3-14B</a>, and <a href="https://huggingface.co/deepseek-ai/dspark_gemma4_12b_block7">Gemma-4-12B-it</a>. For each target, the repo lists Eagle3, DFlash, and DSpark checkpoints, so the open release covers twelve reference draft checkpoints.</p><p>DSpark checkpoints are also available for <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark">DeepSeek V4 Flash</a> and <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark">Pro</a>.</p><h3>Running DSpark in vLLM</h3><p>DSpark support is landing in vLLM through the speculative decoding path. The relevant vLLM PR for DSpark speculators-format checkpoint support was merged on July 2, 2026, so use a nightly build or a release that includes that change.</p><pre><code><code>uv pip install -U vllm --extra-index-url https://wheels.vllm.ai/nightly

vllm serve deepseek-ai/DeepSeek-V4-Flash-DSpark \
  --tensor-parallel-size 8 \
  --speculative-config '{"method":"dspark","num_speculative_tokens":5,"draft_sample_method":"probabilistic"}'</code></code></pre><h3>Training your own DSpark speculator</h3><p>You can train a DSpark speculator on your own data for Qwen3 and GLM 5.2 (only these architectures are support for now) using the <a href="https://github.com/vllm-project/speculators">speculators</a> project. This is very similar to training a DFlash a speculator.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;87ba04d9-9692-4a5f-86bb-a756a43c3148&quot;,&quot;caption&quot;:&quot;Speculative decoding is becoming a popular way to accelerate LLM inference, with approaches such as MTP and DFlash. The idea is simple: a smaller or specialized draft model proposes future tokens, and the full target model verifies them. Matching tokens are accepted and when a mismatch occurs, only the valid prefix is kept and decoding falls back to the target model.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Train and Run DFlash Speculative Decoding&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-18T19:43:12.334Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sbvO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b274a34-f749-4a9f-b68b-0e0f501a9016_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/train-and-run-dflash-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196847181,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/kaitchup.substack.com/subscribe"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="/__u/kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><p></p>]]></content:encoded></item></channel></rss>