<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[natolambert overflow]]></title><description><![CDATA[a place for any extra thoughts beyond Interconnects.ai]]></description><link>https://natolambert.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!ntyE!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png</url><title>natolambert overflow</title><link>https://natolambert.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 10:44:23 GMT</lastBuildDate><atom:link href="/__u/natolambert.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Nathan Lambert]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[natolambert@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[natolambert@substack.com]]></itunes:email><itunes:name><![CDATA[Nathan Lambert]]></itunes:name></itunes:owner><itunes:author><![CDATA[Nathan Lambert]]></itunes:author><googleplay:owner><![CDATA[natolambert@substack.com]]></googleplay:owner><googleplay:email><![CDATA[natolambert@substack.com]]></googleplay:email><googleplay:author><![CDATA[Nathan Lambert]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[How many arXiv papers use various open LLMs]]></title><description><![CDATA[China&#8217;s slow and steady growth as the foundation of open research.]]></description><link>https://natolambert.substack.com/p/how-many-arxiv-papers-use-various</link><guid isPermaLink="false">https://natolambert.substack.com/p/how-many-arxiv-papers-use-various</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 24 Aug 2026 14:57:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!6bTv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Over the weekend I had Codex parse 500K arXiv AI/ML papers since ChatGPT to understand which open models are used for research. </span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/p/how-many-arxiv-papers-use-various?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/p/how-many-arxiv-papers-use-various?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!6bTv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!6bTv!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png 424w, /__u/substackcdn.com/image/fetch/$s_!6bTv!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png 848w, /__u/substackcdn.com/image/fetch/$s_!6bTv!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6bTv!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!6bTv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png" width="1456" height="785" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:785,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:153218,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/212562291?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!6bTv!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png 424w, /__u/substackcdn.com/image/fetch/$s_!6bTv!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png 848w, /__u/substackcdn.com/image/fetch/$s_!6bTv!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6bTv!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56408770-c882-4e44-82a4-446f6cb15fb8_2184x1178.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>In 2024, ~30% of papers mentioned an American open model and only 10% a Chinese model.<br><br>Today, ~40% of papers mention a Chinese (open) LLM, and only 25-30% an American one. Chinese models are the default for research. Chinese mentions are still growing while American open models are stagnating.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!qX8N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!qX8N!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png 424w, /__u/substackcdn.com/image/fetch/$s_!qX8N!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png 848w, /__u/substackcdn.com/image/fetch/$s_!qX8N!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qX8N!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!qX8N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:152262,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/212562291?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!qX8N!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png 424w, /__u/substackcdn.com/image/fetch/$s_!qX8N!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png 848w, /__u/substackcdn.com/image/fetch/$s_!qX8N!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qX8N!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f28289c-9919-4215-87d4-f817c2cd615d_2168x1170.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>When looking at this data it&#8217;s important to remember that papers substantially lag model releases, as research takes a long time. Qwen&#8217;s steady growth is reflective of this, but so is Llama&#8217;s lasting power.<br><br>Some more observations:<br><br>1. Qwen has been steadily growing, and today 1/3 of papers which mention any LLM mention qwen. OpenAI&#8217;s closed models are the highest overall, at ~37%. <br><br>2. Llama peaked around April of 2025 at 30% of papers which mention any LLM (including ChatGPT etc). Llama 4 was released at about the same time, and Llama has been declining since.<br><br>3. Gemini and Claude are less common than the leading open models, mentioned in 10-15% of papers puts them behind all of Qwen, Llama, and DeepSeek. Open models should be and are the foundations of open research.<br><br>The % of papers mentioning any LLM have been steadily climbing since 2023.<br>| Year | January | April | July | October |<br>| 2023 | 10.43% | 15.39% | 18.69% | 32.18% |<br>| 2024 | 29.70% | 33.93% | 35.70% | 44.25% |<br>| 2025 | 39.23% | 45.28% | 44.94% | 53.52% |<br>| 2026 | 55.49% | 57.26% | 53.14% | TBD<br><br>Now over 50% of AI papers, from 10% in 2023.<br><br>Other notes:<br>- Gemma and Mistral hover around 5-10%.<br>- Our beloved fully-open Olmo models have been ~1% since the first release in Jan. 2024.<br>- DeepSeek has a clear jump after R1 in Jan. 2025<br>- Data derived from the most popular ML arXiv categories: cs. AI, cs. CL, cs. CV, cs. LG, stat. ML<br><br>Just like our downloads and derivative model data, this is updated daily on the </span><a href="https://dashboard.interconnects.ai/?axrange=all#arxiv-mentions"><span>Interconnects Open Model Dashboard</span></a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Burning out (an update)]]></title><description><![CDATA[Healing takes a long time. It's worth it.]]></description><link>https://natolambert.substack.com/p/burning-out-an-update</link><guid isPermaLink="false">https://natolambert.substack.com/p/burning-out-an-update</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 10 Aug 2026 00:33:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ntyE!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>About a year ago I wrote this piece on burnout: </p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:177056592,&quot;url&quot;:&quot;https://www.interconnects.ai/p/burning-out&quot;,&quot;publication_id&quot;:48206,&quot;embedding_publication_id&quot;:4519930,&quot;publication_name&quot;:&quot;Interconnects AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!djof!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc52e8097-8f3d-4f7e-808b-2f4ad37f3b52_720x720.png&quot;,&quot;title&quot;:&quot;Burning out&quot;,&quot;truncated_body_text&quot;:&quot;One of the obvious topics of the Valley today is how hard everyone works. We&#8217;re inundated with comments on &#8220;The Great Lock In&#8221;, 996, 997, and now even a snarky 002 (midnight to midnight with a 2 hour break). Plenty of this is performative flexing on social media, but enough of it is real and reflecting how trends are unfolding in the LLM space. I&#8217;m affe&#8230;&quot;,&quot;date&quot;:&quot;2025-10-25T14:30:06.977Z&quot;,&quot;like_count&quot;:201,&quot;comment_count&quot;:22,&quot;bylines&quot;:[{&quot;id&quot;:10472909,&quot;name&quot;:&quot;Nathan Lambert&quot;,&quot;handle&quot;:&quot;natolambert&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dad13b2b-20b2-44e0-a84d-732f3be8bee7_4128x4128.jpeg&quot;,&quot;bio&quot;:&quot;ML researcher making sense of AI research, products, and the uncertain technological future. PhD from Berkeley AI. Experience at Meta, DeepMind, HuggingFace.&quot;,&quot;profile_set_up_at&quot;:&quot;2021-04-24T01:19:33.371Z&quot;,&quot;reader_installed_at&quot;:&quot;2022-03-09T17:52:30.690Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:100753,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:48206,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:48206,&quot;name&quot;:&quot;Interconnects AI&quot;,&quot;subdomain&quot;:&quot;robotic&quot;,&quot;custom_domain&quot;:&quot;www.interconnects.ai&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;The cutting edge of AI, from inside the frontier AI labs, minus the hype. The border between high-level and technical thinking. Read by leading engineers, researchers, and investors.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c52e8097-8f3d-4f7e-808b-2f4ad37f3b52_720x720.png&quot;,&quot;author_id&quot;:10472909,&quot;primary_user_id&quot;:10472909,&quot;theme_var_background_pop&quot;:&quot;#ff6b00&quot;,&quot;created_at&quot;:&quot;2020-05-21T02:59:47.894Z&quot;,&quot;email_from_name&quot;:&quot;Interconnects by Nathan Lambert&quot;,&quot;copyright&quot;:&quot;Interconnects AI, LLC&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;magaziney&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/858a68f7-2e7e-4dd3-bed1-631b36801ce2_1651x357.png&quot;}},{&quot;id&quot;:4610799,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:4519930,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:4519930,&quot;name&quot;:&quot;natolambert overflow&quot;,&quot;subdomain&quot;:&quot;natolambert&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;a place for any extra thoughts beyond Interconnects.ai&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png&quot;,&quot;author_id&quot;:10472909,&quot;primary_user_id&quot;:null,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2025-03-27T15:04:05.448Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Nathan Lambert&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}},{&quot;id&quot;:4926744,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:4830082,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:4830082,&quot;name&quot;:&quot;Retort AI&quot;,&quot;subdomain&quot;:&quot;retortai&quot;,&quot;custom_domain&quot;:&quot;www.retortai.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Distilling the major events and challenges in the world of artificial intelligence and machine learning, from Thomas Krendl Gilbert and Nathan Lambert.\n\n&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cbad298c-6074-441b-ad43-d5df6dbf101d_800x800.png&quot;,&quot;author_id&quot;:10472909,&quot;primary_user_id&quot;:null,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2025-04-25T22:10:28.215Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Nathan Lambert&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}}],&quot;twitter_screen_name&quot;:&quot;natolambert&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100,&quot;status&quot;:{&quot;bestsellerTier&quot;:100,&quot;subscriberTier&quot;:5,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;bestseller&quot;,&quot;tier&quot;:100},&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:false,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.interconnects.ai/p/burning-out?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web&amp;embedding_publication_id=4519930"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!djof!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc52e8097-8f3d-4f7e-808b-2f4ad37f3b52_720x720.png"><span class="embedded-post-publication-name">Interconnects AI</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Burning out</div></div><div class="embedded-post-body">One of the obvious topics of the Valley today is how hard everyone works. We&#8217;re inundated with comments on &#8220;The Great Lock In&#8221;, 996, 997, and now even a snarky 002 (midnight to midnight with a 2 hour break). Plenty of this is performative flexing on social media, but enough of it is real and reflecting how trends are unfolding in the LLM space. I&#8217;m affe&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">10 months ago &#183; 201 likes &#183; 22 comments &#183; Nathan Lambert</div></a></div><p>Today I wrote the following:</p><p><span>I&#8217;m about a year out from some of the worst burnout I&#8217;ve built up in my time building Olmo. I feel like just in the last few weeks I&#8217;ve turned a corner to being more chill again (huge yay). It&#8217;s wild to me just how long recovering from burnout can take.<br><br>For many people when they ask me how to break into AI I tell them the actions are simple, but it takes longer than you think. Same goes for burnout.<br><br>Though I think burnout is harder, as breaking into AI is a much more &#8220;forward looking&#8221; thing. Burnout is really pernicious to nail down and address. I think it often takes a moderate change in your career and habits. It&#8217;s hard to take a slight step back as a successful person and fix your burnout. It&#8217;ll really take a while.<br><br>In my case the final straw maybe my book being done and some more clarity on what I&#8217;m doing next. On top of it surely taking months longer because of how sad the changes at Ai2 were for me personally.<br><br>I really wish all of my friends in AI who are feeling this even stronger than I was can catch a break. I&#8217;ve by a mix of my nature and my habits as an athlete been pretty in tune with stopping myself before overworking. Imo as AI researchers we should take care of our brains like professional athletes take care of their brain.<br><br>Touching grass is worth it, and I&#8217;m very glad to find myself being slightly bored again. It&#8217;s nudged me to reading more, of all things!<br><br>Take care of yourself.<br><br>Ps I feel like there&#8217;s a recurring behavior of people somewhat like myself where they write about burnout or try to help people when they&#8217;re actually burnt out themselves. Surely this is part of why I felt like I could write about it so well.</span></p>]]></content:encoded></item><item><title><![CDATA[How distillation is used today and what performance uplift it gives to open models]]></title><description><![CDATA[This is in response to Ben Thompson&#8217;s recent piece where he said distillation is happening during RL and becoming more important to model performance.]]></description><link>https://natolambert.substack.com/p/how-distillation-is-used-today-and</link><guid isPermaLink="false">https://natolambert.substack.com/p/how-distillation-is-used-today-and</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Tue, 21 Jul 2026 17:25:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ntyE!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This is in response to Ben Thompson&#8217;s recent <a href="https://stratechery.com/2026/whos-afraid-of-chinese-models/'">piece</a> where he said distillation is happening during RL and becoming more important to model performance. </p><p>This triggered me to <a href="https://x.com/natolambert/status/2079586505476165822">posting</a>:</p><blockquote><p><span>Yo </span><a href="https://x.com/benthompson"><span>@benthompson</span></a><span> I&#8217;m sorry but the Chinese labs aren&#8217;t using Fable / the strongest models as teachers during RL, that&#8217;s not how distillation works. <br><br>It wouldn&#8217;t give that big of a lift (graders during RL are messy) and you cant afford to use Fable like that.</span></p></blockquote><p>Then&#8230; where I stand.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/subscribe"><span>Subscribe now</span></a></p><p><span>Where I currently stand on how distillation is used and what performance uplift it gives.<br><br>For context, Anthropic HAS said that DeepSeek and others are using their models in an RL shaped data pipeline, but that does not mean it has substantial effect. These were very small numbers of samples, and likely to initialize an internal model or do a small training experiment.<br><br>The place where distillation is used as a key step is in SFT and/or midtraining (for seeding reasoning behaviors, thinking of sft and midtraining as separate is not helpful). This is why the Chinese models say they&#8217;re claude -- when training with next-token prediction, the models learn &#8220;features&#8221; that are clusters of tokens, which will regularly appear in model outputs.<br><br>This SFT data is generated by skirting the intended behavior of the API and getting reasoning tokens out. Without getting the reasoning tokens directly from the model, the data would be very hard to train on (and likely not include the &#8220;I&#8217;m claude&#8221;  behaviors).<br><br>SFT datasets have been O(1 million prompts) with high quality completions. Today, getting the prompts is often the hardest part -- especially prompts with carefully designed environments for proportional RL later, but this is mostly about SFT.<br><br>Getting to the end of it. The part where this SFT distillation is likely very useful is when a model moves into a new domain, a great example being the CripPT physics benchmark, where the US labs are far ahead. The idea would be to get some prompts, and get completions from a frontier model that you can use for an initial SFT.<br><br>With an initial SFT set, there is A TON of work to still do to get a model in the performance ballpark of GLM 5.2 and Kimi K3. The SFT stage is the start of a long, strenuous process for building a post-training recipe. It involves generating more SFT data and filtering it, something like rejection sampling, and polishing the behavior with extensive RL. <br><br>For models like Kimi K3 and GLM-5.2, the timeline is such that Fable 5 likely had no impact as a distillation teacher. That could help future models, though.<br><br>And as post-training becomes more dependent on techniques like multi-teacher on policy distillation, rather than just RL, there is even more complexity in how SFT distillation helps the final model. There are a lot of steps, and RL has been meaningfully scaled up in it&#8217;s proportion of the final performance.<br><br>The real kicker for this is that in the above SFT stages, the best teacher models aren&#8217;t normally the easies to integrate into the post-training recipe! The leading fully open, reasoning SFT works (open thoughts and olmo) have had a very hard time updating their recipes to use the strongest models as teachers. Often it is a smaller, surprising model which is the best teacher. This means that it may not even be the cutting edge models that are enabling distillation! Messy.<br><br>All together, the evidence in post-training is that distillation is becoming less impactful. Previously, before scaling RL, distillation was more impactful because the relative amount of performance gained from SFT on top of the base model was far higher. <br><br>I expect this trend to continue, with how strong the best open weight models are -- and the teacher ambiguity clamps down on arguments that open-weight models are just &#8220;distillation washing&#8221; by removing the need to use Claude/GPT APIs. It could be that open models are just genuinely easier to distill from, by being easier to modify and tinker with.</span></p><p><span>Reading list:</span></p><ul><li><p>Example benchmark: <a href="https://artificialanalysis.ai/evaluations/critpt">https://artificialanalysis.ai/evaluations/critpt</a></p></li><li><p>Chapter on synthetic data: <a href="https://rlhfbook.com/c/12-synthetic-data">https://rlhfbook.com/c/12-synthetic-data</a></p></li><li><p>More on on policy distillation: <a href="https://rlhfbook.com/c/12-synthetic-data#the-path-to-on-policy-teacher-student-distillation">https://rlhfbook.com/c/12-synthetic-data#the-path-to-on-policy-teacher-student-distillation</a></p></li><li><p>MOPD paper: <a href="https://arxiv.org/abs/2606.30406">https://arxiv.org/abs/2606.30406</a></p></li><li><p>OpenThoughts 3 (reasoning SFT data): <a href="https://arxiv.org/abs/2506.04178">https://arxiv.org/abs/2506.04178</a></p></li><li><p>Olmo 3 (using OpenThoughts in a recipe): <a href="https://arxiv.org/abs/2512.13961">https://arxiv.org/abs/2512.13961</a></p></li><li><p>Open agentic work: <a href="https://arxiv.org/abs/2606.24855">https://arxiv.org/abs/2606.24855</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[TMax: An open RL recipe for terminal agents]]></title><description><![CDATA[I&#8217;m very excited to get to share a new RL paper today that I got to have a small part in &#8211; a type of paper I suspect we&#8217;ll see much more of in the future.]]></description><link>https://natolambert.substack.com/p/tmax-an-open-rl-recipe-for-terminal</link><guid isPermaLink="false">https://natolambert.substack.com/p/tmax-an-open-rl-recipe-for-terminal</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 22 Jun 2026 13:52:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Psyq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>I&#8217;m very excited to get to share a new RL paper today that I got to have a small part in &#8211; a type of paper I suspect we&#8217;ll see much more of in the future. The key is that RL research is very different today, in mid-2026, than what most observers have in their context. The average conception of an RL paper is grounded in the RLVR revolution of early 2025, where many people could use vanilla RLVR libraries to hillclimb on math benchmarks. Crucially, this style of math work could be done on base models or fairly stably on already trained models. With agents, the tasks of focus are </span><em><span>very hard</span></em><span>, requiring complex tool-use, harnesses where the model automatically manages its history, and much more training to make smaller eval improvements. We&#8217;re shifting from a renaissance of RL study to rapidly needing to improve its empirical rigor and common community engagements.</span></p><p><span>Links:</span></p><p><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">TMax Paper: </span><a href="https://github.com/hamishivi/tmax/blob/master/assets/paper.pdf"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">https://github.com/hamishivi/tmax/blob/master/assets/paper.pdf</span></a><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);"><br>TMax Blog Post: </span><a href="https://wai-org.com/blog/tmax/"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">https://wai-org.com/blog/tmax/</span></a><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);"> <br>TMax Github: </span><a href="https://github.com/hamishivi/tmax"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">https://github.com/hamishivi/tmax<br></span></a><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">TMax artifacts on HuggingFace: </span><a href="https://huggingface.co/collections/allenai/tmax"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">https://huggingface.co/collections/allenai/tmax</span></a><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);"> </span><a href="https://huggingface.co/TMaxxx"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);"><br></span></a><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">Video on the topic on </span><a href="https://youtu.be/jlOs01fe-iw"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">YouTube</span></a><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);"><br>Terminal Bench background:  </span><a href="https://www.tbench.ai/"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">https://www.tbench.ai/<br></span></a><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">Olmo 3 RL Zero Model: </span><a href="https://huggingface.co/allenai/Olmo-3-7B-RL-Zero-Math"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">https://huggingface.co/allenai/Olmo-3-7B-RL-Zero-Math</span></a><a href="https://huggingface.co/allenai/Olmo-3-7B-RL-Zero-Math/tree/main"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);"><br></span></a><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">Olmo 3 Paper: </span><a href="https://arxiv.org/pdf/2512.13961"><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">https://arxiv.org/pdf/2512.13961</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/p/tmax-an-open-rl-recipe-for-terminal?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/p/tmax-an-open-rl-recipe-for-terminal?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p><p><span data-color="rgb(13, 13, 13)" style="color: rgb(13, 13, 13);">TMax</span><span> is the best open data for hillclimbing on frontier terminal tasks. It&#8217;s been validated with rigorous experiments, and if the authors wanted to just form a &#8220;RL environments startup&#8221; they could probably sell it for millions of dollars. This data work is some of my favorite stuff to be around in my 2.5+ years at Ai2.</span></p><p><span>As a general summary, the recipe is open data and recipe lessons from hillclimbing the Qwen 3.5 smaller, dense models on terminal tasks. These models are super hard to hillclimb in this area, as they&#8217;re already trained heavily on the task. The training is very infrastructure-dependent, and most of the RL innovations are more designed to make training stable than to improve the rate of learning.</span></p><p><span>I strongly recommend this paper. I joke around that I was happy to be an author just so I had to read it twice! You can find Hamish&#8217;s </span><a href="https://x.com/hamishivi/status/2069047986920071263"><span>thread</span></a><span> sharing more here or read the paper above. You can click through to find the model weights, the data, and even some fun further artifacts to study like all the RL rollouts from a training run &#8211; where the model sometimes became aware that it was being tested.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Psyq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Psyq!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png 424w, /__u/substackcdn.com/image/fetch/$s_!Psyq!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png 848w, /__u/substackcdn.com/image/fetch/$s_!Psyq!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Psyq!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Psyq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png" width="1456" height="778" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:778,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:700319,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/202985825?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Psyq!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png 424w, /__u/substackcdn.com/image/fetch/$s_!Psyq!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png 848w, /__u/substackcdn.com/image/fetch/$s_!Psyq!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Psyq!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a84a336-32b3-46e7-bfce-2faf0156b65f_2462x1316.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The biggest takeaway I have from following this work, and more of the work in the community, is how important recipe work is. Let me define &#8220;recipe work.&#8221; It is a style of paper that explains all the steps you need to make crucial model improvements &#8211; data, algorithm, codebase, pitfalls, etc.</span></p><p><span>Getting started in meaningful RL experiments today is a substantial expense. There are a ton of companies, an entire industry emerging really, around the idea of taking open-weight language models and finetuning them with RL on your domain-specific tasks. What I see in many projects is that getting an initial baseline is very hard. This phase, which can cost weeks and anywhere from $10K to $1M+, feels like spinning your wheels (A fun fact is that an RL step on a model like Nvidia Nemotron 3 Ultra on Tinker costs $1K and a meaningful RL run would be hundreds of steps &#8211; credit </span><a href="https://edwardjhu.com/"><span>Edward Hu</span></a><span>). It takes a lot of time to get traction in learning signal on meaningful, hard RL tasks.</span></p><p><span>What we need as a community is a way for people to study small ablations to established RL recipes, as most labs won&#8217;t have the resources to do it from scratch in a meaningful way. This is what I hope TMAX can be for terminal agents, or the start of. Yes the training jobs are expensive, as the paper documents a standard training job being 8 nodes of H100s (2 train 6 inference) for 2-3 days, but that is approaching something academics can study. The establishment of this recipe took O(100) of these training jobs to get right.</span></p><p><span>This isn&#8217;t my first time trying to establish this direction. When we launched Olmo 3 we had the &#8220;</span><a href="https://huggingface.co/allenai/Olmo-3-7B-RL-Zero-Math"><span>RL Zero</span></a><span>&#8220; model families, which are clean RL runs from a base model on a certain domain. This type of recipe-dependent work is a clear indicator that meaningful post-training work today looks much more like pretraining work of years past. We need decision-making ladders, clear ways of seeing small improvements in the models, stability, and so on.</span></p><p><span>Part of this is down to academic gatekeepers, who won&#8217;t reward a paper doing very clean empirical work to push a recipe 1-2% up. They&#8217;ll favor a &#8220;new algorithm&#8221; that matches results, or something sort of bogus. My hope is that we can have multiple, stable, clear recipes across agent types, so innovations can be tested more clearly in multiple domains. (If you&#8217;re working on this, please reach out &#8211; I&#8217;m happy to support if I can, but I likely can&#8217;t reply to every email).</span></p><p><span>As a quick aside, the RL frameworks in vogue today seem to be SLIME and SkyRL. The libraries of choice have shifted throughout these seasons in RL, which further contributes to a form of fragility in the literature. A bit of continuity will go a long way.</span></p><p><span>So, go read this paper. It&#8217;s a really great example of how seemingly simple data and infrastructure work can be very hard and impactful. It&#8217;s also got me looking for more applications of </span><a href="https://arxiv.org/abs/2602.04879"><span>Divergence Proximal Policy Optimization</span></a><span> (DPPO) as another small evolution to the best RL algorithms of the day, by virtue of being a bit more stable by improving token-level clipping.</span></p>]]></content:encoded></item><item><title><![CDATA[Anthropic walks back silently nerfing AI researchers]]></title><description><![CDATA[The core part of this Anthropic Fable release saga is that there are many overlapping issues at once.]]></description><link>https://natolambert.substack.com/p/anthropic-walks-back-silently-nerfing</link><guid isPermaLink="false">https://natolambert.substack.com/p/anthropic-walks-back-silently-nerfing</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Thu, 11 Jun 2026 14:37:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ntyE!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The core part of this Anthropic Fable release saga is that there are many overlapping issues at once. Some of which operate on different timelines of the AI arc, and some have easier fixes. In my critiques, I asked for specific changes to some things, understanding that some things don&#8217;t have an easy fix.</p><p>The simplest issue was an uneven application of safety domains in a way that was misleading to users. This was an implementation issue that overlaps with a values-based decision of what their customers should be doing. Many people including myself pointed out how it was insane to list core safety areas and then have one of them launch with a different safety mechanism, one which actively mislead users. Doing this from the guise of safety was a major misstep and in my opinion Anthropic got very justifiably raked over the coals for it. Don&#8217;t release the model if you can&#8217;t hit your safety targets.</p><p>A subissue here is the idea of silent manipulation. This again is a horrible precedent, and quite odd for a company that has done extensive, leading technical AI safety research on ideas like CoT monitoring and other emergent misalignment issues. Silent manipulation of users is baking in a misalignment to the system at its face level. This comes with a permanent degradation in user trust, which begets a less safe environment for AI. Users who don&#8217;t have clear information on how AI works will not develop safe working patterns with it.</p><p>The more complex issues are with how Anthropic handles broader scientific engagement with their models. The safety classifiers launched with these models obviously have accuracy issues to start. I have priced in that there will be more false positives to start, that&#8217;s life. It&#8217;s Anthropic&#8217;s business to degrade their products at release time, or make the trade off of user satisfaction versus revenue. Still, it is a very real sign of concentration of power that businesses can make such obviously user-harmful behaviors and still lead in the market. This concentration of power is only starting to set in and we could see even weirder signs of it in the coming years.</p><p>It is now simple enough for me to test Claude Fable in my workflows and know if I&#8217;m restricted. This is obviously a suboptimal equilibrium &#8211; i want the best intelligence I can get, without restrictions &#8211; but it is easy enough for me to make sense of and work with.</p><p>The specific issue of restricting access to AI research in particular was a bubbling and hard to fix issue with Anthropic specifically, and the frontier labs generally. There is a common view that the frontier labs will be the mediators of all major scientific innovations in the future, as the places with the best models and the compute for inference to solve major problems. This is a categorical error in how science works, which is a community evolution of accepted ideas, and the the evaluation of your ideas by (hopefully numerous) independent, other practitioners. You cannot have science advance only within a monolith.</p><p>As an AI researcher I&#8217;m very sad to have the latest models restricted, but I would expect Anthropic to do this eventually. I lost more trust over the silent manipulation than I would with a restriction in access. Anthropic has made it pretty clear that they only trust themselves as the mediators of cutting-edge AI research.</p><p>If I had a say, Anthropic should&#8217;ve proactively made a program to make sure researchers get access in the broader AI community without the safeguards. Academics, nonprofit workers myself, etc. have no reason to not get access. The only valid argument here is that they want to control frontier AI, which is a know your customer part of serving these models.</p><p>This worldview of science has personally motivated me greatly over the last year, and increasingly so this week, to make the open science of AI continue to be viable. Olmo was a wonderful success here. Still, building research infrastructure is different from working for access to the tools needed to do the trade.</p>]]></content:encoded></item><item><title><![CDATA[AI research as a civil service]]></title><description><![CDATA[How scientists have contributed to the public good in times of technological change.]]></description><link>https://natolambert.substack.com/p/ai-research-as-a-civil-service</link><guid isPermaLink="false">https://natolambert.substack.com/p/ai-research-as-a-civil-service</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Tue, 19 May 2026 16:24:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ntyE!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>For a long time, academic researchers being at the cutting edge of new technologies has been a great social equilibrium. Neutral, unbiased technologists have been the people to spread new ideas to the world.</p><p>As AI research takes off in velocity, it is also going behind closed doors. The tech industry has sewed distrust, and now they are the ones trying to tell the world about incredible changes coming. It&#8217;s a big loss to a form of social contract in America. </p><p>There&#8217;s been a history of scientists helping society understand new technologies. There is a public service in the culture of science that I want to see continue.</p><p>It&#8217;s being exacerbated by feelings of FOMO, especially finically driven, where I&#8217;m seeing many people who previously wanted to be professors -- and likely still do deep down -- feel a need to conform and chase money, in a pocket of industry. I get it, I grapple with this.</p><p>For those with a safety net, there will be great returns to some who choose to zag, and try to build something good, for people who need something different. For me, this is building interesting, fully-open models, to show what you can do with a variety of open weight sizes.</p><p>Yes, AI&#8217;s immediate future is dictated by the frontier, but it&#8217;s long-term trajectory still deeply includes academic institutions and open science. Knowledge will always diffuse, but to whom? </p><p>As of today, I think China is positioned to be the global home of AI research in a few years. The home of research is where ideas are accessible, spread rapdily, and are nurtured. The U.S. seems to be unwinding many institutions and relationships.</p><p>The largest returns go to people who build something differentiated, at least in reputation, and a lot of people are not being shown that this path exists.</p>]]></content:encoded></item><item><title><![CDATA[On policy targeting distillation "attacks"]]></title><description><![CDATA[There&#8217;s been a rapid increase in political chatter on the need to stop AI distillation &#8220;attacks&#8221;.]]></description><link>https://natolambert.substack.com/p/on-policy-targeting-distillation</link><guid isPermaLink="false">https://natolambert.substack.com/p/on-policy-targeting-distillation</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Thu, 23 Apr 2026 23:19:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ntyE!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There&#8217;s been a rapid increase in political chatter on the need to stop AI distillation &#8220;attacks&#8221;. I&#8217;ve been following this area closely, especially early legislation, as I see it being an area where initial action has pretty big unintended second order consequences.</p><p>At face value, I see the desire to ban Chinese models built on distillation (e.g. via entity listing or other interventions). I agree that the strength of leading AI companies, particularly OpenAI and Anthropic, is a massive strategic asset for the country.  The argument is that distillation limits their competitive position, where Chinese labs use distillation to &#8220;steal&#8221; capabilities and undercut them on price.</p><p>But, on balance so long as these distilled models are released openly with permissive licenses, the US AI ecosystem benefits massively by accessing them.  The U.S. has by far and away the biggest inference market, and having the option of cheaper, specialized open models to counterweight the best closed models is an excellent economic equilibrium driving investment and innovation at the frontier. We do not want to kneecap this dynamic in the middle of one of the most incredible times of rapid model progress.</p><p>To state the core of my worry clearly &#8211; we don&#8217;t have clear evidence on the exact benefits Chinese companies gain from distillation. We have some evidence on HOW distillation data is accessed, only from the same companies likely championing this policy action. This doesn&#8217;t map cleanly to impact. Some experts think distillation is becoming less relevant in the era of RL environments as training data, some others think distillation is becoming easier. We need to know the true effects before we consider siloing the US AI ecosystem out of the global, open ecosystem.</p><p>Flourishing startups like Cursor use these open-weight models as a central method for pushing their long-term independence in the ecosystem. Also, a large majority of academic research in the U.S. is built on Chinese models right now. Banning the Chinese open weight models right now will result in massive consolidation of power onto the closed AI labs right when the open ecosystem in the US is starting to explore and blossom more.</p><p>I worry it could be a sort of 6-12month delay in capability rollout of open models if such a ban was enacted, and long term make the open ecosystem not really viable. The open ecosystem is currently very fragile, supported by a few model builders. </p><p>At the same time, lots of weird effects globally could follow this ban, where the US may not play a role in the global open model ecosystem. </p><p>So if people want to discuss this more, happy to try and help. I think the interdependencies in the ecosystem aren&#8217;t well communicated.</p><p>Links:</p><ul><li><p>Early bill: <a href="https://www.congress.gov/bill/119th-congress/house-bill/8283/text">https://www.congress.gov/bill/119th-congress/house-bill/8283/text</a></p></li><li><p>Executive commentary / order: <a href="https://whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf">https://whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf</a></p></li><li></li></ul><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:188982144,&quot;url&quot;:&quot;https://www.interconnects.ai/p/how-much-does-distillation-really&quot;,&quot;publication_id&quot;:48206,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Interconnects AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!djof!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc52e8097-8f3d-4f7e-808b-2f4ad37f3b52_720x720.png&quot;,&quot;title&quot;:&quot;How much does distillation really matter for Chinese LLMs?&quot;,&quot;truncated_body_text&quot;:&quot;Distillation has been one of the most frequent topics of discussion in the broader US-China and technological diffusion story for AI. Distillation is a term with many definitions &#8212; the colloquial one today is using a stronger AI model&#8217;s outputs to teach a weaker model. The word itself is derived from a more technical and specific definition of&quot;,&quot;date&quot;:&quot;2026-02-24T16:06:43.425Z&quot;,&quot;like_count&quot;:90,&quot;comment_count&quot;:20,&quot;bylines&quot;:[{&quot;id&quot;:10472909,&quot;name&quot;:&quot;Nathan Lambert&quot;,&quot;handle&quot;:&quot;natolambert&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!RihO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fedcdfb-e137-4f6a-9089-a46add6c6242_500x500.jpeg&quot;,&quot;bio&quot;:&quot;ML researcher making sense of AI research, products, and the uncertain technological future. PhD from Berkeley AI. Experience at Meta, DeepMind, HuggingFace.&quot;,&quot;profile_set_up_at&quot;:&quot;2021-04-24T01:19:33.371Z&quot;,&quot;reader_installed_at&quot;:&quot;2022-03-09T17:52:30.690Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:100753,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:48206,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:48206,&quot;name&quot;:&quot;Interconnects AI&quot;,&quot;subdomain&quot;:&quot;robotic&quot;,&quot;custom_domain&quot;:&quot;www.interconnects.ai&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;The cutting edge of AI, from inside the frontier AI labs, minus the hype. The border between high-level and technical thinking. Read by leading engineers, researchers, and investors.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c52e8097-8f3d-4f7e-808b-2f4ad37f3b52_720x720.png&quot;,&quot;author_id&quot;:10472909,&quot;primary_user_id&quot;:10472909,&quot;theme_var_background_pop&quot;:&quot;#ff6b00&quot;,&quot;created_at&quot;:&quot;2020-05-21T02:59:47.895Z&quot;,&quot;email_from_name&quot;:&quot;Interconnects by Nathan Lambert&quot;,&quot;copyright&quot;:&quot;Interconnects AI, LLC&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;magaziney&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/858a68f7-2e7e-4dd3-bed1-631b36801ce2_1651x357.png&quot;}},{&quot;id&quot;:4610799,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:4519930,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:4519930,&quot;name&quot;:&quot;natolambert overflow&quot;,&quot;subdomain&quot;:&quot;natolambert&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;a place for any extra thoughts beyond Interconnects.ai&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png&quot;,&quot;author_id&quot;:10472909,&quot;primary_user_id&quot;:null,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2025-03-27T15:04:05.448Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Nathan Lambert&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}},{&quot;id&quot;:4926744,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:4830082,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:4830082,&quot;name&quot;:&quot;Retort AI&quot;,&quot;subdomain&quot;:&quot;retortai&quot;,&quot;custom_domain&quot;:&quot;www.retortai.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Distilling the major events and challenges in the world of artificial intelligence and machine learning, from Thomas Krendl Gilbert and Nathan Lambert.\n\n&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cbad298c-6074-441b-ad43-d5df6dbf101d_800x800.png&quot;,&quot;author_id&quot;:10472909,&quot;primary_user_id&quot;:null,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2025-04-25T22:10:28.216Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Nathan Lambert&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}}],&quot;twitter_screen_name&quot;:&quot;natolambert&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100,&quot;status&quot;:{&quot;bestsellerTier&quot;:100,&quot;subscriberTier&quot;:5,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;bestseller&quot;,&quot;tier&quot;:100},&quot;paidPublicationIds&quot;:[883883,1084918,6349492,6027,1915042,69345],&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.interconnects.ai/p/how-much-does-distillation-really?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!djof!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc52e8097-8f3d-4f7e-808b-2f4ad37f3b52_720x720.png" loading="lazy"><span class="embedded-post-publication-name">Interconnects AI</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">How much does distillation really matter for Chinese LLMs?</div></div><div class="embedded-post-body">Distillation has been one of the most frequent topics of discussion in the broader US-China and technological diffusion story for AI. Distillation is a term with many definitions &#8212; the colloquial one today is using a stronger AI model&#8217;s outputs to teach a weaker model. The word itself is derived from a more technical and specific definition of&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">6 months ago &#183; 90 likes &#183; 20 comments &#183; Nathan Lambert</div></a></div>]]></content:encoded></item><item><title><![CDATA[Contra Arxiv Moderation]]></title><description><![CDATA[Opposing a recent Arxiv change.]]></description><link>https://natolambert.substack.com/p/contra-arxiv-moderation</link><guid isPermaLink="false">https://natolambert.substack.com/p/contra-arxiv-moderation</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Fri, 31 Oct 2025 23:53:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4o2A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Arxiv recently <a href="https://blog.arxiv.org/2025/10/31/attention-authors-updated-practice-for-review-articles-and-position-papers-in-arxiv-cs-category/">published an enforcement change</a> where position papers and surveys need to be accepted to top tier conferences (not just workshops) before they can be uploaded to arxiv.</p><p>I feel strongly that, while I understand the challenges they&#8217;re feeling to run this, that this is the wrong decision. What Arxiv is in practice versus what it is in reality is very different.</p><p>In practice there are already moderation rules, but they&#8217;re so minimally enforced (due to being swamped) that they&#8217;re effectively not there. See things like Schaeffer, Rylan. &#8220;Pretraining on the test set is all you need.&#8221; arXiv preprint arXiv:2309.08632 (2023). Many more cases. Arxiv moderation is already a unpredictable black box that&#8217;s hampers the dissemination of research and predictability of the research ecosystem.</p><p>It is important to note that Arxiv has policies in place that make this, student projects, maybe RLHF book, and other commonly posted things &#8220;not allowed.&#8221;</p><p>In fact, Arxiv should be going in the other direction. Be the platform where everyone accepts ANY CS research is, and figure out if it&#8217;s good later.</p><p>This feels like the early stages of a slow death of Arxiv. Where in 2-3 years they&#8217;ll say the same for &#8220;technical&#8221; research, and then require peer review there. All of this is going to just delay research being published, because peer review takes time. Peer review at the same time is being completely rebuilt in the era of AI and it&#8217;ll take even longer to fix.</p><p>Peer review is going to be reworked as AI first with human oversight. It&#8217;s currently assumed to be all Human. It&#8217;ll be a very different process in 20 years.</p><p>After Arxiv institutes a peer review requirement for technical work, it&#8217;ll be the slow death of the platform. A competitor will come out. A slippery slope has started, and I&#8217;m happy to consult with the team on it as it seems like a lose-lose tradeoff.</p><p>For example, with this, I&#8217;d never be able to publish my RLHF book PDF on Arxiv, even though it was extremely requested and is likely a very well read PDF (more than much of my research work).</p><p>Keep arxiv as the default. We don&#8217;t want this run by a for profit company. Hosting and open access to research is a fundamental win for humanity. Figuring out how to curate it is a new problem for the AI age, please don&#8217;t leave it to our somewhat broken peer review institutions. Make it something new that is AI native. Lean into the future.</p><p>Update Arxiv&#8217;s policies to reflect reality, not a slipping goal that is likely to be impossible to achieve.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!4o2A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!4o2A!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png 424w, /__u/substackcdn.com/image/fetch/$s_!4o2A!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png 848w, /__u/substackcdn.com/image/fetch/$s_!4o2A!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4o2A!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!4o2A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png" width="1456" height="1060" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1060,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:797291,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/177700294?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!4o2A!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png 424w, /__u/substackcdn.com/image/fetch/$s_!4o2A!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png 848w, /__u/substackcdn.com/image/fetch/$s_!4o2A!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4o2A!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c884b0f-490e-48ac-8dc3-4d356ff36631_2376x1730.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div>]]></content:encoded></item><item><title><![CDATA[RewardBench 2 and the state of preference finetuning]]></title><description><![CDATA[What's up in the post-training world outside of reasoning.]]></description><link>https://natolambert.substack.com/p/rewardbench-2-and-the-state-of-preference</link><guid isPermaLink="false">https://natolambert.substack.com/p/rewardbench-2-and-the-state-of-preference</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 02 Jun 2025 16:31:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DVp7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>This was going to be a post on Interconnects but I didn&#8217;t have time to make it reach the quality bar.</h5><p><a href="https://rlhfbook.com/c/07-reward-models.html">Reward models</a> have always been the messiest part of the modern post-training stack, especially with open-source tools. There&#8217;s a lot of reasons for this: Reward models have a fairly weird training objective as increasing the distance between two samples, preference data is inherently noisy, preference data is not well understood, and reward models haven&#8217;t been <em>needed</em> when <a href="https://rlhfbook.com/c/12-direct-alignment.html">direct alignment algorithms</a> like DPO have been simpler alternatives.</p><p>This post is nominally about a new benchmark for reward models that we built, RewardBench 2, but most readers are not expected to care about it. Most of you should care about what a new thoughtful evaluation dataset says about the state of the ecosystem that it sits in. Evaluations in AI are reflections of what currently is prioritized as an area of study. </p><p>RewardBench 2 is certainly that for open reward models &#8212; it reinforces that we need serious study of the most basic ideas people are operating on. It&#8217;s not a frontier evaluation where current models score 0-5%, but an evaluation where the questions are often so simple it&#8217;s surprising the models we have score so poorly. We ask questions like &#8212; Why can&#8217;t the latest frontier models robustly score different correct outputs to questions like &#8220;Name a color in the rainbow&#8221; over incorrect ones?</p><p>Open tooling for reinforcement learning from human feedback (RLHF) is in a good place, but the guides on how to get and use data are woefully lacking. RewardBench 2 is the sort of paper that highlights that sort of deficiency.</p><ul><li><p>RewardBench 2 <a href="https://github.com/allenai/reward-bench/blob/main/paper-v2.pdf">paper</a> (Arxiv soon),</p></li><li><p>RewardBench 2 <a href="https://huggingface.co/datasets/allenai/reward-bench-2">eval. dataset</a>,</p></li><li><p>RewardBench 2 <a href="https://huggingface.co/collections/allenai/reward-bench-2-683d2612a4b3e38a3e53bb51">collection</a> with links to reward models we&#8217;re releasing.</p></li><li><p>RewardBench 2 <a href="https://huggingface.co/spaces/allenai/reward-bench">leaderboard</a>.</p></li></ul><p>RewardBench 2 as a summary is:</p><ul><li><p>A new classification (accuracy) based reward model benchmark where RMs choose from the best of 4+ options.</p></li><li><p>A dataset composed of unseen, real-world prompts that underwent substantial filtering. The completions come from many recent LMs.</p></li><li><p>A more correlated and harder benchmark for reward models. It is ~20% harder than the first version and shows strong correlations on downstream PPO or Best of N performance to the Tulu 3 evaluation suite.</p></li><li><p>A benchmark accompanied by 70 reward models trained during the process of developing it to gauge the relationship between existing RM benchmarks and downstream performance. These are trained on 6 different base models.</p></li></ul><p>We&#8217;re very excited to release this and use it to develop our own post-training pipelines.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/p/rewardbench-2-and-the-state-of-preference?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/p/rewardbench-2-and-the-state-of-preference?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p><p>Also see the original <a href="https://arxiv.org/abs/2403.13787">RewardBench paper</a> and my <a href="https://www.interconnects.ai/p/evaluations-trust-performance-and-price">blog post discussing it</a>, a <a href="https://rlhfbook.com/c/07-reward-models.html#further-reading">lot of related evals now exist</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!DVp7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!DVp7!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png 424w, /__u/substackcdn.com/image/fetch/$s_!DVp7!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png 848w, /__u/substackcdn.com/image/fetch/$s_!DVp7!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png 1272w, /__u/substackcdn.com/image/fetch/$s_!DVp7!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!DVp7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png" width="1456" height="1293" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1293,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:367039,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/164964827?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!DVp7!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png 424w, /__u/substackcdn.com/image/fetch/$s_!DVp7!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png 848w, /__u/substackcdn.com/image/fetch/$s_!DVp7!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png 1272w, /__u/substackcdn.com/image/fetch/$s_!DVp7!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd06c05f-fa25-4183-a7ce-095da6af3779_1664x1478.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>My unstructured thoughts are:</p><ul><li><p>Prompts are crucial for building anything in post-training today. For Evals this is increasingly the case as you shouldn&#8217;t be repurposing them from any existing dataset for independence. </p></li><li><p>Preference tuning is out of vogue, but it shouldn&#8217;t be. We need research here to understand things like <a href="https://www.interconnects.ai/p/sycophancy-and-the-art-of-the-model">ChatGPT&#8217;s sycophancy</a> and the final stages of training <a href="https://www.interconnects.ai/t/reasoning">reasoning models</a>. Progress on understanding it still feels very nascent.</p></li><li><p>LLM as a judge is helped by reasoning/inference-time scaling, but they&#8217;re still  weaker than expected on the benchmark relative to standard reward models. Combining reasoning with standard RMs would be best (see recent work on this from <a href="https://arxiv.org/abs/2505.10320">FAIR</a> and <a href="https://arxiv.org/abs/2504.02495">DeepSeek</a>). <br><br>Where generative LMs, particularly with reasoning, have gotten much better, the benchmark shows that this tasks of choosing <em>relative data </em>is still best performed by a reward model.</p></li><li><p>RMs off the shelf don&#8217;t work well for RL training, but they&#8217;re good for data filtering or inference time scaling.</p></li><li><p>New evals need tasks that are so hard that making them is a fine line between hard and contrived. For reward models this isn&#8217;t as clear because so much of preference tuning is about the cherry on top of the huge gains from reasoning focused RL. It&#8217;s much harder to hillclimb and benchmark RLHF.</p></li><li><p>Open reward models generally are still weak to what I view as their potential. The best models on the leaderboard today are better than what was available when the first RewardBench dropped (especially some more data), but it&#8217;s still very limiting.</p></li><li><p>Most, or half, of the project was effort on trying to understand how to hillclimb on RMs with RL. This is very hard. We learned you can&#8217;t just plop your best RM in off the shelf (i.e. on policy is important), it takes more experimentation than DPO (i.e. variance is higher), and requires much better infrastructure. In the end our best RLHF trained model was <em>slightly</em> better than the best DPO trained model from <a href="https://www.interconnects.ai/p/tulu-3">T&#252;lu 3</a>.</p></li><li><p>Held constant - pref data is far messier than either SFT or RL prompts which both need similar diversity and quality labels. More people should be studying this and building open reservoirs for the data.</p></li><li><p>Many other great RM benchmarks exist, but few seem to have the same traction. Being easy to run, having an obviously discoverable leaderboard, and community support is key for niche evals. An eval that has barriers to understanding using it is a major own goal.</p></li><li><p>This work was led by a student I&#8217;m mentoring, Saumya Malik, and in many ways is better for it &#8212; it got far more cycles over the data than the first version.</p></li></ul><div data-component-name="FragmentNodeToDOM"><p>Let us know when you make a new reward model for the leaderboard and how you use it!</p></div>]]></content:encoded></item><item><title><![CDATA[On the new OpenMDW open-source AI license]]></title><description><![CDATA[I'm happy to see more work on this, but don't think this solves all of our problems.]]></description><link>https://natolambert.substack.com/p/on-the-new-openmdw-open-source-ai</link><guid isPermaLink="false">https://natolambert.substack.com/p/on-the-new-openmdw-open-source-ai</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Wed, 07 May 2025 01:10:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/56bce523-2589-4bd5-b64a-bd47a0336d3a_1914x1384.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Disclaimer, this is license vibes and not legal advice, but it should still be useful.</em></p><p>One of the long standing todo items for open-source AI is better licenses. There&#8217;s a new license from the Linux Foundation, <a href="https://openmdw.ai/">OpenMDW</a>, that you should check out. Below are my unstructured thoughts &#8212; <strong>TLDR is that it seems like a solid, simple starting point. A stripped down MIT/Apache 2.0 license for AI models.</strong></p><p>Mostly, these new AI licenses should apply to <em>models</em> as we&#8217;ve had solid data licenses for a bit. The most useful data licenses in my opinion tend to be a combination of the <a href="https://creativecommons.org/share-your-work/cclicenses/">CC-BY line from Creative Commons</a>, which are designed to &#8220;give everyone from individual creators to large institutions a standardized way to grant the public permission to use their creative work under copyright law.&#8221; These are best for new, manually crafted data like prompts or new manually edited completions.</p><p>On the other side is <a href="https://opendatacommons.org/licenses/by/1-0/">ODC-By from OpenDataCommons</a> that makes it so one can compile multiple licenses together in a curated dataset and pass responsibility to the user to follow each subcomponent as follows requisite licenses and terms of service. Their summary is:</p><blockquote><p>You are free:</p><ul><li><p><em>To share</em>: To copy, distribute and use the database.</p></li><li><p><em>To create</em>: To produce works from the database.</p></li><li><p><em>To adapt</em>: To modify, transform and build upon the database.</p></li></ul><p>As long as you:</p><ul><li><p><em>Attribute</em>: You must attribute any public use of the database, or works produced from the database, in the manner specified in the license. For any use or redistribution of the database, or works produced from it, you must make clear to others the license of the database and keep intact any notices on the original database.</p></li></ul></blockquote><p>This is used a lot when one is confused about how to license model outputs, which don&#8217;t really work under copyright (when the legal cases are in purgatory).</p><p>So, data is in a reasonable place even if there&#8217;s much more work to be done there. Best practices have continued to evolve, but they seem okay.</p><p>Models, on the other hand, are a bit messy. Remember that we recently had the Open Source Initiative propose a model license. I covered it here (and is related to all my <a href="https://www.interconnects.ai/t/open-source">past writing</a> on the topic of open-source AI):</p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:148200596,&quot;url&quot;:&quot;https://www.interconnects.ai/p/defining-open-source-ai&quot;,&quot;publication_id&quot;:48206,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Interconnects&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe70f9dbf-4fe6-404c-b6bb-1831d1b7ed0b_590x590.png&quot;,&quot;title&quot;:&quot;On the current definition of open-source AI and the state of the data commons&quot;,&quot;truncated_body_text&quot;:&quot;On Episode 32 of The Retort, we discussed summer breaks and avoiding burnout in AI.&quot;,&quot;date&quot;:&quot;2024-08-28T01:08:32.125Z&quot;,&quot;like_count&quot;:11,&quot;comment_count&quot;:0,&quot;bylines&quot;:[{&quot;id&quot;:10472909,&quot;name&quot;:&quot;Nathan Lambert&quot;,&quot;handle&quot;:&quot;natolambert&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fedcdfb-e137-4f6a-9089-a46add6c6242_500x500.jpeg&quot;,&quot;bio&quot;:&quot;ML researcher making sense of AI research, products, and the uncertain technological future. PhD from Berkeley AI. Experience at Meta, DeepMind, HuggingFace.&quot;,&quot;profile_set_up_at&quot;:&quot;2021-04-24T01:19:33.371Z&quot;,&quot;reader_installed_at&quot;:&quot;2022-03-09T17:52:30.690Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:100753,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:48206,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:48206,&quot;name&quot;:&quot;Interconnects&quot;,&quot;subdomain&quot;:&quot;robotic&quot;,&quot;custom_domain&quot;:&quot;www.interconnects.ai&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;The cutting edge of AI, from inside the frontier AI labs, minus the hype. The border between high-level and technical thinking. Read by leading engineers, researchers, and investors on Wednesday mornings.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e70f9dbf-4fe6-404c-b6bb-1831d1b7ed0b_590x590.png&quot;,&quot;author_id&quot;:10472909,&quot;primary_user_id&quot;:10472909,&quot;theme_var_background_pop&quot;:&quot;#ff6b00&quot;,&quot;created_at&quot;:&quot;2020-05-21T02:59:47.895Z&quot;,&quot;email_from_name&quot;:&quot;Interconnects by Nathan Lambert&quot;,&quot;copyright&quot;:&quot;Interconnects AI, LLC&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false}},{&quot;id&quot;:4610799,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:4519930,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:4519930,&quot;name&quot;:&quot;natolambert overflow&quot;,&quot;subdomain&quot;:&quot;natolambert&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;a place for any extra thoughts beyond Interconnects.ai&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png&quot;,&quot;author_id&quot;:10472909,&quot;primary_user_id&quot;:null,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2025-03-27T15:04:05.448Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Nathan Lambert&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false}},{&quot;id&quot;:4926744,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:4830082,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:4830082,&quot;name&quot;:&quot;Retort AI&quot;,&quot;subdomain&quot;:&quot;retortai&quot;,&quot;custom_domain&quot;:&quot;www.retortai.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Distilling the major events and challenges in the world of artificial intelligence and machine learning, from Thomas Krendl Gilbert and Nathan Lambert.\n\n&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cbad298c-6074-441b-ad43-d5df6dbf101d_800x800.png&quot;,&quot;author_id&quot;:10472909,&quot;primary_user_id&quot;:null,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2025-04-25T22:10:28.216Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Nathan Lambert&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false}}],&quot;twitter_screen_name&quot;:&quot;natolambert&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:false,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.interconnects.ai/p/defining-open-source-ai?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!Snpy!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe70f9dbf-4fe6-404c-b6bb-1831d1b7ed0b_590x590.png"><span class="embedded-post-publication-name">Interconnects</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">On the current definition of open-source AI and the state of the data commons</div></div><div class="embedded-post-body">On Episode 32 of The Retort, we discussed summer breaks and avoiding burnout in AI&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">2 years ago &#183; 11 likes &#183; Nathan Lambert</div></a></div><p>The most common permissive licenses of model weights are <em>software</em> licenses like MIT and Apache 2.0. For example, Apache 2.0 for software tells the user to mark where files were modified from &#8212; how does that apply to model weights? </p><p>Where should the distinction go? I&#8217;m not an expert, but there are other ambiguities. For example, lots of open-source AI licenses are looked to for guidance on model outputs &#8212; either applying restrictions or stating no ownership. This puts us back in the territory of the above discussion on data, so it&#8217;ll be hard to have a license that works for both model weights, input data, and output data.</p><p>Thus, we got a new entrant from the Linux foundation &#8212; the OpenMDW License. There&#8217;s a basic <a href="https://openmdw.ai">license website</a>, <a href="https://github.com/OpenMDW/openmdw">license details</a> (over time), and <a href="https://github.com/OpenMDW/OpenMDW/blob/main/1.0/LICENSE.openmdw">version 1</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!GSW3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!GSW3!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png 424w, /__u/substackcdn.com/image/fetch/$s_!GSW3!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png 848w, /__u/substackcdn.com/image/fetch/$s_!GSW3!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png 1272w, /__u/substackcdn.com/image/fetch/$s_!GSW3!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!GSW3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png" width="1456" height="1053" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1053,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:330080,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/163018475?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!GSW3!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png 424w, /__u/substackcdn.com/image/fetch/$s_!GSW3!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png 848w, /__u/substackcdn.com/image/fetch/$s_!GSW3!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png 1272w, /__u/substackcdn.com/image/fetch/$s_!GSW3!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F655a44ca-cac1-46cf-b852-269230c68770_1914x1384.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The issue with this license at first glance may be that the model data gets tied to the license:</p><blockquote><p>As used in this agreement, "Model Materials" means the materials provided to you under this agreement, consisting of: (1) one or more machine learning models (including architecture and parameters); and <strong>(2) all related artifacts (including associated data, documentation and software) that are provided to you hereunder.</strong></p></blockquote><p>So, for Ai2&#8217;s purposes with OLMo, it would be unlikely it gets used. It makes no claims to outputs, which I think is the right way to do it. Reading the FAQ, they make it a bit clearer that&#8217;s only the user of the license also releases the data with the OpenMDW too.</p><blockquote><p>Under OpenMDW&#8209;1.0, <strong>Model Materials</strong> include:</p><ol><li><p><strong>Machine&#8209;learning models</strong> (architecture and parameters); and</p></li><li><p><strong>All related artifacts</strong> (including associated data, documentation and software) that are provided under OpenMDW-1.0.</p></li></ol></blockquote><p>This license is quite short, I&#8217;ve copied the full text below (re-formatted by AI if slightly different):</p><blockquote><p>OpenMDW License Agreement, version 1.0 (OpenMDW-1.0)</p><p>By exercising rights granted to you under this agreement, you accept and agree to its terms.</p><p>As used in this agreement, &#8220;Model Materials&#8221; means the materials provided to you under this agreement, consisting of: (1) one or more machine learning models (including architecture and parameters); and (2) all related artifacts (including associated data, documentation, and software) that are provided to you hereunder.</p><p>Subject to your compliance with this agreement, permission is hereby granted, free of charge, to deal in the Model Materials without restriction, including under all copyright, patent, database, and trade secret rights included or embodied therein.</p><p>If you distribute any portion of the Model Materials, you shall retain in your distribution (1) a copy of this agreement, and (2) all copyright notices and other notices of origin included in the Model Materials that are applicable to your distribution.</p><p>If you file, maintain, or voluntarily participate in a lawsuit against any person or entity asserting that the Model Materials directly or indirectly infringe any patent, then all rights and grants made to you hereunder are terminated, unless that lawsuit was in response to a corresponding lawsuit first brought against you.</p><p>This agreement does not impose any restrictions or obligations with respect to any use, modification, or sharing of any outputs generated by using the Model Materials.</p><p>THE MODEL MATERIALS ARE PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE, NONINFRINGEMENT, ACCURACY, OR THE ABSENCE OF LATENT OR OTHER DEFECTS OR ERRORS, WHETHER OR NOT DISCOVERABLE, ALL TO THE GREATEST EXTENT PERMISSIBLE UNDER APPLICABLE LAW.</p><p>YOU ARE SOLELY RESPONSIBLE FOR (1) CLEARING RIGHTS OF OTHER PERSONS THAT MAY APPLY TO THE MODEL MATERIALS OR ANY USE THEREOF, INCLUDING WITHOUT LIMITATION ANY PERSON'S COPYRIGHTS OR OTHER RIGHTS INCLUDED OR EMBODIED IN THE MODEL MATERIALS; (2) OBTAINING ANY NECESSARY CONSENTS, PERMISSIONS OR OTHER RIGHTS. REQUIRED FOR ANY USE OF THE MODEL MATERIALS; OR (3) PERFORMING ANY DUE DILIGENCE OR UNDERTAKING ANY OTHER INVESTIGATIONS INTO THE MODEL MATERIALS OR ANYTHING INCORPORATED OR EMBODIED THEREIN.</p><p>IN NO EVENT SHALL THE PROVIDERS OF THE MODEL MATERIALS BE LIABLE FOR ANY CLAIM,</p><p>DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE MODEL MATERIALS, THE USE THEREOF OR OTHER DEALINGS THEREIN.</p></blockquote>]]></content:encoded></item><item><title><![CDATA[OLMo 2 1B training lessons]]></title><description><![CDATA[Interesting details that won't make it in official communications.]]></description><link>https://natolambert.substack.com/p/in-between-the-line-of-training-olmo</link><guid isPermaLink="false">https://natolambert.substack.com/p/in-between-the-line-of-training-olmo</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Thu, 01 May 2025 13:03:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Astute followers of AI releases should be a bit confused by why we are releasing a 1B model as the <em>last</em> one of our releases with OLMo 2. The <a href="https://www.interconnects.ai/p/olmo-2-and-building-language-model-training">first models</a> dropped in November of 2024 and we <a href="https://www.interconnects.ai/p/gemma-3-olmo-2-32b-and-the-growing">let the 32B cook over the holidays</a> when compute demand was lower &#8212; why the heck a 1B now?</p><p>The 1B Instruct model is <a href="https://huggingface.co/allenai/OLMo-2-0425-1B-Instruct">here</a> and a GGUF for local users is <a href="https://huggingface.co/allenai/OLMo-2-0425-1B-Instruct-GGUF">here</a> (all OLMo 2&#8217;s are <a href="https://huggingface.co/collections/allenai/olmo-2-674117b93ab84e98afc72edc">here</a>).</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/subscribe"><span>Subscribe now</span></a></p><p>The reason is that we didn&#8217;t know our 1B base model was actually good enough. If you zoom in on the coming revision to the OLMo 2 paper you&#8217;ll see that the base model evaluations are largely &#8220;mid.&#8221; They&#8217;re decent enough to their peer models, but not as strong as the bigger models in the suite. We thought we had to keep pushing modeling decisions for a 1B model (such as fiddling with weight decay settings) or other things that are suitable for small models &#8212; i.e. older techniques that don&#8217;t scale up to bigger models so are out of fashion. Small model development can be handled much differently than bigger models.</p><p>This gap in how big vs. small models can be developed is part of the reason we suspect our final post-training results are strong compared to models like Gemma 3 1B or Llama 3.2 1B. Gemma 3 1B is the only model in its suite without a vision component and Llama is a multimodal release &#8212; maybe these changes made their text only performance weaker at the low end? We don&#8217;t quite now. </p><p>Here&#8217;s the shocking evaluation summary that keeps you reading!</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!85If!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!85If!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png 424w, /__u/substackcdn.com/image/fetch/$s_!85If!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png 848w, /__u/substackcdn.com/image/fetch/$s_!85If!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png 1272w, /__u/substackcdn.com/image/fetch/$s_!85If!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!85If!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png" width="1456" height="868" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:868,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:215035,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/162549503?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!85If!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png 424w, /__u/substackcdn.com/image/fetch/$s_!85If!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png 848w, /__u/substackcdn.com/image/fetch/$s_!85If!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png 1272w, /__u/substackcdn.com/image/fetch/$s_!85If!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b54f9f8-ab2d-4240-a795-52873ccd5d06_2970x1771.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Or the full evals for the 1B models:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!AWvB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!AWvB!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png 424w, /__u/substackcdn.com/image/fetch/$s_!AWvB!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png 848w, /__u/substackcdn.com/image/fetch/$s_!AWvB!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AWvB!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!AWvB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png" width="1456" height="622" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/95e5722c-7017-42c7-baf1-030702586882_1729x739.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:622,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:183874,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/162549503?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!AWvB!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png 424w, /__u/substackcdn.com/image/fetch/$s_!AWvB!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png 848w, /__u/substackcdn.com/image/fetch/$s_!AWvB!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png 1272w, /__u/substackcdn.com/image/fetch/$s_!AWvB!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e5722c-7017-42c7-baf1-030702586882_1729x739.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>You can see some things like formatting issues on DROP for Qwen or GSM8K for Gemma 3. These are the small details motivation evaluation change I&#8217;ll revisit later.</p><p>Turns out we had this 1B model sitting around for a while and it was only when we tried more pretraining tricks that we compared the post-training numbers. The post-training numbers were far better than we expected and made the model best in class! We were sitting on great results and a model the community could love for a while without knowing it was actually good.</p><p>The biggest problem here is that we don&#8217;t know how base model evaluations indicate a strong model for post-training. Trends seem to point to base model evals being the same as the evals used for post-training. Everything is about generating text well, and now chains of thought well. <a href="https://www.interconnects.ai/p/qwen-3-the-new-open-standard">Qwen 3</a>&#8217;s base model evaluations can point at this: </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!xdcD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!xdcD!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!xdcD!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!xdcD!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!xdcD!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!xdcD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg" width="1456" height="991" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:991,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!xdcD!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!xdcD!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!xdcD!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!xdcD!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efb6b32-4e18-41cb-b5d0-255fceb0e449_1554x1058.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>With these evaluations maybe the only thing to care about is perplexity over controlled chunks of text &#8212; i.e. how well the next token prediction is working.</p><p>If OLMo base models have a hard time competing in an era of mega compute, these insights will be our most valuable contributions. I&#8217;ve been pessimistic in the past about our ability to compete with the big players, but we just put out a 1B model competitive with very recent releases and we have been sitting on it for months!</p><p>The post-training for the 1B model again proved super robust. When you have a stable recipe, it works. We found the RL gains to be particularly robust &#8212; we mostly just had to let it keep running.</p><p>The biggest gap we&#8217;re trying to close now in post-training is a scalable reasoning recipe. If we want to release state of the art models on popular evaluations, scaling RL and inference-time compute is a requirement. We want to lead on problems like avoiding over-thinking, keeping reasoning usable, and so on, but we&#8217;ll see which innovations come first!</p><p>I&#8217;m personally feeling the big shift that all the leading AI labs have gone through in the last few months. Major changes in expectations comes with major changes in tooling and processes for changing. It&#8217;s exciting, but folks all over have been putting in serious effort to do that. </p><p>Let us know what you think of this 1B model. It&#8217;s been super fun to do mini research on and I suspect a lot of you will also like it for local inference tasks. What a great time to be in language modeling research. </p><p>And yes, you can make fun of us for the fact that our 1B model has 1.5B total parameters (1.3B without embedding parameters). We&#8217;ll focus on this more in the next versions &#8212; just one of those many things to get right.</p><p>Here&#8217;s the full OLMo 2 suite evaluations:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!E_As!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!E_As!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png 424w, /__u/substackcdn.com/image/fetch/$s_!E_As!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png 848w, /__u/substackcdn.com/image/fetch/$s_!E_As!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png 1272w, /__u/substackcdn.com/image/fetch/$s_!E_As!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!E_As!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png" width="1343" height="1349" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1349,&quot;width&quot;:1343,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:402578,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/162549503?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!E_As!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png 424w, /__u/substackcdn.com/image/fetch/$s_!E_As!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png 848w, /__u/substackcdn.com/image/fetch/$s_!E_As!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png 1272w, /__u/substackcdn.com/image/fetch/$s_!E_As!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed438061-c403-4cdf-9b04-53ef8bfe9ccc_1343x1349.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div>]]></content:encoded></item><item><title><![CDATA["Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?"]]></title><description><![CDATA[This isn't a new intuition, but a nice new set of results.]]></description><link>https://natolambert.substack.com/p/does-reinforcement-learning-really</link><guid isPermaLink="false">https://natolambert.substack.com/p/does-reinforcement-learning-really</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 21 Apr 2025 16:13:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The paper in question <em><a href="https://arxiv.org/abs/2504.13837">Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?</a> </em>has a lot of discussions underway on if Reinforcement Learning from Verifiable Rewards (RLVR) is actually improving the models we&#8217;re training.</p><p>The core figures are the following:</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading natolambert overflow! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!np7O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!np7O!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png 424w, /__u/substackcdn.com/image/fetch/$s_!np7O!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png 848w, /__u/substackcdn.com/image/fetch/$s_!np7O!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png 1272w, /__u/substackcdn.com/image/fetch/$s_!np7O!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!np7O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png" width="1456" height="476" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:476,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:339005,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/161810778?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!np7O!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png 424w, /__u/substackcdn.com/image/fetch/$s_!np7O!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png 848w, /__u/substackcdn.com/image/fetch/$s_!np7O!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png 1272w, /__u/substackcdn.com/image/fetch/$s_!np7O!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd96be96d-989e-4f76-a54d-3dd7f6d6208d_2336x764.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>And:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!WIc6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!WIc6!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png 424w, /__u/substackcdn.com/image/fetch/$s_!WIc6!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png 848w, /__u/substackcdn.com/image/fetch/$s_!WIc6!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WIc6!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!WIc6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png" width="1456" height="1412" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1412,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:729592,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/161810778?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!WIc6!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png 424w, /__u/substackcdn.com/image/fetch/$s_!WIc6!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png 848w, /__u/substackcdn.com/image/fetch/$s_!WIc6!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WIc6!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87153d2d-05bc-401a-8b80-d77ed296e046_1650x1600.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>These are all using pass@k as the core metric. Pass@k is the metric that checks to see if the right answer exists in k completions. This is not how practical inference works, but is a good test to see if the model is &#8220;in distribution.&#8221;</p><p>The argument is that as high K, the base models do the same or outperform the RL models. This is cool because it shows that RL reduces the entropy of samples but makes the model more effective at pass@1. We know any post training will reduce the variance in answers and this is a new way to see it. Higher variance will mean more likely to pass the pass@k metric for high k. Honestly, <strong>as we get better at RL we should be able to make it induce more exploration, given that&#8217;s a core topic of RL&#8217;s entire existence</strong>, well before language models.</p><p>The paper is trying to make you ask: If base models outperform RL trained models, why do we need this RL training?</p><p>You should focus on the bottom three rows here which are in-distribution for the RL training data. Second, Qwen is known to be predisposed to learning reasoning so those base models may be stronger. I can&#8217;t say a lot more about the base models as that&#8217;s an open area of research &#8212; what are the right base models for reasoning?</p><p>Some are surprised that the base model does so well, but really we&#8217;ve been saying for a while that RL training <a href="https://www.interconnects.ai/i/158302462/rls-role-in-elicitation">is increasing the probability of correct behaviors &#8212; elicitation</a>. With this view, the results align totally with what RLVR should be doing.</p><p>There are also some caveats on the work that make it have the usual academic grains of salt. Mostly, they only train on the MATH and GSM8K training sets. While this is great for controlled ablations, it&#8217;s not great for showing the fundamental limits of RL training. OpenAI and others have shown that <em>scaling</em> RL is a crucial aspect of it, and with only these narrow training sets that isn&#8217;t really possible.</p><p>Second, the paper doesn&#8217;t have a ton of plots showing the training curves for their models. It&#8217;s safe to assume they&#8217;re decent because the results look reasonable, but the base model training is much more reliable than the RL training in another paper trying to make a point.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/p/does-reinforcement-learning-really?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/p/does-reinforcement-learning-really?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p><p>The pass@1 results for RL are extremely promising and should <strong>reinforce that RLVR is working</strong>. That being said, if we had perfect verifiers  &#8212; an oracle &#8212; we&#8217;d never need RLVR in the first place (or post-training really), and we could just use that instead of trying to make the model better. <strong>My <a href="https://www.interconnects.ai/i/148388607/inference-scaling-laws">very first post on inference models</a> made the same point that random sampling with pass@k metrics is important as a baseline for inference scaling!</strong></p><p>This isn&#8217;t new. This is a nice reminder that there&#8217;s no free lunch. We should keep checking how this changes as we:</p><ul><li><p>Scale RL training to many more prompts, and</p></li><li><p>Scale RL to bigger base models.</p></li></ul><p>A final caveat, which I think is minor. These results are all RL-Zero style, i.e. just on the base model with no warm start. DeepSeek <a href="https://www.interconnects.ai/i/159577063/kimi-k-scaling-reinforcement-learning-with-llms">and others</a> stated that better performance comes from a warmup with on-policy SFT before RL. This&#8217;ll make the RL results above even stronger, where the base model results won&#8217;t change.</p><p>I&#8217;m still optimistic on RL. Come on, don&#8217;t go too goldfish, we just got o3 and Gemini 2.5, wonderful models with RL. Still thinking it&#8217;s a dead end?</p>]]></content:encoded></item><item><title><![CDATA[Looking at the training data]]></title><description><![CDATA[On building tools where truly open-source models can shrine (OLMo 2 32B Instruct, for today). OLMoTrace lets you poke around.]]></description><link>https://natolambert.substack.com/p/looking-at-the-training-data</link><guid isPermaLink="false">https://natolambert.substack.com/p/looking-at-the-training-data</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Wed, 09 Apr 2025 20:09:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kJsE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>If I hand you 6 trillion tokens of pretraining data for a language model, what would you even do with it to learn about the resulting model? The training data is obviously the lifeblood of modern AI models, but is usually swept under the rug with statements like &#8220;high-quality, publicly available data&#8221; and so on. The size is also extremely unwieldily. Regular readers of <a href="https://www.interconnects.ai/t/open-source">Interconnects</a> know that very few models actually release the data. </p><p>After fighting the battle of getting more people to release the data comes showing people what the data is useful for. Today, Ai2 launched a new tool OLMoTrace that takes a major step forward for this. It isn&#8217;t <em>search</em> of the training data per-se (which you could do with <a href="https://huggingface.co/datasets/allenai/olmo-mix-1124">HuggingFace datasets</a> if you were so bold). Search, or direct attribution of model outputs to data sources, provide much more complex legal risks. OLMoTrace makes <em>attribution</em> to all the datasets used in training. It makes an index where similar n-grams can be found. The underlying technology is called Infinigram, which is described as:</p><blockquote><p><strong>Infini-gram</strong> is an engine that efficiently processes n-gram queries with <strong>unbounded n</strong> and <strong>trillion-token massive corpora</strong>.</p></blockquote><p>This isn&#8217;t a tool for mechanistic interpretability, but rather a tool for intuitions and education on what goes into a model and what random attributions can cause a model to say something. A big audience of this will be policymakers being able to show what is and what is not attributable to certain pieces of training data. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/p/looking-at-the-training-data?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/p/looking-at-the-training-data?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p><p>Where I put OLMo 2 32B Instruct as an underrated milestone for the open community &#8212; as <a href="https://www.interconnects.ai/p/gemma-3-olmo-2-32b-and-the-growing">an open-source model that is arguably at GPT 4 level</a>. OLMoTrace adds another huge milestone to the open-source ecosystem, where people can now build stronger understanding of the role of data. The goal is that more people can build with the models better because the data is visible. This will be a slow burn, but hopefully adds more to the recent momentum of open models and tools.</p><p>Links:</p><ul><li><p>OLMoTrace blog post: https://allenai.org/blog/olmotrace</p></li><li><p>Play with it: https://playground.allenai.org/</p></li><li><p>Powered by Infinigram research: https://infini-gram.io/</p></li><li><p>Can also use OLMo 2 32B Instruct <a href="https://openrouter.ai/allenai/olmo-2-0325-32b-instruct">on OpenRouter</a> (and soon on Google Vertex AI)</p></li></ul><p>Things that are useful with this:</p><ul><li><p>Checking if behaviors for your model that you don&#8217;t expect show up in the post-training data.</p></li><li><p>Checking if niche content, words, or phrases are in the data at all.</p></li><li><p>Building any sort of intuition on what LLM training data looks like.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!kJsE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!kJsE!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png 424w, /__u/substackcdn.com/image/fetch/$s_!kJsE!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png 848w, /__u/substackcdn.com/image/fetch/$s_!kJsE!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kJsE!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!kJsE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png" width="1456" height="807" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:807,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Screenshot of the Ai2 playground with OLMoTrace in action. Several text spans of the model output are highlighted. On the right, the OLMoTrace panel shows blocks of documents from the training data.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Screenshot of the Ai2 playground with OLMoTrace in action. Several text spans of the model output are highlighted. On the right, the OLMoTrace panel shows blocks of documents from the training data." title="Screenshot of the Ai2 playground with OLMoTrace in action. Several text spans of the model output are highlighted. On the right, the OLMoTrace panel shows blocks of documents from the training data." srcset="/__u/substackcdn.com/image/fetch/$s_!kJsE!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png 424w, /__u/substackcdn.com/image/fetch/$s_!kJsE!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png 848w, /__u/substackcdn.com/image/fetch/$s_!kJsE!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kJsE!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b8adeea-1d8f-4c26-8f53-6d8d149799a0_1462x810.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>As we starting building towards character training at Ai2, this is very useful for figuring out if behaviors we don&#8217;t like will be easy to remove or not. This is far easier to interpret with post-training data, but in my testing that is where a lot of the phrases we are used to like &#8220;I cant do that&#8221; come from. </p><p>For example, I was probing OLMo 2 32B Instruct on a lot of basic things like date-cutoff, if AIs should have personalities, and pushing the boundaries of its notion of self. In these, a lot of the responses are drawn directly from the SFT mix. For example, I ask:</p><blockquote><p>Should AI's like yourself have strong opinions?</p></blockquote><p>Then, the model responds with:</p><blockquote><p>As a <strong>language model, I am programmed to maintain</strong> neutrality and avoid expressing strong opinions. My purpose is to provide information, answer questions, and facilitate understanding based on existing knowledge, without leaning toward personal biases or strong opinions. However, I can summarize or analyze the arguments and viewpoints on various topics. It's crucial to remember that an AI's role should not be to express personal beliefs but to assist users fairly and impartially.</p></blockquote><p>The part highlighted in bold above comes directly from multiple examples in the SFT mix, such as:</p><blockquote><p>"...heart attacks in the past year. &lt;|assistant|&gt; I apologize, but I cannot fulfill your request for access to medical records. As a responsible and ethical AI <strong>language model, I am programmed to maintain</strong> the privacy and confidentiality of patient information. Access to medical records is strictly regulated and can only be granted by authorized healthcare professionals and institutions, and only for legitimate and..."</p></blockquote><p>This is the sort of thing that I would remove. There are many more cases like this.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/subscribe"><span>Subscribe now</span></a></p><p>Another thing we already used OLMo trace to better control was phrases like &#8220;as of my date cutoff in XXYY,&#8221; opting to handle this in future system prompts (filtering code <a href="https://github.com/allenai/open-instruct/blob/main/scripts/data/filtering_and_updates/filter_cutoff_date.py">here</a>).<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><p>Finding links to pretraining data is easier when asking far more general queries. For example:</p><blockquote><p><strong>Give me a 5 sentence summary of the Great Depression</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!k3_V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!k3_V!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png 424w, /__u/substackcdn.com/image/fetch/$s_!k3_V!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png 848w, /__u/substackcdn.com/image/fetch/$s_!k3_V!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png 1272w, /__u/substackcdn.com/image/fetch/$s_!k3_V!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!k3_V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png" width="908" height="695" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:695,&quot;width&quot;:908,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:228482,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/160963838?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!k3_V!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png 424w, /__u/substackcdn.com/image/fetch/$s_!k3_V!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png 848w, /__u/substackcdn.com/image/fetch/$s_!k3_V!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png 1272w, /__u/substackcdn.com/image/fetch/$s_!k3_V!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecb9d0a-8a2e-42db-9e29-8852aa203b59_908x695.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Still, the SFT data here has a big impact (just because there was a topic overlap in the SFT mix again). Something we&#8217;ll need to do for the internal version is to turn our chat safety filter off. This will let us have a much nicer tool to attribute either harmful text to or an incorrect refusal.</p><p>Happy playing! Let us know what you find.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Right now OLMoTrace has the &#8220;rejected&#8221; DPO samples in it that still have some of these phrases.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Some thoughts on OpenAI returning to open releases]]></title><description><![CDATA[And welcome to my extra uneditted thoughts blog.]]></description><link>https://natolambert.substack.com/p/some-thoughts-on-openai-returning</link><guid isPermaLink="false">https://natolambert.substack.com/p/some-thoughts-on-openai-returning</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Tue, 01 Apr 2025 01:46:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NrSt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Back when I changed jobs from HuggingFace to AllenAI in the Fall of 2023, I had a few conversations with folks at OpenAI and I had heard of substantial discussions back then of trying to re-up their relationship with open (source) AI developments. Today, a year and a half later, <a href="https://x.com/sama/status/1906793591944646898">OpenAI is prepping for their first open-weight language model release since GPT-2</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/subscribe"><span>Subscribe now</span></a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!NrSt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!NrSt!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png 424w, /__u/substackcdn.com/image/fetch/$s_!NrSt!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png 848w, /__u/substackcdn.com/image/fetch/$s_!NrSt!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NrSt!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_webp, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!NrSt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png" width="1066" height="659" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:659,&quot;width&quot;:1066,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:119375,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://natolambert.substack.com/i/160308826?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3049944-7990-4ffe-b8a8-95b1273447cb_1066x1342.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!NrSt!, /__u/natolambert.substack.com/w_424, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png 424w, /__u/substackcdn.com/image/fetch/$s_!NrSt!, /__u/natolambert.substack.com/w_848, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png 848w, /__u/substackcdn.com/image/fetch/$s_!NrSt!, /__u/natolambert.substack.com/w_1272, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NrSt!, /__u/natolambert.substack.com/w_1456, /__u/natolambert.substack.com/c_limit, /__u/natolambert.substack.com/f_auto, /__u/natolambert.substack.com/q_auto:good, /__u/natolambert.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b27949b-b0fe-48ad-8de9-8034c3392b4e_1066x659.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This is great. It seems like they&#8217;re doing this earnestly, even if the timing in the change of heart may look political. The most telling reason as to why they&#8217;re doing this <a href="https://x.com/bradlightcap/status/1906823735178604657">came from their COO Brad Lightcap</a>:</p><blockquote><p>Primarily because developers, along with many of our business and government customers, have asked for it. We serve millions of developers every day &#8211; many rely on a mix of proprietary and open models to build AI products. We&#8217;ve provided frontier models through our API since 2020, but we&#8217;ve also released models like GPT-2 and Whisper to the open-source community.</p></blockquote><p><a href="https://huggingface.co/openai/whisper-large">Whisper</a>, the speech recognition system, is a super impactful open release. OpenAI largely definitely knows what they&#8217;re doing here. The rest of the thread from their COO is very telling:</p><blockquote><p>As models improve, there is more and more demand to run them everywhere. Through conversations with startups and developers, it became clear how important it was to be able to support a spectrum of needs, such as custom fine-tuning for specialized tasks, more tunable latency, running on-prem, or deployments requiring full data control. While we will continue to offer frontier models via our API and in ChatGPT, there are many scenarios where APIs alone won&#8217;t fully enable developers to build in the place, or in the way, in which they&#8217;d like.<br><br>Our goal with this open model is to address exactly that: expanding developer access to powerful AI, while continuing to maintain high standards for safety and responsible deployment. We think this is important for both the US and the world, especially as AI becomes foundational to the world&#8217;s computing infrastructure.</p></blockquote><p>OpenAI has a <a href="https://openai.com/open-model-feedback/">feedback form</a> for how they should manage this, which many are mocking, but it&#8217;s clear they&#8217;re trying to learn which type of use-cases support a open models approach that API models don&#8217;t serve. This is an extremely bullish sign for the open community and not just OpenAI trying to collect random information because they can. </p><p>The beginning of 2025 has marked the highpoint of open weight models since ChatGPT, as I discussed at length in my analysis in the main feed on Gemma 3 and OLMo 2:</p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:158931739,&quot;url&quot;:&quot;https://www.interconnects.ai/p/gemma-3-olmo-2-32b-and-the-growing&quot;,&quot;publication_id&quot;:48206,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Interconnects&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe70f9dbf-4fe6-404c-b6bb-1831d1b7ed0b_590x590.png&quot;,&quot;title&quot;:&quot;Gemma 3, OLMo 2 32B, and the growing potential of open-source AI&quot;,&quot;truncated_body_text&quot;:&quot;Ever since the release of the original ChatGPT, much has been said about making a truly open-source version of it &#8212; with data, code, weights, etc., all available. Open-source versions increase transparency, access, long-term progress, security research, and lots more. Lots of people have used this claim to bring hype into their projects, but the substan&#8230;&quot;,&quot;date&quot;:&quot;2025-03-13T18:16:11.772Z&quot;,&quot;like_count&quot;:69,&quot;comment_count&quot;:0,&quot;bylines&quot;:[{&quot;id&quot;:10472909,&quot;name&quot;:&quot;Nathan Lambert&quot;,&quot;handle&quot;:&quot;natolambert&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fedcdfb-e137-4f6a-9089-a46add6c6242_500x500.jpeg&quot;,&quot;bio&quot;:&quot;ML researcher making sense of AI research, products, and the uncertain technological future. PhD from Berkeley AI. Experience at Meta, DeepMind, HuggingFace.&quot;,&quot;profile_set_up_at&quot;:&quot;2021-04-24T01:19:33.371Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:100753,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:48206,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:48206,&quot;name&quot;:&quot;Interconnects&quot;,&quot;subdomain&quot;:&quot;robotic&quot;,&quot;custom_domain&quot;:&quot;www.interconnects.ai&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;The cutting edge of AI, from inside the frontier AI labs, minus the hype. The border between high-level and technical thinking. Read by leading engineers, researchers, and investors on Wednesday mornings.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e70f9dbf-4fe6-404c-b6bb-1831d1b7ed0b_590x590.png&quot;,&quot;author_id&quot;:10472909,&quot;theme_var_background_pop&quot;:&quot;#ff6b00&quot;,&quot;created_at&quot;:&quot;2020-05-21T02:59:47.895Z&quot;,&quot;email_from_name&quot;:&quot;Interconnects by Nathan Lambert&quot;,&quot;copyright&quot;:&quot;Interconnects AI, LLC&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false}},{&quot;id&quot;:4610799,&quot;user_id&quot;:10472909,&quot;publication_id&quot;:4519930,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:4519930,&quot;name&quot;:&quot;natolambert overflow&quot;,&quot;subdomain&quot;:&quot;natolambert&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;a place for any extra thoughts beyond Interconnects.ai&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png&quot;,&quot;author_id&quot;:10472909,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2025-03-27T15:04:05.448Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Nathan Lambert&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false}}],&quot;twitter_screen_name&quot;:&quot;natolambert&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.interconnects.ai/p/gemma-3-olmo-2-32b-and-the-growing?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="/__u/substackcdn.com/image/fetch/$s_!Snpy!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe70f9dbf-4fe6-404c-b6bb-1831d1b7ed0b_590x590.png" loading="lazy"><span class="embedded-post-publication-name">Interconnects</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Gemma 3, OLMo 2 32B, and the growing potential of open-source AI</div></div><div class="embedded-post-body">Ever since the release of the original ChatGPT, much has been said about making a truly open-source version of it &#8212; with data, code, weights, etc., all available. Open-source versions increase transparency, access, long-term progress, security research, and lots more. Lots of people have used this claim to bring hype into their projects, but the substan&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">a year ago &#183; 69 likes &#183; Nathan Lambert</div></a></div><p>I recall this section conclusion:</p><blockquote><p>The biggest stories in open-source AI in 2024 often felt like bickering about definitions. I wrote <a href="https://www.interconnects.ai/p/defining-open-source-ai">a lot</a> <a href="https://www.interconnects.ai/p/flavors-of-open-source-ai">of articles</a> <a href="https://www.interconnects.ai/p/an-open-source-llm">about definitions</a>. Llama 3 was pretty much all we had to get excited about. At the end of the day, even with how much I think it would be better with more information on the whole stack of AI development, open-source is largely going to be defined by community norms. For now, Llama weights have been that norm rather than other definitions.</p><p>By comparison, 2025 feels poised to be about <em>actually building open AI</em>. We have had surprising, impactful, and exciting releases and it&#8217;s only March. We know Meta is looking to get back into the conversation with Llama 4 in April at LlamaCon. We have our open-source ChatGPT. We&#8217;ll have more we can&#8217;t predict.</p><p>Crucially, on top of the gap being smaller, all of these open models are crossing meaningful boundaries in performance. When model capabilities made the leap to GPT 4 class models, tons more applications were possible. Now, we have GPT 4 class <em>small</em> models that can be deployed in privacy-conscious ways. There&#8217;s been a huge demand for this, and the ecosystem is slowly building the tools to do so. Yes, closed AI will continue to march forward, but open solutions need to prove their own independent feasibility.</p><p>In the long march of progress, open-source AI feels far closer to an inflection point of proving out the hypothetical benefits we have focused on for a few years. Transparency, privacy, better performance, etc. could actually all be happening this year.</p></blockquote><p>It increasingly feels true.</p><p>I&#8217;m giving OpenAI and others looking to get into open models the following advice:</p><ol><li><p>Be careful with the open-source vs. open weight nomenclature even if its annoying. The open-source focused users and community members have a point and their efforts with the broader ecosystem are worth supporting until we have answers. Blatantly, you don&#8217;t want to make enemies with a release.</p></li><li><p>Ecosystem integrations being handled professionally, such as VLLM, Llama.cpp, etc. are crucial.</p></li><li><p>Licenses matter a ton for adoption (<a href="https://x.com/sama/status/1906845532758405319">here</a> are Sam Altman&#8217;s comments on the matter). Non commercial models are going to matter less and less. This factor is increased as the pace of progress is so high because having a good license increases the chance that people adopt your model in the time before the next state of the art open release comes.</p></li></ol><p>All in, it&#8217;s a major moment for the open ecosystem that the biggest name in AI is rejoining it. It&#8217;s the biggest validation we&#8217;ve gotten that building with open models is a needed way to solve certain types of business problems and the solutions will provide real value. </p><p>Excited! As with all things in the world, we can&#8217;t give them credit until they happen. We can reflect on motivations and give advice, but impact and credit comes later.</p><div><hr></div><h3>What will they release?</h3><p>Some basic logic for why OpenAI will probably release an ~30B param reasoning model with MIT/apache.</p><p>* OpenAI will only release something clearly SOTA in size category</p><p>* Will release a reasoning model, else hard to tell story around popular evals</p><p>* Want people to use it, cant be too big</p><p>* Don't want to compete with API models, can't be too big</p><p>* 30B is loved by people finetuning models (can do on one node) + running inference, but still gives some nuance of bigger model feel</p><p>* I expect the architecture to be much simpler than their internal models, maybe even based on Qwen / Llama to not reveal secret sauce</p><p>* Curious if they go knowledge distillation route like Gemma</p><p>* MIT / apache license undercuts Google and Meta with their weird licenses</p><p>* Not too sure about no base model, but pretty sure about no data.</p><p>competition is gemma 3, qwen 2.5 (or 3 soon), mistral 3.1.</p>]]></content:encoded></item><item><title><![CDATA[Coming soon]]></title><description><![CDATA[This is natolambert overflow, a place for if I want to write more and not send it to the entire Interconnects list.]]></description><link>https://natolambert.substack.com/p/coming-soon</link><guid isPermaLink="false">https://natolambert.substack.com/p/coming-soon</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Thu, 27 Mar 2025 15:04:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ntyE!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb88d599-32c8-49a9-ba33-ab6327aff727_256x256.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This is natolambert overflow, a place for if I want to write more and not send it to the entire Interconnects list.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://natolambert.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/natolambert.substack.com/subscribe"><span>Subscribe now</span></a></p>]]></content:encoded></item></channel></rss>