<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Interconnects AI]]></title><description><![CDATA[The cutting edge of AI, from inside the frontier AI labs, minus the hype. The border between high-level and technical thinking. Read by leading engineers, researchers, and investors.]]></description><link>https://www.interconnects.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!djof!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc52e8097-8f3d-4f7e-808b-2f4ad37f3b52_720x720.png</url><title>Interconnects AI</title><link>https://www.interconnects.ai</link></image><generator>Substack</generator><lastBuildDate>Wed, 29 Jul 2026 06:44:44 GMT</lastBuildDate><atom:link href="https://www.interconnects.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Interconnects AI, LLC]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[mail@interconnects.ai]]></webMaster><itunes:owner><itunes:email><![CDATA[mail@interconnects.ai]]></itunes:email><itunes:name><![CDATA[Nathan Lambert]]></itunes:name></itunes:owner><itunes:author><![CDATA[Nathan Lambert]]></itunes:author><googleplay:owner><![CDATA[mail@interconnects.ai]]></googleplay:owner><googleplay:email><![CDATA[mail@interconnects.ai]]></googleplay:email><googleplay:author><![CDATA[Nathan Lambert]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next]]></title><description><![CDATA[A podcast with Florian Brand.]]></description><link>https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3</link><guid isPermaLink="false">https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Wed, 22 Jul 2026 14:09:04 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207969620/a82e3d40f39345b0609610ffedd72469.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<h5>Exciting news! My <a href="https://rlhfbook.com/">book</a> trying to share post-training knowledge with the world is done and shipping soon. Order on <a href="https://hubs.la/Q03TsMBq0">Manning</a> or <a href="https://amzn.to/4cwCDJQ">Amazon</a>. Thanks for the support. It&#8217;s currently the #1 AI book on Amazon :).</h5><p><span>Nathan and Florian sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating &#8212; geopolitics of US v China, economics of open vs. closed models, security at the frontier of AI, and so on.<br><br>Chapters:<br>00:00 Welcome &amp; context<br>04:38 Living with / using Kimi K3<br>08:53 GLM 5.2&#8217;s continued role<br>12:47 How are the Chinese models this good?<br>17:41 Data, environments, and a tour of the Chinese labs<br></span>19:47 Roundup of Chinese providers: Qwen, DeepSeek, MiniMax&#8230;<span><br>24:08 The US open-model ecosystem<br>30:25 Frontier vs. near-frontier, and the cybersecurity case against bans<br>34:58 Distillation and the Ben Thompson debate<br>44:12 Predictions and a frontier tier list<br>48:36 Wrap-up</span></p><p><span>Listen on </span><a href="https://podcasts.apple.com/us/podcast/interconnects-audio/id1719552353">Apple Podcasts</a><span>, </span><a href="https://open.spotify.com/show/6XNzfJULeVxR7SneeesDUs">Spotify</a><span>, and </span><a href="https://www.interconnects.ai/podcast">where ever you get your podcasts</a><span>. For other Interconnects interviews, </span><a href="https://www.interconnects.ai/t/interviews">go here</a><span>.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/open-models-recap-more-on-kimi-k3?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p><span>For more educational post-training videos, see the </span><a href="https://rlhfbook.com/course">course</a><span> I&#8217;m putting together.</span></p><div id="youtube2-XsBy8UGIY-I" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;XsBy8UGIY-I&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/XsBy8UGIY-I?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2>Transcript</h2><p><strong><span>00:00:06 Nathan Lambert:</span></strong><span> Okay, welcome back to Interconnects. We&#8217;re doing our quarterly open model roundup, which is mostly us just making fun of or explaining, not making fun, why so many distillation takes are bad and understanding the state of where things stand. I think last Thursday was when Kimi K3 was released. I think we will see much much more in the near future. It seems pretty inevitable. Like over the weekend, Xi gave his speech where he directly committed to openness and open source as a strategy. It wasn&#8217;t a detailed layout state of affairs.</span></p><p><span>Qwen announced their next big model is going to be open weight, which is a big change of things. I think there&#8217;s just so much to get into. I think Flo you kind of were already going off on some of the performance gap and distillation takes. So we could probably start there and then as I go I have a little bit a little list and we could always go through the topics and the blog that I wrote which all are very nuanced. So I think we have infinite to talk about. So continue rant kind.</span></p><p><strong><span>00:01:17 Florian Brand:</span></strong><span> Yeah, I think, or the biggest thing at every model release at least at every open model release is how much or how many months it is behind the closed frontier. Um and people love to put a definite uh definitive number onto this uh which is really really mudding because we have so many different benchmark providers these days and such uh so many different benchmarks as well that every site and I&#8217;m not innocent in that either um pulls up their favorite benchmarks to show that the current model or the newly released model is at the frontier which is then counted by the other side pulling up another benchmark and showing oh it&#8217;s actually a year behind or something. Um and it like a lot of it seemingly hinges on that question how many months we open models are behind.</span></p><p><strong><span>00:02:26 Nathan Lambert:</span></strong><span> Yeah. So I my provocation is that some of the benchmarks are actually reasonably correlated with what people are doing and this is agentic coding and agentic computer use tasks and some of the benchmarks are correlated with the long tail which is where I think Claude and GPT is so valuable. But if it&#8217;s it&#8217;s like what is the market for Claude Code and Codex right now and if it is software engineering then like the fact that the models are say a couple months behind on that can be a very, very big deal and then I suspect that this model will be okay disclaimer the model weights aren&#8217;t out yet supposedly on 20 July 27th and a lot of the discussion will impinge on the assumption that they come.</span></p><p><span>But like people could post-train this model to very likely match Opus and GPT in many of these kind of niche domains that people want I think watching I mean we both have different views into the post-training open model industry, but there is a ton ton of excitement in progress on making these models like fine-tuned for specific high-value tasks and this has been historically done on a mix of like Qwen and GLM and GLM 5.2 really accelerated this and I I curious on the first person that puts out a blog post like we fine-tuned Kimi K3 on our task because I bet you could get big gains. I think even the you use Kimi K3 more than I do, but my hunch is that it would be a bit of a um rough edged post-training just by how big of a scale up it is and that normally means there&#8217;s a lot of performance that could still be extracted from it. No.</span></p><p><strong><span>00:04:03 Florian Brand:</span></strong><span> Yeah. Running, running, running and especially post-training that one will be super hard because you need like one node of B300s to just load the weights which is crazy in terms of scale. So will probably take some time and uh a lot of engineering I&#8217;ve heard to actually get this into a state where it&#8217;s fine-tunable. Um but people you want to talk about using the model like you actually signed up for the the coding program and used it. So like getting this out there is good context.</span></p><p><strong><span>00:04:38 Nathan Lambert:</span></strong><span> Yeah. So I signed up on day after release or so uh for the $200 plan uh which is their biggest one similar to to all the others but they have um like I think $40 and $100 as well. Uh but the biggest plan has uh 1 million context and I think or at least it feels like it has also some priority in terms of the API requests because so many people um online are saying that they hit API errors constantly and so far I&#8217;ve been I&#8217;ve been uh pretty well off if I&#8217;m uh going to say that. Um and in terms of model capabilities, it at some ways aside from front end where it is really good, it in some ways it really shines and it excels.</span></p><p><span>Um even my um expectations even with things like uh some research tasks like I have or at interconnects we now have over a year of data on on open models um and I ask the frontier models to come up with some interesting analysis which we haven&#8217;t done before in uh because we do our own analysis and have this published uh and I asked them all right do something new and um surprise me, basically. And a lot of the models or the frontier models u or basically all models latch onto the things we&#8217;ve done redo the data analysis part and then do some weird esoteric parts.</span></p><p><span>Uh Kimi K3 did some more interesting things um I&#8217;ve told it explicitly to scrape Reddit um and then it found uh some subreddits I haven&#8217;t even considered and then found out for example that the Reddit discussions are um one or two months in uh more recent or they found they find the interesting models one or two months before the download numbers usually take off like they are all onto Qwen and then the people download more Qwen models like those kind of analysis is groundbreaking um but it is something that Kimi surprised me at compared to to all the other frontier um models.</span></p><p><span>A simple question like can you use this for most of the core work you do in terms of like the exp you you have a distribution of stuff you tend to do most of them are with Codex I think you&#8217;re a Codex person rather than Claude person like what percentage do you think the Venn diagram overlaps where this model would be fine</span></p><p><strong><span>00:07:24 Florian Brand:</span></strong><span> uh it&#8217;s really depends on how much leeway I give it like, the big thing I have seen with Kimi K3 right now I&#8217;m I&#8217;m working on u the framework we are doing at uh at Prime Intellect, where I work, and the main thing I found with Kimi is its code is a lot simpler uh which makes it way more readable uh but it misses some things that Codex just or like we&#8217;re talking 56, 55 and especially 54 would be on those levels. So, I would say that Kimi K3 is like 54-55 level for these kind of tasks.</span></p><p><span>But if I like I read the code and I say all right that&#8217;s really good code and then I give it a pass over with with Codex and it finds all these niche niche cases where it doesn&#8217;t excel but for supervising runs or for running uh some experiments it is actually really usable. Um and for some other niche things like you can just let it run. The one downside is but that&#8217;s also because the API is completely swamped in terms of users and it their servers are in China. The wall clock time is significantly significantly higher than GPT. But I would say like if I was to to push it and use it in my daily workflow, I would be slower, but I wouldn&#8217;t be slowed down by so much that I would say, &#8220;All right, that&#8217;s unusable.&#8221;</span></p><p><strong><span>00:08:53 Nathan Lambert:</span></strong><span> And how does this compare to GLM 5.2? Because GLM 5.2 was still a story unfolding in my opinion where like I would go bop around SF and people are like yeah I genuinely use this for this part of my like agentic coding and/or workflow. Um, how do you like I feel like were you in that camp using GLM at all or</span></p><p><strong><span>00:09:20 Florian Brand:</span></strong><span> Yeah. like where do you I also use used and use uh GLM mostly because we have an internal endpoint which is really fast and we have or or before that I I also used an API which had I don&#8217;t know 200 or 300 tokens per second. Um and if you can do a lot of task at a good enough level like really fast you just use that model compared to going to Codex then selecting the lesser model then selecting the right reasoning effort then selecting fast like I just use GLM get the same result and uh and it&#8217;s uh pretty fine like it it definitely is Sonnet-ish level in terms of capabilities and for a lot of cleanup task for a task that just is grunt work. It really works. Like I I would say you could probably go really far for a lot of the work uh with Kimi K3 as the main agent and GLM for for sub agent work.</span></p><p><strong><span>00:10:24 Nathan Lambert:</span></strong><span> Something that&#8217;s pretty different with Kimi&#8217;s announcement and the scale of models this is. I think it&#8217;ll take a bit longer for these open models to really be optimized and available across the inference providers. Like GLM 5.2 is pretty fast, but one, we don&#8217;t have the weights yet, and then two, like I don&#8217;t think it&#8217;s going to be as fast of a roll out on adoption as the like 500B, 700B MoE. like I I there&#8217;s going to be more problems there which is a very different regime where in the past the Chinese models would finish their RL run and release the model with open weights within hours to days maybe a week and then like immediately the ecosystem kind of knew how to do this.</span></p><p><span>I think there&#8217;s a lot bigger of an infrastructure kind of uplift on this next scale of open weight models which I think we have to factor in like the closed labs do this behind the scenes before announcing the models. So it&#8217;s just like that is kind of manipulating the time gap in a way where it could be like an extra month before people can actually post-train and use Kimi at scale for their workflows. And like we love to say as an open weight fan like oh it&#8217;s only when the closed model is available that you could take the time gap but like now there&#8217;s similar dynamics in open models where it&#8217;s like the Kimi API is totally broken. There&#8217;s too much supply. There&#8217;s too much demand. There&#8217;s not enough supply. So it&#8217;s not like this model is immediately diffusing like the I&#8217;m just I&#8217;m just thinking about this as it relates to the performance time gap</span></p><p><strong><span>00:11:54 Florian Brand:</span></strong><span> because that that is true but on the other hand the open ecosystem has professionalized quite a lot in the last few months. uh like during or in your initial roll out they all come with some partners which have the weights beforehand. They have the vLLM patches out days or or even weeks before these days which is completely different from from a year ago where basically weights got dropped and the model makers were like all right you got to figure this out. So I expect like the general availability on day one will be pretty okay and then race starts of all the providers starting to optimize to get even higher and higher speeds because it&#8217;s so much prestige.</span></p><p><strong><span>00:12:47 Nathan Lambert:</span></strong><span> Yeah. Okay. Two directions to go. Why do we think the Chinese models are able to be this good? I think I&#8217;ve wrote about there&#8217;s a debate in our Discord with in with JSD at at </span><a href="https://epoch.ai/"><span>Epoch</span></a><span> and I think it&#8217;s very good and I had this section in my piece that I&#8217;m like coming around to think that the Chinese labs are more capital efficient and you can turn capital into compute data and talent in a way that makes the models better and this is really I think this is super important if it actually is some structural advantage whatever the cause I think the cause could be talent is better trained for whatever their education system was to work on problems that make LLMs better.</span></p><p><span>It could just be that all the compute and talent and everything cost way less in China somehow. Whether it&#8217;s a subsidy, whether it&#8217;s just average pay being lower. But this is a very big deal as we turn the crank in the model iterations. And if a next generation model costs $10 billion for Anthropic but only $4 billion for Kimi like this this is like could be very huge but it&#8217;s not clear why this is the case. For example, I think Big Eagle the Kimi engineer replied to my tweet on this and was like it helps because we&#8217;re not trying to push the frontier. are just trying to catch up, which really could be a mindset thing where how the the goals of the labs are scoped in in China so that it cost them way less money to build these models.</span></p><p><span>But I in the last year we&#8217;ve asked a lot of questions on like will the Chinese models fall off. I have thought that the gap between closed and open bottles would grow due to this kind of capital intensity of training and it seems like it&#8217;s going the opposite direction which is just like it&#8217;s it&#8217;s hard to unpack but like do you agree that the labs are keeping up a bit more than we would have expected as in the Chinese labs and why?</span></p><p><strong><span>00:14:43 Florian Brand:</span></strong><span> Well, I I actually looked at our uh predictions for uh for this year based on our last year&#8217;s recap and we basically said that the gap will stay with within a few months. Uh so that prediction seems to largely hold. Um luckily for us, we didn&#8217;t put a concrete number whether it&#8217;s 3 months, 6 months or 9 months. So we are safe on that side. Um but I think like the general thing we both felt when we were in China and talking to these people like they are like the researchers themselves are teams of two or 300 people all mid20s and all just want one model to be really good like they don&#8217;t seem to do any side quests.</span></p><p><span>They don&#8217;t seem to do anything that uh deviates from from these things. And um they might or in terms of compute which is a really hard question for for us to answer especially as uh these Chinese uh chips are now coming online. We have I also think chips I think chip smuggling increased substantially in the last like 6 to 9 months or the chips that have been smuggled started to become online.</span></p><p><strong><span>00:15:58 Nathan Lambert:</span></strong><span> Smuggling is a general term for getting around export restrictions. If the chips are in Malaysia and they&#8217;re using them, I I count that similar and I think that that has massively increased in the last six to nine months, this is the partially the result of that and and you&#8217;re saying but I just wanted to put that out there of like I do think that they have a lot more compute though than they did when they were training the previous generation of models.</span></p><p><strong><span>00:16:24 Florian Brand:</span></strong><span> Yeah. like we like or just for for context two weeks ago I think LongCat released their model which they uh claim and I we know that it is very likely true uh is trained entirely on uh on Chinese chips. uh they didn&#8217;t specify publicly which ones but people speculate that it&#8217;s uh that it&#8217;s some uh Ascends from Huawei um and as the domestic production ramps up and you can they&#8217;re probably used most or they are used for for training but they are especially useful for inference which is a huge part of training as well.</span></p><p><span>So they probably use some mix of uh of Nvidia and other chips for the training part and then an increasingly larger part for the inference part during which during the stage is is really important. So I think their overall compute is increasing and also they don&#8217;t actually have a lot of users. So they don&#8217;t need to power 1 billion users like ChatGPT has to do, hundreds or thousands of enterprises like Anthropic has to do because they don&#8217;t have that magnitude of uh of of paying customers.</span></p><p><strong><span>00:17:41 Nathan Lambert:</span></strong><span> Yeah. And I think even those paying customers also, at least on the enterprise side, there&#8217;s just like there is company time and chatter when you&#8217;re supporting these things. Even if you&#8217;re like not a research, even if it&#8217;s not in your job, it like does change the attention of the company. And if SSI comes out with a good model, it&#8217;ll be the ultimate validation that distractions are are a problem, but that&#8217;s an aside that we we can wait on. I think the there&#8217;s also rumblings of the data and environments industry starting to appear there.</span></p><p><span>Do you remember any specific ones? Because when we were in China, it was kind of shocking how little they seem to utilize external data. So just a few months hearing a whole bunch of a month months after our trip we went in April and then just months later in July, we&#8217;re are hearing a few things of like new companies in China and them wanting to buy data and things. And that is uh like a funny timeline of how that changes.</span></p><p><strong><span>00:18:40 Florian Brand:</span></strong><span> And I would put error bars on what they actually told us.</span></p><p><strong><span>00:18:44 Nathan Lambert:</span></strong><span> And that cuz it&#8217;s like so close in time that I don&#8217;t know.</span></p><p><strong><span>00:18:49 Florian Brand:</span></strong><span> Yeah. That that that might that might be true. Uh but like those things are hard to to pinpoint. I but I would say it it seems like the buying of external data is becoming more of a factor. Um which will help the open models catch up to the closed ones if they just buy the same data maybe at a discount because um they buy the the data environments later. But it is it is a factor. How big of a factor like we don&#8217;t know. we don&#8217;t have any public insights and I doubt that we will get those insights uh from from anyone b uh really uh so that&#8217;s definitely one of the parts why um why we are able to to catch up or improve their their model scores.</span></p><p><strong><span>00:19:47 Nathan Lambert:</span></strong><span> Okay, roundup of other Chinese model providers. We&#8217;ve talked about Kimi, we talked about Zhipu / GLM. I think there will be more GLM models soon that are very good. They might call it like GLM 5.5. Um Qwen, we talked about their biggest model coming. Qwen&#8217;s biggest models I will say have tended to relative to the excellence of their small models not had the same like absolute ranking in performance which is a probably a cost of focus. I think it goes with a cloud companies. It&#8217;s it&#8217;s almost like it&#8217;s if you squint it&#8217;s almost like Google.</span></p><p><span>It&#8217;s like Qwen has Alibaba has so much opportunity here and the opportunity of getting developers associated with Alibaba Qwen with these small models is such a huge opportunity for their cloud that I think they&#8217;re succeeding wildly. But their big models have always not been as excellent as their small models. So I don&#8217;t expect their model to be as breakthrough as Kimi K3 or GLM 5.2. I expect it to be covered in the news as major open quite as the open bottle name in China drops giant bottle but I don&#8217;t think it will be as sustained as a um news story um DeepSeek you can go if chime in whatever</span></p><p><strong><span>00:21:01 Florian Brand:</span></strong><span> the the interesting thing is don&#8217;t know how how much you follow this but they are have or they have an endpoint which you can use for a preview version and they&#8217;ve updated this endpoint daily so they have some really fast iteration cycle because we the we progress in all these um Twitter um benchmarks. So a lot of these SVG things and three.js like all these visual generation tasks the model has been improving a lot over the last few days. So they have figured out some kind of fast feedback mechanism um which other companies have as well. Uh we we know this or cursor has a lot of blogs about this how they iterate really fast. Um but they seem to continuously upload new checkpoints and make them available.</span></p><p><strong><span>00:21:47 Nathan Lambert:</span></strong><span> Um but I agree. I&#8217;m guessing it&#8217;s like a time gated within their final RL run. It&#8217;s like still slightly improving at the end of their RL run and they&#8217;re just like checking the box.</span></p><p><strong><span>00:22:03 Nathan Lambert:</span></strong><span> Okay. Qwen DeepSeek V4 is supposed to come out a preview version. Um the thing about DeepSeek V4 I think is that the flash model is actually way more popular which is their smaller which seems to be an absolute workhorse for people. So that I think is the model to watch for them. I don&#8217;t expect V4 Pro to be a dramatic breakthrough. This is similar to anything like if Xiaomi were to release a new MiMo Pro model soon. I don&#8217;t expect it to be as big of a drop but it would probably be a very solid model. It&#8217;s just like it&#8217;s hard to know. They&#8217;re still a pretty new entrance. MiniMax, I think, is playing a different game. I don&#8217;t think MiniMax is chasing this um Kimi/GLM moonshot to AGI type vibe.</span></p><p><strong><span>00:22:46 Florian Brand:</span></strong><span> Oh, I would, I would disagree there.</span></p><p><strong><span>00:22:49 Nathan Lambert:</span></strong><span> You think, Do you think MiniMax is still in this?</span></p><p><strong><span>00:22:52 Florian Brand:</span></strong><span> Yeah, I I I I think they they are seeing the tension especially because they are a public company similar to GLM and if you look at the stock performance RIP those stocks in the last few days um it it it it make it seems to make a huge difference and the interesting part will be uh the license because they&#8217;ve changed the license a lot uh to be more and more restrictive and um if there&#8217;s now a change of heart again after the Xi, uh, speech.</span></p><p><span>Uh it will be interesting to see whether MiniMax goes back to completely open licenses. It&#8217;s also an interesting thing to see um which license will be the license for for K3 because they have said they will open source it but I don&#8217;t think they have done any commitments in terms of the actual license where you put on top.</span></p><p><strong><span>00:23:45 Nathan Lambert:</span></strong><span> Yeah. I mean that&#8217;s it&#8217;s super important is the thing. Yeah, we we&#8217;ll see. Um, Ling, Meituan, LongCat kind of similar, very strong models, probably getting a lot of value out of them internally. Aren&#8217;t don&#8217;t have the same developer breakthrough. Um, so what that&#8217;s like seven seven to eight Chinese labs. I might have forgotten some. And we can also talk about US labs. Aside um, Gemini 3.6 flash dropped. It looks fine. It&#8217;s like it&#8217;s like it&#8217;s it&#8217;s a tiny bump. It&#8217;s faster. It&#8217;s less of a yapper, but like doesn&#8217;t really matter. We&#8217;re going to stop we&#8217;ll stop sharing this. Um that&#8217;s that&#8217;s the amount of mention that Gemini gets for us.</span></p><p><span>But I do think it&#8217;s worth talking about the US ecosystem a bit. I think there are emerging players. Thinking machines released their first model. I&#8217;ve talked to some of them. they&#8217;re very on board for figuring out this how to make a fine-tunable model with Tinker and I think that&#8217;s a research area that I really really recommend for most of the open model builders. I think if you can get mind share there you will get massive adoption because it&#8217;s more about being fine-tunable for real tasks than it is about having that be best best numbers. Um, so this was their Inkling model which is a one trillion parameter which has like decent but not frontier scores.</span></p><p><span>I think kind of like DeepSeek V4 they&#8217;re going to they&#8217;re planning to release a smaller which is like a quarter of the size in total parameters which has really really good performance and if Inkling small preview comes out in a few weeks I do think that that will be a really used model. It&#8217;s a good size for kind of automating tasks and kind of domain specific tasks and might not be a like general agent type thing like Kimi and GLM 5.2 but I think that suits their business really well. Um I know that there&#8217;s some other the I would say like the smaller players in the US seem well like Arcee released their models earlier this year still chugging along. Poolside has started releasing some models.</span></p><p><span>They&#8217;ve gotten a few in the last few months and seem poised to release more models on top of that. So they&#8217;re really going Reflection is perpetually in the model coming soon camp and it really behooves them to get some models or some code or something out so that they can just start getting the developer flywheel going if they&#8217;re really committed to open source. It just takes a lot this it&#8217;s hard to get the models out. Like I talked to some people at Thinking Machines and it&#8217;s like kind of like oh that&#8217;s a lot of it&#8217;s a lot of work to actually do this I think. And um Nvidia chugging along. I think they&#8217;re at the stable player at this point. They&#8217;re keeping to release models. They&#8217;ll release more soon. They release a lot of data. I&#8217;m bullying them to try to get them to release Qwen style small models, which is like Gemma.</span></p><p><span>Gemma only has these like Qwen competitor models that are super popular. Um, the Gemma models are a little they&#8217;re all over the place in sizes or in architectures for the sizes and things like this, but the Gemma models are really really matching the Qwen models in terms of adoption. Um, I&#8217;m not sure they&#8217;re as easy to use for research, which could take a while. It could take multiple iterations. Like so much of language model research is now designed around small Qwen models and Qwen-based models that like it takes a while. Like people know how to use these models really well and if with the research results. So I hope Gemma keeps coming and can kind of compete in that niche. I don&#8217;t know any anyone that I missed here.</span></p><p><strong><span>00:27:22 Florian Brand:</span></strong><span> No, I think both are the big players. Uh it&#8217;s, it is becoming broader. Uh in terms of model creators like last year, did we have any release aside from Gemma 3 and um GPT-OSS?</span></p><p><strong><span>00:27:41 Nathan Lambert:</span></strong><span> was GPT-OSS 2 would go hard and obviously and obviously Nemotron as well. Um, oh, and I think Llama 4 at the start of the year, but uh, I don&#8217;t want that to be forgotten, but we are seeing like more players are are are now joining and turning out models at a really incredible rate.</span></p><p><strong><span>00:27:59 Florian Brand:</span></strong><span> like Poolside has been releasing three or four models in the last two or three months. Uh and they seem to have figured out some way to turn out models pretty consistently. Um and that&#8217;s also something we are seeing on the open source side as well. we are talking about GLM like I think their iterations uh times for the model releases are now between 1 or 2 months with each new iteration becoming better and better which closely resembles what the closed labs are doing like we get a new GPT we get a new Claude every uh 6 weeks or so these days uh so in terms of having uh good enough pipeline uh to release stronger and stronger models they have to or the open source ecosystem has really figured it out or seemingly figured it out.</span></p><p><strong><span>00:28:59 Nathan Lambert:</span></strong><span> Yeah, I agree. It&#8217;s it&#8217;s promising, but it is also so funny that like the US ecosystem started releasing some models and then then you have like Xi on the mic and these two models. It&#8217;s just like it&#8217;s so hard to catch up because it takes a lot of institutional expertise to train models that people actually use. And I think this is is what the American companies that are releasing models are now realizing is like these are not just benchmaxxed distilled IP theft models.</span></p><p><span>These are like genuinely good models that people are comparing to on their internal trading benchmarks and then like seeing how hard it is to beat them on measurable things. And I think that that is like I I&#8217;ve I&#8217;ve picked this sentiment up from a few people in the US trading models and it is just like there&#8217;s some I I think people should innovate on like size and fine-tunability and try to like use this potential market that is really close to home but also the pressures for every company is so high to release a model that you can claim as Frontier. I think investors expect that out of so many of these players that they&#8217;re kind of trying to do a a pretty hard thing and it&#8217;ll be interesting how the next year unfolds for the US China balance.</span></p><p><strong><span>00:30:25 Florian Brand:</span></strong><span> Yeah, I think or in general I and a lot of other people have talked about the general ecosystem and that&#8217;s also something you&#8217;ve talked about at the very beginning. I think we are seeing more and more of a split between the capabilities of models that is good enough for a lot of tasks like uh for a lot of coding tasks the current frontier models both open and closed are good enough. um improvements feel less and less uh important here.</span></p><p><span>But if we look at the frontiers frontier, so finding new math proofs, finding uh new uh cures, finding new drugs, and inventing new things, that seems to be a whole different beast and probably will be dominated by the very frontier for quite a long time. The big question then becomes how much does that matter uh in terms of the addressable market and also how much of a focus will this be. I think, or my general base case is that we are seeing the frontier close down more and more. We have seen this with Mythos for cyber security GPT... or for biotech that those models won&#8217;t be accessible for everyone um and maybe not even external partners if we consider the reports that Anthropic is now spawning or or creating some internal labs to develop drugs.</span></p><p><span>Um so the very frontier is inaccessible for everyone and then the near frontier capabilities is becoming more and more commoditized um which has a lot of different implications especially if you think about things like uh cyber security. There was that report from Hugging Face two or three days ago that they had some agent trying to to hack their system. um and they tried to analyze it with GPT and with Claude but were unable to because all the guardrails blocked them. So they had to use GLM, a lesser capable model, but it had no guardrails for this kind of defensive action. And they had to use a worse model to defend themselves or to analyze the data, which is a horrible state to be in that we have US-based companies now relying on lesser models because the closed frontier is inaccessible to them.</span></p><p><strong><span>00:33:10 Nathan Lambert:</span></strong><span> Yeah. And I think this is actually one of the best arguments for not doing anything. It&#8217;s like if the rest of the world has access to these open models and we ban them for the companies in the US to use and it&#8217;s just like a growing disparity between US companies ability to defend and the attackers all over the world in terms of cyber and we could debate like how much of an immediate risks the cyber stuff is at the current capability levels but if you&#8217;re setting it up structurally so that the defenders get don&#8217;t get better over time and the attackers can like that seems like when the Why would cyber risk become more real? And that would to be very clear that would be if you ban the best Chinese openweight models from being used at companies in the US.</span></p><p><span>And this ban would likely be a kind of shadow ban, which is the threat of legal threat of legal action or punishment without it being clear on exactly what the pathway to do it is. And there are a lot of talks about this right now. I don&#8217;t like like I don&#8217;t know if we&#8217;re going to have a ton to say about this, but it&#8217;s clear that DC is flirting with different ways of restricting the best Chinese openweight models in the US. This is I think downstream of some fear-mongering. We&#8217;ll transition into the distillation question too. It&#8217;s like all these things from the primary AI media narrative in the US that is pointing towards Chinese models as stealing IP or being dangerous or being affiliated with the Chinese government, an authoritarian government.</span></p><p><span>And it&#8217;s like all these things are leading up to this moment of interest in taking action on AI and then not really knowing where to do it. So potentially taking a crude instrument to the like quote unquote enemy and we could transition into distillation. I think there&#8217;s a lot of discussion on it. Most recently Ben Thompson finally chimed in on distillation. I think Ben is probably one of the is probably the highest read blog in tech (</span><a href="https://stratechery.com/"><span>Stratechery</span></a><span>). I think that the the debate let&#8217;s see where do we even start the debate. The core question is like how much does distillation help and what should you do about it? I&#8217;ve been of the opinion that distillation has becoming less and less impactful over time as the Chinese models get closer to the frontier and the trading regime shifts to RL. The way that distillation tends to happen is that the Chinese labs hack the APIs. Hack is like maybe a strong word, but they jailbreak the APIs of Claude and GPT to extract the reasoning tokens.</span></p><p><span>When you have the reasoning tokens with the tool calls, that is perfect SFT data and or mid-training data to train the base model with to seed some agentic behaviors in an important domain. And now after that the core part of post-training is to do large-scale RL in agentic domains to so like push the frontier and everything that they&#8217;re doing today and RL is only becoming more prevalent with this as SFT becomes less prevalent in previous generations you could get very close to the frontier just by scaling up SFT and that would be what really impactful if you could say take a million agentic rollouts from Claude or GPT have that be your SFT set and train on it.</span></p><p><span>I think in previous years that would have done a lot more to get you to the frontier. What Ben Thompson has said which made me really annoyed is that he very strongly proclaimed that distillation is getting more impactful as you do RL. He did this in his article who&#8217;s afraid of Chinese models. We can link it below. It&#8217;s a public one. And then he was also on his own podcast tour. He has also podcast as well saying the same things. And I think it&#8217;s really important to say that distillation during the RL stage is a lot harder.</span></p><p><span>What he said was that the kind of grading models that can be used during RL, which is essentially you can have a model check over the agentic trajectory of a roll out and grade different parts on if it completed the reward, what actions it took. And he&#8217;s insinuating that the Chinese labs are using Fable and GPT 5.6 and the strongest models to actually do this supervision in RL. The problem is that big RL runs are millions and millions of rollouts. I think Thinking Machines blog post had like 20 to 40 million or something for their final RL run. So to do this on an API like Fable or GPT 5.6 would be insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift versus using your own tailored greater model or and many things like this.</span></p><p><span>And so I just think the argument that distillation is helping more because RL is becoming more prevalent is not grounded in literature that we have today. This is tough for me because Ben&#8217;s article also concludes that we should like make terms of service disallowing distillation illegal, which I kind I like want to support his radical conclusion to make distillation legal for US companies, but I can&#8217;t support any conclusion that I think is on um infactual mis like misguided information. So, I&#8217;m also a fan of Ben. If you&#8217;re a fan of Ben and could also nudge him on this, I would you really should because there&#8217;s probably one more podcast. What is he going to record it on? Like when does he record Sharp Tech? Thursday.</span></p><p><span>We We got to get on and get him to correct the record because I I don&#8217;t know. I I find it so annoying that the most prominent voice in tech is trying to be an ally for our point of view on distillation is that we should do nothing. Um but like it&#8217;s hard. It&#8217;s like he has such wide reach that this is now going to be the status quo that we have to debunk which I guess it&#8217;s a better status quo than I don&#8217;t know actually no it&#8217;s not helpful because he&#8217;s saying that distillation is more important which means the people who are afraid about that are going to use that as a data point to say that we should take action even if they don&#8217;t because they probably won&#8217;t agree with his conclusions. I don&#8217;t know. That was my rant. Ben, you&#8217;re wrong.</span></p><p><strong><span>00:39:16 Florian Brand:</span></strong><span> Yeah, I-I do think it is important to to differentiate these phases. Um and especially like there is no doubt that it is used during the SFT stage which is the first stage of or one of the stages for post-training and that&#8217;s also where the model picks up its manners like that&#8217;s why the models say oh I am Claude because they learn this during the SFT stage that&#8217;s where uh this personality is formed but the strong capabilities come during the RL stage which is where the money is spent which is where you need to have a fast enough judge which in the best case just runs in at the same GPUs or very close to your GPUs with smallish or with a fast enough model so you can uh are not bottlenecked by this.</span></p><p><span>Um and it in terms of impact it is also very hard to say how much impact or how much of a boost the better model SFT data gives you versus a lesser model. Um so if you are able to to have 10 million tokens from the latest Claude model versus two generations behind open model how much of a boost that really gives you if you keep the stage right uh the same and the pre-training stage the same is an open question which I don&#8217;t think we will see answered in a paper because then you have to showcase your uh SFT and your jailbreaking capabilities</span></p><p><strong><span>00:40:48 Nathan Lambert:</span></strong><span> but I I wanted to double down on this like there&#8217;s been a good amount of literature on generating SFT reasoning traces whether it&#8217;s the most prominent ones have been opens line of work they did </span><a href="https://www.open-thoughts.ai/"><span>Open Thoughts 3</span></a><span> and Open Thoughts Agent have kind of been the foundational like scaling reasoning SFT works in the last few years and whenever somebody revisits this question they have not found the answer that the strongest model on performance in your domain is the best teacher for SFT people have try I&#8217;ve tried many people have tried the idea is so simple is like the state-of-the-art open SFT data set is built on QwQ-32B like an ancient reasoning model or something.</span></p><p><span>Why can we not just generate completions from GLM 5.2 do SFT on it and improve the model? We don&#8217;t know. It&#8217;s like the research so many people have tried and it is not an answered research question. There might be something like the base model the mid-training is too close to Qwen. So therefore it&#8217;s like hard to break. You have to redo the mid training. I think you have to redo the mid-training for reasoning. I think reasoning mid-training and reasoning SFT are so closely intertwined. It almost doesn&#8217;t make sense to have different words for them. That could be the issue. But the literature doesn&#8217;t even know how to ext like if I had a magical API that gave me reasoning traces from Claude/Gemini. I actually don&#8217;t know if me like fine-tuning an OLMo model on that would make OLMo smarter.</span></p><p><span>It&#8217;s one of the most wild unanswered research questions. And this is just makes the distillation thing so funny where it&#8217;s like yes the Chinese labs I think are using strong models like Opus for some SFT data but they&#8217;re also innovating. I was like I would love to them to tell us how to make this freaking work. And I think it&#8217;s the the paradigm I think is like open AI and anthropic find a niche domain that they do so well at and then the Chinese labs can get some samples there to kind of bootstrap their data engine and that&#8217;s where you will gain you will gain a few months on a specific domain.</span></p><p><span>But a hill climbing on these core domains like math and code and like Terminal-Bench like they&#8217;re just doing the same thing which is like so hard to generate prompts which are problems with environments that are hard for the current models and provide real nonreward hacking um learning behavior. And like that is what frontier data research looks like right now. And it is like it&#8217;s hard to generate these hard problems. And I&#8217;m sure the Chinese labs are doing the same the same things. And I don&#8217;t I don&#8217;t know. That&#8217;s that&#8217;s my rant. I&#8217;m kind of lost the context of our conversation.</span></p><p><strong><span>00:43:23 Florian Brand:</span></strong><span> No, no, I would I would agree. Or to to to recap, yeah, SFT or distillation has some effect. Yeah, it gives them a boost, but not that much uh as people would like or or seem to think it gives.</span></p><p><strong><span>00:43:39 Nathan Lambert:</span></strong><span> I think that&#8217;s that&#8217;s a good good summary of the of the of the conversation. It also becomes kind of tiresome because it says uh that open all all these open models are just good because they are distilling. um which definitely isn&#8217;t the case cuz if if it were the case, everyone would would be easily able to catch up to a GLM or to a K3 um by using its data for distillation. But we have not or we won&#8217;t see this from SFT alone.</span></p><p><strong><span>00:44:12 Florian Brand:</span></strong><span> Yeah, I agree. Do you have any predictions or or more topics you want to get to?</span></p><p><strong><span>00:44:18 Nathan Lambert:</span></strong><span> Um in terms of predictions, I think we are or I I revisited uh ours from from last year and it basically said everything will continue uh like it did uh the previous year. Uh we predicted that we will see bigger models uh up over two trillion parameters which it did and I don&#8217;t think we will see a much bigger explosion in terms of model size this year. we might see something or some model a bit bigger than three trillion parameters uh total but I don&#8217;t expect a five or 10 trillion parameter model and we open this year that would really surprise me um then list from last year can we redo this we don&#8217;t have to do the whole thing</span></p><p><strong><span>00:45:03 Florian Brand:</span></strong><span> oh sure</span></p><p><strong><span>00:45:03 Nathan Lambert:</span></strong><span> this is where we were at the end of 2025 who do you put in frontier now well it is Kimi and it is Zhipu DeepSeek is kind of a hard nut these days. Like I think they would be in close competitors. So I would put DeepSeek and Qwen the one as close competitors with Kimi and Zhipu as Frontier. Do you think anyone else would deserve close competitor? Cuz after that noteworthy and below like there&#8217;s so many.</span></p><p><strong><span>00:45:38 Florian Brand:</span></strong><span> I I think we will see a surprise from MiniMax by end of the year. I think we will see a big model which like a really big model not uh M3 size but trillion parameters plus which will surprise us in terms of uh the outputs of MiniMax compared to before uh so I would still put them at close competitors by the end of the year</span></p><p><strong><span>00:45:54 Nathan Lambert:</span></strong><span> do you think any US companies will be in the closed competitors by end of the year Nemotron I don&#8217;t think I would put there Thinking Machines closer especially if the smaller model really breaks through. But I don&#8217;t think I would put them there yet. Reflection is supposedly like only wants to release if they have a model that&#8217;s frontier. But then the question is will we get it? Like do we think that any US companies will get into this what is roughly like our top five by the end of the year? So the top five are the same but reshuffled.</span></p><p><strong><span>00:46:36 Florian Brand:</span></strong><span> I would say it is possible uh that they are really close. Um it it also depends on what we think matters for closeness. Like I think uh Nemotron and um uh Thinking Machines will release models which act as really good base to be fine-tuned for your domain which doesn&#8217;t mean they are usable like a frontier model but they have so much utility uh that I would put them into close competitors because you would just need to find your data and uh to push the model into the right direction.</span></p><p><strong><span>00:47:12 Nathan Lambert:</span></strong><span> Um as a I was going to think that we would make this a group of six with a US company by then like if we do this in late November I would guess that a US company pro most likely Nvidia thinky or Reflection mo does stuff that gets us to say that there is an American company in this like top cluster which would be a first time for a while.</span></p><p><strong><span>00:47:43 Florian Brand:</span></strong><span> Yeah, I I I think that is realistic. My my one wild card is Tencent, which I think we might see something by end of the year. Uh they got some new leadership. Uh they released their Hunyuan model under Apache this time. Wait, so Tencent always had these custom licenses which disallowed anyone in the UK and South Korea and the entirety of the EU to to use their their model and also had acceptance use policy and so on. Um, and with Hunyuan and their new leadership, they got a really competent model at 250ish billion parameters. Um and I think by end of the year we might see a big model release which will surprise the people not following the ecosystem.</span></p><p><strong><span>00:48:36 Nathan Lambert:</span></strong><span> Yeah, I I am also sure we will be in for some surprises. This is always the thing with AI and especially open models. It&#8217;s very very unpredictable. Okay, I I think this is a good place to stop. We probably should really do this quarterly. It&#8217;s not that hard and people will enjoy it. Um, but good to see you and we&#8217;ll talk soon. Hopefully in person soon.</span></p><p><strong><span>00:49:02 Florian Brand:</span></strong><span> Peace.</span></p>]]></content:encoded></item><item><title><![CDATA[Kimi K3: The open-weights escalation]]></title><description><![CDATA[The global implications on the AI ecosystem.]]></description><link>https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation</link><guid isPermaLink="false">https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 20 Jul 2026 15:48:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b7967fef-8d08-4180-a3fd-8c2239bc9619_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>On Thursday July 16th, Moonshot AI released their latest flagship model </span><a href="https://www.kimi.com/blog/kimi-k3"><span>Kimi K3</span></a><span>. K3 is a 2.8T parameter MoE model which will have its weights released on July 27th. Much of this article follows as a reflection on the state of the ecosystem, under the assumption that Moonshot keeps their promise of the weights release date. This is a more extreme view of the equilibrium, and many of the results end up in a middle ground if the state of affairs is that China has similarly powerful, but closed models (i.e. K3 is never released). </span></p><p><span>The key fact is that either the open-to-closed or American-to-Chinese model performance gap has been reduced from the debated 6-9 months to something shorter, say 3-5 months.</span></p><p><span>From the release materials, it is clear that K3 is a true frontier model. It will be the closest open models have been to the frontier since </span><a href="https://www.interconnects.ai/p/deepseek-r1-recipe-for-o1"><span>DeepSeek R1</span></a><span>. DeepSeek R1 was a different story. This was a Chinese lab being extremely quick to pivot to reasoning models and release one faster than many American companies. Kimi K3 an example of a Chinese lab executing on scaling the known areas: data, algorithms, architecture, tools, environments, etc.  </span></p><p><span>Kimi K3 comes in at </span><a href="https://x.com/ValsAI/status/2077834614668988586"><span>#2 overall on the Vals AI index</span></a><span>, </span><a href="https://x.com/ArtificialAnlys/status/2077832874183860404"><span>#3 overall on Artificial Analysis&#8217;s Intelligence Index</span></a><span> (only beaten by Claude Fable and GPT-5.6 Sol Max while being </span><a href="https://x.com/ArtificialAnlys/status/2077832885021835289?s=20"><span>cheaper</span></a><span>), </span><a href="https://x.com/arena/status/2077824029126504525"><span>#1 overall in Frontend Code Arena</span></a><span>, and </span><a href="https://x.com/cramforce/status/2078574147333152957?s=46"><span>more impressive results</span></a><span>.  Moonshot AI is going toe to toe with Anthropic and OpenAI with far, far fewer resources.</span></p><p><span>It is clearly the strongest open model ever released. It should be clear looking at this model that if adversarial distillation from the closed frontier models in the U.S. contributed, it is at most to a relatively small degree. AI observers who followed the </span><a href="https://www.interconnects.ai/p/the-distillation-panic"><span>distillation panic</span></a><span> and came away with the wrong conclusion that Chinese AI labs are only producing good models due to IP theft are in for an awakening &#8211; that Chinese companies are extremely good at building models in the same way the leading American companies are. Moonshot AI is solving many of the same problems that folks at OpenAI or Anthropic are solving. I&#8217;m confident there will be more distillation discussion, and </span><a href="https://www.interconnects.ai/p/the-distillation-panic"><span>pressure</span></a><span>, but the evidence is now out that Chinese companies can do </span><a href="https://www.interconnects.ai/p/how-much-does-distillation-really"><span>more than just fast following</span></a><span>.</span></p><p><span>Meeting some of the core Kimi team on </span><a href="https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs"><span>my trip to China</span></a><span>, it was clear to me that they had incredible culture, some would say aura, and a freedom to express it &#8211; within the constraints of a GPU-limited environment. Where building models is so much of a scaling game, much of the ability to build a good model still comes down individual execution, motivation, and expression. Having visited them, this result is less surprising. Having visited many AI companies, very few have a culture that you can immediately pick up like this.</span></p><p><span>At the same time, China&#8217;s AI adoption trends started later than those in the U.S. So, while all the Chinese labs have way less compute than their counterparts in the U.S., more of it can certainly go to training. When I joked around about how much compute an average researcher at OpenAI could have &#8211; say a few thousand H100 equivalent machines &#8211; the researchers at Kimi were shocked. The org chart and approach to building the Kimi models surely reflect this, but it is difficult to tease out what this looks like without substantial proprietary information.</span></p><p><span>The state of affairs on peak model performance is roughly as </span><a href="https://x.com/finbarrtimbers/status/2077816750574539161"><span>follows</span></a><span>:</span></p><ol><li><p><span>Anthropic &#8211; Claude Fable 5</span></p></li><li><p><span>OpenAI &#8211; GPT 5.6 Sol</span></p></li><li><p><span>Moonshot AI &#8211; Kimi K3 (open weights*)</span></p></li><li><p><span>SpaceXAI &#8211; Grok 4.5</span></p></li><li><p><span>Zhipu (Z.ai) &#8211; GLM 5.2 (open weights)</span></p></li><li><p><span>Meta &#8211; Muse Spark 1.1</span></p></li><li><p><span>DeepMind &#8211; Gemini Flash 3.5</span></p></li><li><p><span>Alibaba &#8211; Qwen 3.7 Max (3.8 announced, also to be open-weights, when writing)</span></p></li></ol><p><span>It is astonishing to see DeepMind, and some of the other American giants this low. In many ways, the X AI team deserves more credit. A visual summary from Artificial Analysis is below:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NQJl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NQJl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg 424w, https://substackcdn.com/image/fetch/$s_!NQJl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg 848w, https://substackcdn.com/image/fetch/$s_!NQJl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!NQJl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NQJl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg" width="1456" height="612" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:612,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:431327,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.interconnects.ai/i/207699639?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NQJl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg 424w, https://substackcdn.com/image/fetch/$s_!NQJl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg 848w, https://substackcdn.com/image/fetch/$s_!NQJl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!NQJl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>This release and other recent events have caused a major change in direction for the most likely outcomes in the balance between open and closed models. I&#8217;ll unpack them individually.</span></p><p><span>In many ways, it feels like the start of a new era. An era with much more competition, but also a much higher need for coordination, as we rollout incredibly powerful technologies around the world.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><span>1. China&#8217;s recommits to open-source AI &#8211; showing a different read on near-term risks</span></h2><p><span>Many people started following China&#8217;s AI scene relatively recently, so they can reach the conclusion that releasing models openly is their core strategy. In fact, I think most labs have a core strategy far closer to Anthropic or OpenAI &#8211; build the best intelligence possible. Having followed and engaged with the Chinese labs for years now, the best explanation for their original turn to releasing their models openly is practicality. They needed to release the models openly to get adoption, attention, and feedback </span>(especially in the high-value, Bay Area market)<span>.</span></p><p><span>For a long time, there had been </span><a href="https://www.chinatalk.media/p/chinas-new-ai-plan"><span>very limited policy in China</span></a><span> explaining the role of open-source AI, and what could be the &#8220;country-level strategy.&#8221; To my knowledge, no senior leaders had commented on open-source AI publicly. This changed this week too, as Xi Jinping gave a keynote address at the World AI Conference (WAIC), and very directly </span><a href="https://mattsheehan.substack.com/p/xi-jinpings-big-ai-speech-annotated"><span>committed the future of China&#8217;s AI ecosystem to open-source and global diffusion</span></a><span>. This commitment to the status quo, the same week as the announcement of the strongest open-weight model to date, is a clear mark in the early history of modern AI.</span></p><p><span>This comes during a time period where many potential paths forward have been discussed for the Chinese AI industry &#8211; Will they stay open? Can they keep up with the American labs in scaling? Is there a growing revenue market in China? With these, the focus has been on China&#8217;s risk tolerance, the companies&#8217; ability to monetize, and any closely related reason for a company to stop releasing their </span><em><span>best</span></em><span> models openly.</span></p><p><span>In tying Xi&#8217;s commitment in time to a very strong model, China has implicitly commented on its risk tolerance with respect to releasing open-weight models. For the time being, it is a read into the perceived risks of topics like strong cybersecurity capabilities (or bio-dangers) within the Chinese system.</span></p><p><span>The simplest explanation is that China&#8217;s government is definitely following potential risks from the models closely &#8211; likely with more technical scope than the US government&#8217;s vibe regulation &#8211; and would take action if it measured risk. The simple explanation is that they do not find current frontier models to have meaningful risk.</span></p><p><span>At the same time, China&#8217;s economic decision makers think having AI adoption is good, so they can make profits on the industry later &#8211; after growing distribution (as China has done for cars, solar, advanced manufacturing, and many areas in recent history).</span></p><p><span>These can seem somewhat shocking, in an American AI media landscape that has gone through months of hype and fearmongering over the Claude Mythos model. This surprise should be excellent grounding &#8211; the world does not have a unanimous agreement with the narratives about AI that we hear most in the U.S.</span></p><h2><span>2. Open models as the economic Achilles heel of frontier labs</span></h2><p><span>Many of the narrators guiding the discussion on AI have clear incentives to depress the perceived capabilities of the best, open AI models. Dean Ball &#8211; who is personally supportive of open models, but now works at OpenAI &#8211; had a widely commented on </span><a href="https://x.com/deanwball/status/2078133895766114412"><span>post</span></a><span> with some reflections on Kimi, where he said the following on open models. It is important to understand the statement, as it focuses the role of open models in the economic side of the AI buildout. Dean says:</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><blockquote><ol start="3"><li><p><span>Open-weight models are inherently decelerationist, and I&#8217;m continually surprised to see the so-called &#8220;accelerationists&#8221; so excited about open-weight models.</span></p></li></ol></blockquote><p><span>Explaining why open-weight models are </span><em><span>a form of decelerationism</span></em><span> is important to understanding the coming world order. He is right.</span></p><p><span>Open models are decelerationist economically for the frontier labs, which will slow the net investment and capex rollout for AI. This is due to the fact that strong open-weight AI models massively reduce the margin potential for the closed labs. This has two effects. First, the AI labs have fewer profits to re-invest into future models. Second, the market sees the terminal value of these companies as being lower, so they will kneecap future fundraising rounds. These together will slow timelines to the most transformative AI models, but I do not see them as strong enough effects to stop OpenAI and Anthropic from being a few of the top valued companies in the world.</span></p><p><span>These, to me, are a net good for society. As open-weight models are </span><a href="https://x.com/natolambert/status/2078164535262040448"><span>accelerationist</span></a><span> for </span><em><span>AI diffusion across the economy</span></em><span> by having the entry price for intelligence at a certain level of performance be lower. Open models also encourage customization. The thing is that this type of diffusion is by its nature far slower than the frontier AI labs products, who sell tools used directly by developers. The potential for open models is for nearly every business to use them to craft domain-specific agents. This economic diffusion takes an extremely long time! I&#8217;ve described this as </span><a href="https://www.interconnects.ai/p/open-and-closed-models-are-on-different"><span>open-weight models being on a much slower starting, but potentially bigger exponential</span></a><span>. The problem is, if closed models get too far ahead in raw capabilities, this ability to customize can be moot.</span></p><p><span>The combination of increased diffusion and decreased concentration of power in the AI labs I see to be very positive for the AI transition. It gives us more time to figure out the hard problems of new capabilities and lets more stakeholders impact the story &#8211; any one company is very likely to have issues with controlling the world&#8217;s most important technology safely. </span></p><p><span>It is, of course, important to me in this world for the best models to still be made by the U.S. companies, which will allow the US to control the trajectory of the technology and its values. I also expect this to be the case, as the U.S. has larger capital markets that are willing to invest in AI (and a growing share of profits), but it is not a given.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>3. China&#8217;s efficiency advantage</span></h2><p><span>Kimi&#8217;s launch blog has some technical details that confirm the sort of improvements that are supplying the consistent model improvements we feel. To select one:</span></p><blockquote><p><span>Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth. We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework. </span><strong><span>Together with refined training and data recipes, these structural changes yield an approximate 2.5&#215; improvement in overall scaling efficiency compared to Kimi K2</span></strong><span>, allowing the model to convert compute into intelligence more effectively.</span></p></blockquote><p><span>Training efficiency really adds up. They will result in continued, incredible steps for the models.</span></p><p><span>As an aside, tracing the path of this particular innovation through the ecosystem is an interesting example. Kimi Delta Attention (KDA) was introduced in the Kimi Linear </span><a href="https://arxiv.org/abs/2510.26692"><span>paper</span></a><span>, which is similar to the Gated DeltaNet used for </span><a href="https://arxiv.org/abs/2604.03444"><span>Olmo Hybrid</span></a><span> (my last Olmo model while at Ai2). Qwen&#8217;s latest models switched to a related architecture and the recent Nemotron models also are hybrid (but still closer to Mamba than Gated DeltaNet). It&#8217;s awesome to see new architecture ideas like these, which were heavily progressed by academia, get so quickly translated into frontier-scale models. </span><a href="https://arxiv.org/abs/2412.06464"><span>Gated Delta Networks</span></a><span> were introduced in late 2024, building on ideas from Mamba. By mid 2026, they&#8217;re in frontier models.</span></p><p><span>I chose to focus on this example, partially because the Kimi team put a cool number to innovations between models, but primarily to give space to a broader discussion of China&#8217;s resource efficiency.</span></p><p><span>It is becoming clear that the Chinese labs are far more capital efficient. In a world where scaling laws dictate that intelligence is proportional to effective capital &#8211; which buys compute, data, &amp; talent &#8211; that may be the greatest strength your AI industry could ever have. There are many possible explanations for why this is the case, such as Chinese researchers being paid less while being more effective at LLM research puzzles, but we will probably never get such specific reasons.</span></p><p><span>Since writing my </span><a href="https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs"><span>notes</span></a><span> on China, I&#8217;m hearing more about an emerging data industry in China (far behind the billion dollar budgets of Anthropic for data) and that Chinese labs have access to meaningful training compute (by skirting export controls). Chinese companies </span><em><span>do not</span></em><span> have the same inference demand (until recently, as Moonshot AI had to </span><a href="https://x.com/Kimi_Moonshot/status/2078855608565207130"><span>pause new subscriptions</span></a><span> for access to their K3 model - while the API is still live), so much more of their compute could go to training. These areas impinge heavily on the truth of the ability of the labs, but we have very limited measurement into them.</span></p><p><span>The facts on the ground are that these Chinese labs have raised orders of magnitude less capital than any slice of the American AI ecosystem. The most direct comparisons are to OpenAI and Anthropic, who have slightly better public models. Others, such as Google and Meta have the largest cash flows in the history of business, and are behind on building models. As for American neolabs, the picture is even more competitive &#8211; Thinking Machines released their first model recently, </span><a href="https://thinkingmachines.ai/news/introducing-inkling/"><span>Inkling</span></a><span>, which is strong but not in the same class as Kimi K3.</span></p><p><span>These American companies with more resources could still catch up, but you need to strongly weigh the public measurements we have of model quality and not resort to hope &#8211; which often reflects a bias. If the Chinese labs do have a latent advantage, they could continue to </span>utilize that to build even stronger models<span> than </span><em><span>all</span></em><span> the competitors! Many outcomes are plausible and K3 should increase most people&#8217;s probability that China can outright lead in AI capabilities in the near future on the back of more efficient training efforts &#8211; even if it&#8217;s not your most likely predicted outcome. </span></p><p><span>A big contributor to the capital efficiency is likely in the approach, where American labs are spending meaningful energy in pushing the frontier in dramatic, big steps, and the Chinese labs are more focused on catching up &#8212; this catch-up is cheaper. Just as the student model can outperform the teacher in distillation generally (not limited to the </span><em><span>adversarial </span></em><span>distillation of the Chinese labs), an approach of &#8220;trying to catch up&#8221; rather than &#8220;invent the next paradigm&#8221; could lead to stronger models.</span></p><h2><span>4. A growing ecosystem of frontier, open models</span></h2><p><span>The weekend after the Kimi K3 release, while writing this and discussing the events broadly, Alibaba </span><a href="https://x.com/Alibaba_Qwen/status/2078759124914098291"><span>announced</span></a><span> that a 2.4 trillion parameter Qwen 3.8 model is coming soon with open-weights. Historically, Alibaba has kept their largest models as API-only offerings via their cloud business, so this is another big vibe shift opening the doors to the next chapter of the open model economy. Even if the model is behind Kimi K3 on benchmarks, it signifies that Chinese companies may not only be maintaining the status quo for their open model strategy, but leaning further into it.</span></p><p><span>If this Qwen 3.8 model releases soon, i.e. before the next Gemini model, it could push Google to the 8th position on the leaderboard of labs with the smartest models &#8211; a list that China has been climbing. There are other rumors of more strong Chinese models soon, with DeepSeek V4 expected to graduate out of it&#8217;s &#8220;preview&#8221; version.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation/comments"><span>Leave a comment</span></a></p><h2><span>5. The very beginning of a long story of frontier open-weight policy</span></h2><p><span>I think that if Claude Mythos was released as an open-weight model today, the negative outcomes would be relatively minor. This is a somewhat challenging opinion to hold, as we have very limited public cybersecurity evaluations and it is a complicated ecosystem (and because I trust many people at Anthropic). I still stand by it. The risks have been over-hyped.</span></p><p><span>The problem is that this will not always be the case for the strongest AI models. Far stronger models are coming &#8212; and with them increased risks &#8212; so it is an incredibly safe equilibrium for the best models to be accessed </span>in a controlled, closed manner several months ahead<span> of similar open-weight models. </span></p><p><span>Open-weight models which are very controllable by the user will always be coming &#8212; you cannot effectively ban digital products, especially from bad actors &#8212; as AI training has proven globally accessible longer than many analysts expected.</span></p><p><span>Still, </span>as I write this<span>, the government </span><a href="https://www.interconnects.ai/p/6-months-to-live-for-open-models"><span>continues to flirt with more </span>measures aimed at restricting open-weight model<span>s</span></a><span> in the U.S. The latest is from </span><a href="https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi"><span>Axios</span></a><span>:</span></p><blockquote><p><span>Behind the scenes: The Commerce Department last year considered adding multiple Chinese AI labs to its &#8220;Entity List,&#8221; which would effectively cut off U.S. access without a license, a source close to the administration told Axios.</span></p><ul><li><p><span>The National Security Agency and White House Office of the National Cyber Director also considered putting out an advisory on Chinese AI lab threats last year, practically discouraging U.S. companies from using their tech, the source said.</span></p></li><li><p><span>The White House considered implementing an executive order saying U.S. companies could only host Chinese models if they could guarantee security and take liability if it were breached, the source added.</span></p></li><li><p><span>Commerce last summer also circulated draft rules within the administration leveraging its authorities to secure domestic supply chains to target Chinese open-source models, another source close to the administration said.</span></p></li></ul></blockquote><p>This would leave the U.S. in a very asymmetric state where the best models in the U.S. have guardrails on cybersecurity tasks, but global actors have access to great Chinese open-weight models to probe our defenses. This is one of many examples where banning open-weight models is not only harms the free markets of AI but also makes the ecosystem less safe in the short-term. There are other very bad outcomes, such as slowing the diffusion of AI applications and AI research, as I discussed above. </p><p><span>These equilibriums are very hard to maintain, especially as AI tools accelerate progress in the models, but it is important to maintain this status quo between open and closed. </span>Having a model that is truly alone at the frontier<span> in capabilities &#8212; something like Mythos when it was announced &#8212; </span>also be open-weight poses <span>serious risks as we go into the unknown of capabilities. Models are going to progress very fast and it is increasingly hard to measure their total capabilities. </span></p><p><span>We are then stuck in a world where we are trying to thread the needle on open models. It&#8217;s reasonable to not want something so powerful to be diffused globally in an instant, but meanwhile the makers of the models are </span>incentivized to hype their capabilities<span>, and their competitors are incentivized to hype their risks. It all comes down to careful measurement and proactive hardening of society to risk vectors.</span></p><p><span>This careening train of policy debates, model releases, and raucous reactions is only going to continue from today. We&#8217;ve been on a train of rapid progress, where all the key ideas of how AI should play out are tested, </span>since the release of Claude Opus 4.5 last December, which sent us<span> down the agentic pathway. The key to making good decisions here is evaluation capabilities</span>, independent of the companies<span> with the largest financial stakes. One of many actions needed then, as we enter the </span><a href="https://www.interconnects.ai/p/welcome-to-the-agi-era-of-ai-governance"><span>AGI era of AI governance</span></a><span>, is an Operation Warp Speed style approach of bootstrapping state capacity (and other independent actors) that can evaluate models accurately, and study emerging risks.</span></p><h2>Conclusion: The wake-up call</h2><p>Open-weight models, by accelerating the diffusion of capabilities, are a massive escalation in the good and the potential bad of AI. For now, the bad side of frontier language models has been largely hypothetical, but that will not always remain the case. </p><p>Having open weight models be slightly behind the closed frontier is our natural buffer to mitigate the risks. The key point is that we must collectively act to mitigate potential harms as they appear, and whether open-weight models are 3 or 6 or 9 months behind, that is still a very short timeline. If we regulate open-weight models heavy-handedly, I suspect much of the world will be lulled into thinking we no longer need to act. All we would&#8217;ve done is slightly delayed the inevitable &#8212; open models will continue to cross all the key capability thresholds eventually and regardless of legality.</p><p>Understanding and benefiting from this open-closed dance must be a collective action from the AI community across all sectors of power and influence over the coming years.</p><p>Kimi K3 is a watershed moment because frontier open-weight models are now real. Many hypotheses will be tested on where risks of open-weight models truly land &#8212; I suspect it&#8217;ll be narrower than many expect, and many risks of AI will still be proliferated by closed and &#8220;safer&#8221; APIs. The evaluation of these risks will evolve in time with an acceleration of AI&#8217;s integration in our economy. We cannot get one without the other, and we will continue to get both.</p><p><span>With this, 2025 was when open models started to be taken more seriously &#8212; especially when </span><a href="https://atomproject.ai/"><span>China leaped ahead with such a clear lead</span></a><span> &#8212; as people realized that it would not be a unipolar world, with only American, closed AI labs determining the trajectory. 2026 is when those previously discussed, potential risks and accelerations due to truly frontier, open-weight models landed. </span></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>A large portion of the response was for the fourth bullet, which compared the inevitable outcome of open models to AI communism, which I think missed the mark. Specifically, the use of the word communism without explanation caused much of the blowback.</p></div></div>]]></content:encoded></item><item><title><![CDATA[6 months to live for open models]]></title><description><![CDATA[The most serious test to date of open source AI&#8217;s viability is happening right now.]]></description><link>https://www.interconnects.ai/p/6-months-to-live-for-open-models</link><guid isPermaLink="false">https://www.interconnects.ai/p/6-months-to-live-for-open-models</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Sun, 12 Jul 2026 16:47:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2e40ce14-a532-4db6-a855-caee778250f7_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>The most serious test to date of open source AI&#8217;s viability is happening right now. I&#8217;ve seen </span><a href="https://www.interconnects.ai/p/sb-1047-and-open-weights?utm_source=publication-search"><span>many waves</span></a><span> of anti open-source AI rhetoric come and go since ChatGPT was launched, but none of them had obvious analogues in their potential enforcement to </span><em><span>real</span></em><span> action already in place targeting the peer, closed models of the day. It is more real because new forms of regulation are being tested and implemented, with minimal oversight. I will be doing far more policy-facing writing than usual until this passes.</span></p><p><span>As of writing this, many sources are citing White House </span><a href="https://x.com/jacob_wendler/status/2075565694322651376"><span>discussions</span></a><span> on how to manage open models via a new executive order. There is no official information here, and it would likely impact a) Chinese-origin models and b) government uses only, but this is how the dominoes start to fall.</span></p><p><span>Open models lack the central economic champion to represent the potential downside of action against them. From </span><a href="https://www.theinformation.com/articles/trump-administration-asks-openai-stagger-release-new-model-security-concerns"><span>recent coverage</span></a><span> of the events surrounding model licensing agreements for Fable (and then GPT-5.6), more was said about what unfolded on June 9th:</span></p><blockquote><p><span>At the meeting, the topic of how the program will deal with open-source AI models came up, according to a person familiar with the session. A representative from Reflection AI, a U.S.-based open-source model provider, argued that open-source models should have exemptions from the framework based on their capabilities, the person said. Currently, Chinese open-source models such as DeepSeek have a substantial lead over other available open models, and Reflection has not yet launched a public model.</span></p></blockquote><p><span>A ban of any form here would be a </span><a href="https://www.interconnects.ai/p/banning-open-source-ai-would-be-a"><span>big mistake</span></a><span> for the long-term trajectory of AI.</span></p><p><span>The most likely incoming action is to ban or indefinitely delay any open-weights model meaningfully above the capability level in the range of GPT 5.5, Claude Opus 4.8, or </span><a href="https://www.interconnects.ai/p/glm-52-is-the-step-change-for-open"><span>GLM-5.2</span></a><span>. With the consistent </span><a href="https://www.interconnects.ai/p/reading-todays-open-closed-performance?utm_source=publication-search"><span>capability gap</span></a><span>, this should be within the next 6 months.</span></p><p><span>As it stands, these would most likely be from a Chinese company, which is how this conversation of frontier open model capabilities inextricably becomes linked to other issues such as distillation. The capability threshold for a &#8220;right to review&#8221; from the government will shift over time, but once in place will likely progress far slower for open models rather than their closed counterparts. This is partially due to closed models being easier to secure but also due to the closed model companies having far more effective lobbying.</span></p><p><span>So, this leaves us in a place where there are two crucial policy discussions unfolding at once impacting open models &#8211; distillation &amp; frontier capabilities. They&#8217;re very different in their nature, the necessity of response, and the potential response space. Still, together they represent the talking points of a surging platform of support for a potential ban of open models in the next 6 months.</span></p><p><span>The primary driver motivating regulation today is the inevitable truth that an open-weights model will soon reach the capabilities of Claude&#8217;s Mythos model. The actual performance of this openly released model will likely be more jagged, but all it takes is the model getting flagged in the nascent White House AI model checker. It&#8217;s hard to unwind new habits motivated by fear.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/6-months-to-live-for-open-models?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/6-months-to-live-for-open-models?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3><span>The current distillation debates are regulatory capture and doing nothing for now is fine</span></h3><p><span>Distillation is largely a regulatory capture campaign at this point, as the only solutions on the table massively benefit the organizations pushing for it.</span></p><p><span>To elaborate, the anti-Chinese models political campaign is led by Anthropic, where they are sharing a mix of blog posts and </span><a href="https://www.bbc.com/news/articles/cwyklykn5dwo"><span>letters to representatives</span></a><span> detailing what the Chinese companies are doing. Anthropic has detected use from foreign companies, which is people coming and paying for its API, and eventually turned off usage and then written strongly worded recommendations of policy action and shared minimal technical evidence. This campaign may have started through a genuine business concern, but it has progressed to be the definition of regulatory capture, as Anthropic would gain substantial economic security in its products if the Chinese model makers they accused were banned.</span></p><p><span>If Anthropic was presenting information in a more neutral &#8220;you decide what to do&#8221; way, the community would have a lot more sympathy. It is more of a policy recommendation than an information sharing exercise at the frontier of a rapidly evolving technology. If Anthropic&#8217;s technology is as powerful as they say it is &#8211; so powerful that open models like it should likely be banned &#8211; then they should be able to secure their API. I continue to wait for them to explain why they cannot. One of their statements would need to be walked back.</span></p><p><span>Anthropic is also </span><a href="https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety"><span>pulling up the ladder for access to intelligence</span></a><span> in other ways &#8212; so the political recommendations they make in the vein of China competition are consistent with a much broader pattern of restriction of access to competitors of related technology in the vein of safety. It is easy to buy into company culture like an extra safety focus when many employees are on track for generational wealth. I do not blame the employees for this, but Anthropic&#8217;s corporate strategy should be understood in these broad, contextual lenses. Ben Thompson&#8217;s piece, </span><em><a href="https://stratechery.com/2026/anthropics-safety-superpower/"><span>Anthropic&#8217;s Safety Superpower</span></a></em><span>, is the best writing on the subject.</span></p><p><span>The action that Anthropic is effectively asking for is the wholesale banning of pretty much all the Chinese open weight models in the U.S. &#8212; as any products built around open models are predicated on their continued improvement, increasing product market fit and compute efficiency as models get better. This would demolish the open model economy that is emerging in the US with inference companies, fine tuning companies, new products, and everything in between. We as a community desperately need to hold the line that conceding anything with respect to the distillation conversation is not acceptable.</span></p><p><span>Anthropic should try to protect their IP, but they shouldn&#8217;t ask the government to cement their position and in the process potentially isolate the US from the global open-source community. There are no good solutions to distillation with the information we have, other than letting the labs self-enforce it.</span></p><p><em><span>I&#8217;ve written at length on distillation in particular. I&#8217;ve written directly on its impact on </span><a href="https://www.interconnects.ai/p/how-much-does-distillation-really?utm_source=publication-search"><span>model capabilities</span></a><span> and the </span><a href="https://www.interconnects.ai/p/the-distillation-panic?utm_source=publication-search"><span>regulatory environment</span></a><span>, and you can search for more in between the lines mentions on Interconnects.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>APIs aren&#8217;t magically secure</h3><p>There&#8217;s a particular messiness to the distillation discussion, where there&#8217;s likely meaningful worry of Chinese labs distilling the narrow <em>cybersecurity </em>capabilities of Mythos into an open model. This paints more on the insecurity of current model APIs more than it does on the distillation risk. Even when Claude Mythos was in its most-limited private beta, <em><a href="https://www.wired.com/story/security-news-this-week-discord-sleuths-gained-unauthorized-access-to-anthropics-mythos/">Discord Sleuths Gained Unauthorized Access to Anthropic&#8217;s Mythos</a></em>. To date, model APIs continue to be jailbroken and accessed in unintended ways. </p><p>This proliferates the risk of capabilities falling to bad actors far more rapidly than anything that involves complex fine-tuning of a 1T+ parameter model. I&#8217;m not a cybersecurity expert, but the dichotomy that only open-weight models are insecure and APIs are safe has been very overblown in recent years. APIs in concept <em>should</em> be able to be more secure, but that is yet to be demonstrated. </p><p>If Anthropic has a truly dangerous capability in their models, the only coherent action would be to not host it in a directly queryable API &#8212; before the discussion of it being distilled. </p><h3><span>Ready or not, open models are coming</span></h3><p><span>Where this becomes hard is that at the same time as these distillation questions, we&#8217;re staring down the barrel of &#8220;how do we handle frontier open weight models at the general capability level of Mythos?&#8221; This is a hard question, and it&#8217;s a natural human tendency to want to take a proposed solution to another problem (distillation) and apply it to the new, far more real problem (frontier capabilities).</span></p><p><span>We need to figure out the right policy for open frontier capabilities, but a flat out ban is likely not the answer. If the models are not banned in China as well, it is VERY easy for a bad actor to still use said banned open weight model, which negates the safety potential.</span></p><p><span>At the same time, if we alone ban the import of certain models, the global open-source community will continue. A world where the U.S. bans these models before China, or other more risk-sensitive cultures, would likely indicate that a form of AI fearmongering (or other social, political momentum) pushed the U.S. government to act early. It feels like speedrunning dystopia in the U.S., as our tech industry &#8211; the crown jewels of the economy &#8211; </span><a href="https://interconnect.substack.com/p/the-great-american-tech-crackdown"><span>looks far more like a Chinese system</span></a><span> with control and government investments. These are very bad outcomes!</span></p><p><span>The only way to add a ceiling on open-source progress is a global agreement on the management of risks of AI models, which we aren&#8217;t close to. Other delays make the AI rollout less predictable in the US and increasingly messy. It&#8217;ll feel a lot like the GPT-5.6 rollout, but instead of just limiting the upsides of open, cheaper intelligence to a few companies all the bad actors will immediately have access, too. Open models increase safety by broad access and understanding, not kneecapping the positive actors only.</span></p><p><span>In reality, there is no stopping the open-source ecosystem. The people building the best models are assessing risk as well, as China is very risk-sensitive and for example Z ai is already a public company exposed to a comprehensive set of pressures &#8211; keeping their rocketing stock up, for one. There will only continue to be more models that cross worrying levels of capabilities. Training AI models is not magic, and we haven&#8217;t seen access to building them drop off as the required investment has increased (yet).</span></p><p><span>One of the short-term off-ramps for this policy death spiral for open-source, a double-helix of related and complementary, scary issues, is for a company in the U.S. to release a similarly capable open model. This will shift the focus away from &#8220;only China is building the open models via distillation&#8221; to &#8220;we&#8217;re all in this together, and we should focus on the complex, moving frontier issues within our ecosystem.&#8221; This is an existential priority for open-source, and the companies who have the business reasons to release open-weight models like Microsoft and Meta (commoditizing their complements) should do it ASAP. I have less faith in Meta, with the new leadership, but they benefit from mass access to AI and should not get in their own way. If Reflection is sitting on an okay, but not frontier, model, they may need to release it to save their proposed business direction.</span></p><p><span>The shorter term solution than training a new model is to build the coalition. Open-source is a diffuse technology, without clear ownership, but the benefits will be shared by so many. This &#8220;everyone else&#8221; outside the frontier labs needs to start working today, on how to continue the safe rollout of open-weight models (and lobby for their principles and values) to the powers that be.</span></p><h3><span>Rising temperatures in the AI discourse</span></h3>
      <p>
          <a href="https://www.interconnects.ai/p/6-months-to-live-for-open-models">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Latest open artifacts (#22): Zyphra, Cohere, and Poolside are expanding the breadth of the ecosystem]]></title><description><![CDATA[An assessment of the open ecosystem and the motivations behind releasing models]]></description><link>https://www.interconnects.ai/p/artifacts-22-zyphra-cohere-and-poolside</link><guid isPermaLink="false">https://www.interconnects.ai/p/artifacts-22-zyphra-cohere-and-poolside</guid><dc:creator><![CDATA[Florian Brand]]></dc:creator><pubDate>Sun, 28 Jun 2026 17:03:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Yzio!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f6a1684-d8af-4a82-a0fc-ba65d6a6dde0_1024x576.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>A trend we </span><a href="https://www.interconnects.ai/p/latest-open-artifacts-20-new-orgs"><span>continue to see</span></a><span> in open model releases is that the ecosystem is becoming more diverse, with an increasing number of organizations releasing a wide range of models. A year ago, open artifacts and the open model landscape more broadly were dominated by a handful of (Chinese) players. This has shifted, with us increasingly featuring more niche companies all over the world.</span></p><p><span>While it is hard to know the exact motivations of the companies themselves, we can broadly observe the following categories:</span></p><ul><li><p><strong><span>&#8220;Pure&#8221; model makers</span></strong><span>: These are companies whose stated goal is to train models that are at the frontier, or at least close to it. This includes many Chinese companies, such as DeepSeek, Zhipu, and Minimax, but also Western ones like Poolside, Arcee, and Zyphra. It also increasingly includes sovereign AI players, such as Cohere, Sovereign, Mistral, and Trillion Labs. The recent Mythos episode has woken up some policymakers, which may lead to increased interest in sovereign model training.</span></p></li><li><p><strong><span>Big Tech</span></strong><span>: For Big Tech companies, including Alibaba&#8217;s Qwen, Google&#8217;s Gemma, and, to some extent, NVIDIA, the motivations are more diverse. Alibaba uses model releases to upsell its closed models, while NVIDIA benefits from a flourishing open model ecosystem as it increases interest in and usage of its GPUs. This vested interest is different from the Llama era of open Western models, where the motivations for open releases were less clear (and ultimately did not hold).</span></p></li><li><p><strong><span>Product companies</span></strong><span>: Some companies, such as JetBrains, Zed, Krea, and Photoroom, mainly sell products that use AI as a core component. As they don&#8217;t want to be cut off from </span><a href="https://techcrunch.com/2025/06/03/windsurf-says-anthropic-is-limiting-its-direct-access-to-claude-ai-models/"><span>accessing closed models</span></a><span> or want to offer something unique, they can train highly specialized, small models that fit their product needs. Thus, open-sourcing those model weights does not hurt their bottom line.</span></p></li></ul><p>This diversity of makers and models <a href="https://www.interconnects.ai/p/the-next-phase-of-open-models?utm_source=publication-search">fits our hypothesis</a> that more companies will develop a long-tail of models and the number of companies chasing the absolute, open frontier will diminish.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/artifacts-22-zyphra-cohere-and-poolside?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/artifacts-22-zyphra-cohere-and-poolside?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p><span>While not every model release fits neatly into one of these categories, the broader point is that open model development is not driven by a single type of actor or motivation. This diversity is one of the strengths of the open ecosystem and can be seen in the tech reports of model releases, which reuse training methods, architecture choices and data from other open model releases.</span></p><p><span>Attempts to slow or ban this ecosystem are not only futile, as the history of tech-related bans has shown, but also </span><a href="https://www.interconnects.ai/p/banning-open-source-ai-would-be-a"><span>unsafe and anti-freedom</span></a><span>. Such restrictions would concentrate AI development and usage among the select few, which ultimately endangers outsiders&#8217; ability to freely adopt one of the most important technologies of our lifetime.</span></p><h3><span>Our Picks</span></h3><ul><li><p><strong><a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16"><span>NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16</span></a></strong><span> by </span><a href="https://huggingface.co/nvidia"><span>nvidia</span></a><span>: The big version of the Nemotron series, which uses LatentMoE to be even faster than comparable models. Just like the other Nemotron models, the vast majority of the data is open source. And, to top it all off: NVIDIA commits to using the </span><a href="https://www.linuxfoundation.org/press/linux-foundation-releases-openmdw-1.1-nvidia-adopts-openmdw-for-cosmos-isaac-gr00t-ising-and-nemotron-ai-model-families"><span>OpenMDW</span></a><span> license, which is tailored specifically for model weights (and data) and drops its custom license. While MIT and Apache are in the same spirit as OpenMDW, only the latter really covers model weights, while the former are software licenses that do not really apply to model weights.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1nXJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1nXJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png 424w, https://substackcdn.com/image/fetch/$s_!1nXJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png 848w, https://substackcdn.com/image/fetch/$s_!1nXJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png 1272w, https://substackcdn.com/image/fetch/$s_!1nXJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1nXJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png" width="1223" height="707" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:707,&quot;width&quot;:1223,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1nXJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png 424w, https://substackcdn.com/image/fetch/$s_!1nXJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png 848w, https://substackcdn.com/image/fetch/$s_!1nXJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png 1272w, https://substackcdn.com/image/fetch/$s_!1nXJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90449d43-2181-4090-9e4b-8c94a05b111c_1223x707.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></li><li><p><strong><a href="https://huggingface.co/CohereLabs/command-a-plus-05-2026-bf16"><span>command-a-plus-05-2026-bf16</span></a></strong><span> by </span><a href="https://huggingface.co/CohereLabs"><span>CohereLabs</span></a><span>: Cohere, which is becoming more of a regular entrant into Artifacts lately, released their flagship, Command A+, under Apache 2.0. Previous iterations of the series have been released under a non-commercial license, so this change is more than welcome! Command A+ combines multi-modal, multi-lingual and agentic capabilities as a 218B-A25B MoE, making it usable with a single B200 (when using 4-bit).</span></p></li><li><p><strong><a href="https://huggingface.co/zai-org/GLM-5.2"><span>GLM-5.2</span></a></strong><span> by </span><a href="https://huggingface.co/zai-org"><span>zai-org</span></a><span>: The biggest story in this Artifacts is GLM-5.2, which we have covered in a separate </span><a href="https://www.interconnects.ai/p/glm-52-is-the-step-change-for-open"><span>blog</span></a><span> as well. The model continues to impress and is genuinely usable for everyday work, not a huge regression compared to the best closed models available right now. Interestingly enough, the raw download numbers since release are more in line with other model releases, with GLM-5.2 being roughly in line with GLM-5 after release.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sooJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sooJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png 424w, https://substackcdn.com/image/fetch/$s_!sooJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png 848w, https://substackcdn.com/image/fetch/$s_!sooJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png 1272w, https://substackcdn.com/image/fetch/$s_!sooJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sooJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png" width="1435" height="747" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png&quot;,&quot;srcNoWatermark&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bea2a2c0-471e-4eaa-867d-a2e95ee56edb_1435x747.png&quot;,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:747,&quot;width&quot;:1435,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:785336,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.interconnects.ai/i/203955384?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbea2a2c0-471e-4eaa-867d-a2e95ee56edb_1435x747.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sooJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png 424w, https://substackcdn.com/image/fetch/$s_!sooJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png 848w, https://substackcdn.com/image/fetch/$s_!sooJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png 1272w, https://substackcdn.com/image/fetch/$s_!sooJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d1ecb70-0e3d-435e-aabf-0e2b6e7d7138_1435x747.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></li><li><p><strong><a href="https://huggingface.co/Zyphra/ZAYA1-74B-preview"><span>ZAYA1-74B-preview</span></a></strong><span> by </span><a href="https://huggingface.co/Zyphra"><span>Zyphra</span></a><span>: Zyphra, which trains on AMD GPUs and is known as some sort of insider tip in the research community due to their tech reports with interesting architecture choices, has released some new models, with a 74B-A4B MoE and an </span><a href="https://huggingface.co/Zyphra/ZAYA1-8B"><span>8B-A0.6B MoE</span></a><span> (</span><a href="https://arxiv.org/abs/2605.05365"><span>tech report</span></a><span>) being their current flagship releases.</span></p></li><li><p><strong><a href="https://huggingface.co/poolside/Laguna-M.1"><span>Laguna-M.1</span></a></strong><span> by </span><a href="https://huggingface.co/poolside"><span>poolside</span></a><span>: Poolside, which we covered in </span><a href="https://www.interconnects.ai/p/latest-open-artifacts-21-open-model"><span>the last Artifacts</span></a><span>, also released their flagship model under Apache 2.0! They also commit to open releases </span><a href="https://x.com/poolsideai/status/2067623663562637683?s=20"><span>going forward</span></a><span>:</span></p><blockquote><p>Open weights are now our default. We&#8217;ll keep building toward the frontier and releasing increasingly capable models in the open.</p></blockquote></li></ul><h3><span>Models</span></h3><h4><span>General Purpose</span></h4><ul><li><p><strong><a href="https://huggingface.co/moonshotai/Kimi-K2.7-Code"><span>Kimi-K2.7-Code</span></a></strong><span> by </span><a href="https://huggingface.co/moonshotai"><span>moonshotai</span></a><span>: An update to Kimi focusing a lot on token efficiency.</span></p></li><li><p><strong><a href="https://huggingface.co/stepfun-ai/Step-3.7-Flash"><span>Step-3.7-Flash</span></a></strong><span> by </span><a href="https://huggingface.co/stepfun-ai"><span>stepfun-ai</span></a><span>: An update to Step-Flash, which is really strong in Math in particular.</span></p></li><li><p><strong><a href="https://huggingface.co/nvidia/Nemotron-Labs-Diffusion-14B"><span>Nemotron-Labs-Diffusion-14B</span></a></strong><span> by </span><a href="https://huggingface.co/nvidia"><span>nvidia</span></a><span>: An experimental model which can be used in three different modes: autoregressive, diffusion, and self-speculation. Each of these modes is suitable for a different use case.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ylaJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ylaJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png 424w, https://substackcdn.com/image/fetch/$s_!ylaJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png 848w, https://substackcdn.com/image/fetch/$s_!ylaJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png 1272w, https://substackcdn.com/image/fetch/$s_!ylaJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ylaJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png" width="338" height="240.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1036,&quot;width&quot;:1456,&quot;resizeWidth&quot;:338,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;An illustration of Tri-Mode LMs&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="An illustration of Tri-Mode LMs" title="An illustration of Tri-Mode LMs" srcset="https://substackcdn.com/image/fetch/$s_!ylaJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png 424w, https://substackcdn.com/image/fetch/$s_!ylaJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png 848w, https://substackcdn.com/image/fetch/$s_!ylaJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png 1272w, https://substackcdn.com/image/fetch/$s_!ylaJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e166405-1e61-4b66-8774-0ab29072e3f4_2401x1709.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></li></ul>
      <p>
          <a href="https://www.interconnects.ai/p/artifacts-22-zyphra-cohere-and-poolside">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[GLM-5.2 is the step change for open agents]]></title><description><![CDATA[A capability threshold I've been carefully monitoring.]]></description><link>https://www.interconnects.ai/p/glm-52-is-the-step-change-for-open</link><guid isPermaLink="false">https://www.interconnects.ai/p/glm-52-is-the-step-change-for-open</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 22 Jun 2026 14:52:45 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/267cdc82-6fbc-4f68-bcba-160091396dfd_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>Housekeeping: Following my &#8220;<a href="https://www.interconnects.ai/p/state-of-the-blog-mid-2026">State of the blog</a>&#8221; post last week, noting a slight increase in paid features, it&#8217;s a good time to remind folks that I offer <a href="https://www.interconnects.ai/about#&#167;group-paid-subscriptions">group subscriptions</a> with larger discounts proportional to the number of seats. <br>I also released a new paper today on open RL recipes for terminal agents, read more <a href="https://natolambert.substack.com/p/tmax-an-open-rl-recipe-for-terminal">here</a>.</h5><p>A bit over a week ago, when the AI world was still reeling from the shocking <a href="https://www.interconnects.ai/p/welcome-to-the-agi-era-of-ai-governance">export restriction, and effective banning</a>, of <a href="https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety">Claude Fable 5</a>, Z.ai released their latest model, GLM-5.2. This model was <a href="https://x.com/Zai_org/status/2065704919299235870">rolled out</a> unusually on a Saturday, June 13th, to GLM Coding Plan members. This is an unusual release practice, normally when an AI model is released on a weekend it&#8217;s for a weird reason (most famously, <a href="https://www.interconnects.ai/p/llama-4">Llama 4</a>).<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> In this case, it seemed like Z.ai was excited to capitalize on the zeitgeist of &#8220;Anthropic being anti open-science&#8221; with their silent safeguards on AI researchers. For the past year or two, the Chinese open-weight labs have taken every opportunity they have for easy marketing wins like this.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/glm-52-is-the-step-change-for-open?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/glm-52-is-the-step-change-for-open?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>GLM-5.2, in a common naming convention across the industry, looked potentially like an incremental update following the popular GLM-5.1 model. At this point, Moonshot AI, makers of the Kimi models, and Z.ai, makers of the GLM models, have consolidated the top of the reputational market with the most beloved open-weight models among AI researchers. What unfolded is a common lesson in tracking AI models that often minor version numbers can have AI models crossing meaningful user experience thresholds. A small change in benchmarks and training can open a wide range of new use-cases.</p><p>What has followed is a slow, groundswell of hype for GLM-5.2. The official, MIT-licensed <a href="https://huggingface.co/zai-org/GLM-5.2">model weights</a> and <a href="https://z.ai/blog/glm-5.2">release blog</a> dropped three days after the initial rollout, on June 16th. One could ramble many technical details, such as the strong benchmark scores, the very popular RL framework that Z.ai uses (<a href="https://github.com/THUDM/slime">SLIME</a>), the recommendation of always using the model on Max thinking effort, and so on, but the initial release blogs usually aren&#8217;t the thing to focus on. You can wait and read the ecosystem reaction to know if it&#8217;s the real deal. <a href="https://www.interconnects.ai/p/opus-46-vs-codex-53">Benchmarks are half dead these days</a>, anyways.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xhhJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xhhJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png 424w, https://substackcdn.com/image/fetch/$s_!xhhJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png 848w, https://substackcdn.com/image/fetch/$s_!xhhJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png 1272w, https://substackcdn.com/image/fetch/$s_!xhhJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xhhJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png" width="1456" height="961" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:961,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;img_v3_0212o_51684a16-c33f-4429-aea5-9f5f7cdfc30g&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="img_v3_0212o_51684a16-c33f-4429-aea5-9f5f7cdfc30g" title="img_v3_0212o_51684a16-c33f-4429-aea5-9f5f7cdfc30g" srcset="https://substackcdn.com/image/fetch/$s_!xhhJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png 424w, https://substackcdn.com/image/fetch/$s_!xhhJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png 848w, https://substackcdn.com/image/fetch/$s_!xhhJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png 1272w, https://substackcdn.com/image/fetch/$s_!xhhJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>What followed on the 16th was a slew of community benchmarks showing better-than-expected results for GLM-5.2. <a href="https://x.com/arena/status/2066943450914943025">Arena&#8217;s agent leaderboard</a> had it as the only open model mixing it up with OpenAI and Anthropic&#8217;s latest models (notably matching Opus 4.8&#8217;s no-thinking effort to GLM-5.2&#8217;s max mode). This is one of many evals GLM-5.2 is crushing Gemini on, but that&#8217;s a topic for another time. A benchmark that has mixed perception in the community (particularly among actual designers), <a href="https://x.com/Designarena/status/2066940737011560652">Design Arena</a> even had GLM-5.2 besting Claude Fable itself &#8212; the recently banned hype machine!</p><p>Pretty much everyone I respect among the AI commentariat and researcher class has praised the model after using it personally. Such a focal point of discussion among the community has only been so clear with an open model release once before &#8212; <a href="https://www.interconnects.ai/p/deepseek-r1-recipe-for-o1">DeepSeek R1</a>. This is not a comparison I make lightly, and when I compared <a href="https://www.interconnects.ai/p/kimi-k2-and-when-deepseek-moments">Kimi K2&#8217;s release to a &#8220;DeepSeek Moment,&#8221;</a> GLM-5.2 has well exceeded that. What made Kimi K2 impressive was that big steps in open model performance could seemingly come from <em>anywhere</em> in China. The step that GLM-5.2 has taken is more of a one way door for AI progress.</p><p>Anthropic&#8217;s record revenue growth rate on the back of Claude Code is heavily driven by being the best model, and the only model that can really do this. GLM-5.2 is the first of many (coming soon) open weight models to offer credible alternatives. The parallel is very clear, to when DeepSeek R1 showed that open-weight labs, with far fewer resources, could also replicate the chain-of-thought reasoning models that OpenAI championed with o1. As AI systems get more complex and far more expensive to build, with tools, integrated harnesses, and scaled model weights, it was not a given that this GLM-5.2 moment would happen at all.</p><p>The key point is that <strong>GLM-5.2 is the open weight model that <a href="https://www.interconnects.ai/p/claude-code-hits-different?utm_source=publication-search">feels right</a> in coding harnesses as a general agent</strong>. <strong>It&#8217;s the first one.</strong> I was personally overdue in trying some of the recent peer models, such as Kimi K2.7 or GLM-5.1, but the hype was too much for me to ignore. I put it to work helping make content for my <a href="https://github.com/natolambert/rlhf-book/pull/457">post-training course</a> with Fireworks&#8217; API in Claude Code (<a href="https://docs.fireworks.ai/ecosystem/fireconnect/claude-code">setting this up</a> was <em>very</em> easy). There were some minor knife cuts, such as the Claude Code harness / my repo documentation trying to send images to the model, which would brick Fireworks API for the session &#8212; forcing a manual context clear. Overall, the model capabilities immediately felt right, and I still have some tinkering to do in which harness and inference provider to use. </p><p>For more hype, you can sample the Z.ai founder telling Elon that &#8220;<a href="https://x.com/pmarca/status/2067640859957539104">open-weight Fable capabilities will be here sooner than Q1 2027</a>,&#8221; the CEO of Vercel <a href="https://x.com/rauchg/status/2068517095818809770">saying</a> &#8220;Genuinely impressed, almost shocked, at how good GLM-5.2 by @zai_org is at coding. This changes things,&#8221; and much <a href="https://x.com/ArtificialAnlys/status/2067135640249209175">more</a> from a mix of people whose opinions I <a href="https://x.com/gneubig/status/2067936197888930263?s=20">deeply</a> <a href="https://x.com/_xjdr/status/2068422921249529916">respect</a> and others I&#8217;m <a href="https://x.com/matvelloso/status/2067791546335019439?s=20">new to</a>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>So, this is a good model, where does this leave us?</p><p>There are many trends at play. To start, let&#8217;s ground things in the open-closed capabilities gap. I&#8217;ve written how I expect an &#8220;<a href="https://www.interconnects.ai/p/some-ideas-for-what-comes-next-may">explosion in usage</a>&#8221; if open models crossed the Opus 4.5 in Claude Code threshold from around the start of 2026. Here we are. With Claude Opus 4.5&#8217;s release on November 24th, 2025, the gap in time to GLM-5.2&#8217;s release on June 16th, 2026 is 204 days &#8212; or about 6.8 months. This puts us square in the 6-9 month time gap that many people claim as the performance lag between the U.S.&#8217;s closed labs and China&#8217;s open counterparts.</p><p>Upon writing this, I&#8217;m surprised. As the U.S. labs have so rapidly ramped compute in the last ~year, I&#8217;ve expected the gap in performance to grow in time. A very meaningful step in this trajectory will also be Claude Fable 5&#8217;s release &#8212; which was more reliant on scale, and therefore the most advanced GPUs, relative to the Claude Opus models. Still, that&#8217;s not a satisfactory answer. Continuing to unpack the trajectory here involves more nuance than I can afford to fit in a signposting article.</p><p>The most immediate meaning of this is far more serious pricing pressure within the organizations tokenmaxxing, sending Anthropic&#8217;s revenue to the moon. Some would predict Anthropic doesn&#8217;t realize its forecasted ARR numbers, but I don&#8217;t think that prices in the true demand for these models and the inevitable growth. This model existing is a huge boon for the open model <em>economy</em>. All the likes of Fireworks, Together, Thinky (via Tinker), Prime Intellect, and whoever else sells open model inference or finetuning just hit another inflection point. </p><p>It&#8217;ll take a long time for the effects here to diffuse into the broader economy (and use-cases). Workflows are becoming more complex, with people using different models for planning, primary coding, and subagent dispatch. I expect the hype to continue to grow, and heck, as I&#8217;m writing this on a Sunday evening, I could see the media and market reaction on the Monday being a thing just like the DeepSeek R1 release. This diffusion happening while Anthropic&#8217;s, and by extension the U.S.&#8217;s flagship model, is still banned is a severe economic dagger. GLM-5.2 is being given time to carve out the economic underbelly of the frontier labs when they want to be pushing forward into higher margin, higher revenue domains enabled only by the absolute frontier models.</p><p>The economic concern mirrors a story that has been told many times in AI, so it&#8217;s unclear when it&#8217;ll stick.</p><p>The conversation that feels more core to the trajectory of AI is that of regulation and control of open models. I think it is an economic good for cheap intelligence to diffuse widely, and our default position should be to cheer for open models, but this model&#8217;s release date will have it be permanently associated with Claude Fable &#8212; and therefore Claude Mythos &#8212; in the mental map of AI power structures. We are at a point where Mythos-class model capabilities are deemed not safe for release by the U.S. Government and the Chinese model makers are charging forward in capabilities available to all. </p><p>These trend lines aren&#8217;t necessarily causally linked, as we don&#8217;t know the cyber performance of GLM-5.2 versus its predecessors, but the capabilities are definitely correlated. Without anything changing, this points to a potentiality where the U.S. Government decides a certain open-weights Chinese model is not safe for the public. There are many other potential scenarios here too, but what is clear is that we have a lot of work to do in mapping them out, preparing our infrastructure, and messaging to society. </p><p>It&#8217;ll take a lot more people than just me to imagine and communicate a world to decision makers for how to manage evermore capable open models.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> We have years more of AI progress to come, with Nvidia&#8217;s next generation chips already in production and a constant stream of algorithmic advancements. It feels like a narrow path for open model advocates to take, but we need to figure out how to make them viable so the massive leaps in performance don&#8217;t only go to closed models. </p><p>I totally see why it is scary to imagine an openly accessible Mythos class model, but if open models get banned now and only closed models get 10 or 100X better in 2 years in the hands of one or two companies, I think we will have bigger problems on our hands.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Something that has always stood out to me is how fast the Chinese labs release their models. I&#8217;ve heard from multiple labs that the time to upload the weights publicly to HuggingFace after the model finishes training could be measured in hours rather than days. This has at least slowed a bit, now that they need to prepare to serve the model to a wider inference market. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Something that will need to be discussed more is how even closed models, e.g.  Mythos preview, are regularly in the hands of unauthorized users or jailbroken. So, the open vs. closed dichotomy on access isn&#8217;t totally black and white. </p></div></div>]]></content:encoded></item><item><title><![CDATA[Banning Open Source AI Would Be A Mistake]]></title><description><![CDATA[This post was originally an op-ed co-authored with Kevin Xu of Interconnected for a general, non-technical audience.]]></description><link>https://www.interconnects.ai/p/banning-open-source-ai-would-be-a</link><guid isPermaLink="false">https://www.interconnects.ai/p/banning-open-source-ai-would-be-a</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Fri, 19 Jun 2026 13:02:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2cf763f4-5f1e-47ec-b6e1-b14a69443d00_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><span>This post was originally an op-ed co-authored with </span><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Kevin Xu&quot;,&quot;id&quot;:9714824,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8724733-4f91-46b4-a37d-652026b382ae_400x400.jpeg&quot;,&quot;uuid&quot;:&quot;d8ef2a3e-75e3-4f28-91b6-5869a8879d8d&quot;}" data-component-name="MentionToDOM"></span></em> <em>of <a href="https://interconnect.substack.com/">Interconnected</a></em> <em><span>for a general, non-technical audience. The gatekeepers &#8212; the many media outlets we pitched it to &#8212; passed on publishing it. Luckily, we have our own platforms to get the message out. Please help us forward this op-ed to any one you know who is on the fence about open source AI or new to the topic and want to learn more. Thank you.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/banning-open-source-ai-would-be-a?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/banning-open-source-ai-would-be-a?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><p><span>The energy to regulate AI is in the air in Washington. With the recently signed executive order to review AI models, </span><a href="https://www.politico.com/2026/06/04/obernolte-trahan-ai-draft"><span>a congressional proposal</span></a><span> to legislate AI further, the government </span><a href="https://www.cnbc.com/2026/06/05/trump-open-ai-altman-stake.html"><span>possibly taking shares</span></a><span> of frontier AI labs, and last Friday&#8217;s </span><a href="https://www.wsj.com/tech/ai/amazon-ceos-talks-with-u-s-officials-triggered-crackdown-on-anthropic-models-dcc90578?mod=Searchresults&amp;pos=1&amp;page=1"><span>action</span></a><span> prohibiting foreign nationals anywhere from accessing Anthropic&#8217;s most advanced models, this may be the opening salvo of more AI regulation to come.</span></p><p><span>We are afraid future actions could inadvertently or intentionally regulate or even ban open source, a much maligned and misunderstood topic in AI. That would be a grave mistake.</span></p><p><span>Open source &#8211; simply a process that allows technology to be shared, built, and distributed publicly and transparently &#8211; is safe, secure, and drives economic growth. More than </span><a href="https://github.blog/news-insights/research/octoverse-2022-10-years-of-tracking-open-source/"><span>90% of the world&#8217;s software</span></a><span> was already built on open source and produced more than </span><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4693148"><span>8 trillion dollars</span></a><span> worth of economic benefits, long before AI entered the picture. Today, open source technology is quietly training, improving, deploying, and securing AI everywhere.</span></p><p><span>For more than three decades, open source has been powering three trends, and upholding three values, which the American society holds dear &#8211; education, competition, and innovation.</span></p><p><strong><span>Open source is pro-education</span></strong><span> because its origin was rooted in academic institutions trying to make technology free and open, not held hostage to the profit-maximizing zeal or the menacing lawyers of large corporations.</span></p><p><span>The precursor of open source is the free software movement, which started in 1983 on the campus of MIT. It was a time when every small act of using software, whether it was teaching students or doing research or improving a printer&#8217;s performance, meant paying or dealing with big corporations like AT&amp;T or Xerox. After this struggle gave birth to open source, every student in every university, community college, and coding bootcamp in America now taps into the freedom that open source enables to learn how to program, engineer, and build. Open source is at the heart of technical education everywhere.</span></p><p><strong><span>Open source is pro-innovation</span></strong><span> because it essentially provides a set of tools plus a community of other users to help anyone turn an idea into reality, for free. Combined with its role in education, it has watered most of the seeds of innovation in recent memory. Some of these seeds stayed as hobbies that brought joy and personal learning to the hobbyists. Others blossomed into huge companies, like Meta, where the initial version of Facebook was built entirely on a stack of open source software.</span></p><p><span>Every day, new ideas or solutions are being coded up in a dorm room, garage, or basement, all because open source lets innovators create without fear of a lawsuit or an expensive bill.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/subscribe?"><span>Subscribe now</span></a></p><p><strong><span>Open source is pro-competition</span></strong><span> because it helps the underdogs challenge and compete with the large incumbents, keeping monopolistic threats at bay. Linux, the open source operating system that now runs more than 90% of the world&#8217;s cloud computing infrastructure, was the antidote to the Windows monopoly (so much so that former Microsoft CEO, Steve Ballmer, called Linux &#8220;cancer&#8221;). Android, the open source mobile system, fostered a long string of competitive smartphones before Apple&#8217;s iPhone could control the market. Many other examples exist in the more niche, but no less important, segments of self-driving, databases, and semiconductor design.</span></p><p><span>Without the equalizing and democratizing nature of open source, we would all be living with the rent-seeking consequences of more monopolies and less free market competition.</span></p><p><span>Does AI change any of this? No.</span></p><p><span>The duopoly of Anthropic and OpenAI are rapidly concentrating power between them with their closed, proprietary models. Anthropic, in particular, has flexed its monopolistic muscle recently by </span><a href="https://www.wired.com/story/anthropic-responds-to-backlash-on-claudes-secret-sabotage-on-ai-research/"><span>reducing its most advanced model&#8217;s capability</span></a><span> when it is being used to improve someone else&#8217;s model. While the capabilities of their models are undeniable, so are their price tags and market concentration. Open source AI, mostly in the form of open weight models, has been the only counterweight for startups, educational institutions, and enterprises looking for alternatives.</span></p><p><span>Does open source lead to more safety or security concerns? Not quite.</span></p><p><span>We acknowledge it is worth monitoring the security implications of open source models that may reach frontier capabilities. But for the most part, the transparency that is inherent to open source makes them safer and more secure, because more engineers and researchers can tune out unwanted model behaviors, like censorship, or fix bugs in the software that runs these models. As one popular </span><a href="https://en.wikipedia.org/wiki/Linus%27s_law"><span>saying goes</span></a><span>, &#8220;given enough eyeballs, all bugs are shallow.&#8221; An open source model also does not transfer data, when installed on your own company&#8217;s infrastructure as Airbnb CEO, Brian Chesky, </span><a href="https://www.bloomberg.com/news/articles/2026-05-20/airbnb-s-chesky-says-us-misunderstanding-use-of-chinese-open-source-ai-models"><span>explained</span></a><span>. Open source AI is the most secure and privacy friendly path.</span></p><p><span>What about China? Beware of unintended consequences.</span></p><p><span>China is certainly a fierce competitor with the US on many dimensions &#8211; economically, militarily, diplomatically &#8211; but using this dynamic as a pretext to regulate open source will backfire.</span></p><p><span>Open source models are actually improving the efficiency and profitability of many American startups, who cannot afford to pay the monopoly-level premium to Anthropic or OpenAI. AI companies working in </span><a href="https://cursor.com/blog/composer-2-technical-report"><span>coding</span></a><span>, </span><a href="https://fireworks.ai/blog/open-source-agents-frontier-advisors"><span>legal</span></a><span>, and other domains are using open source models, including ones from China, every day. The fact that these models are made by Chinese labs should be a wake-up call that open source is under-invested and under-appreciated in America! The response should be more support for open source at home. Regulating or limiting open source because of China would achieve the opposite: putting a chilling effect on education, innovation, and competition, while pushing the rest of the world &#8211; much of which wants open source&#8217;s benefits as much as we do &#8211; to adopt China&#8217;s.</span></p><p><span>Former Supreme Court Justice Louis Brandeis, famously said, &#8220;sunlight is said to be the best of disinfectants,&#8221; when it comes to removing corporate or societal misconduct. Open source is that &#8220;sunlight&#8221; in technology and AI. America should always be on the side of light.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/banning-open-source-ai-would-be-a?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/banning-open-source-ai-would-be-a?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[State of the blog, mid-2026]]></title><description><![CDATA[About 3 years since I started writing weekly.]]></description><link>https://www.interconnects.ai/p/state-of-the-blog-mid-2026</link><guid isPermaLink="false">https://www.interconnects.ai/p/state-of-the-blog-mid-2026</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Wed, 17 Jun 2026 14:29:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/06077e73-d656-4396-bda7-34b45e920822_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As I navigate my career change <a href="https://www.interconnects.ai/p/farewell-ai2">after Ai2</a>, I wanted to share my views of how this blog relates to my missions and broader work. In my farewell post, I summarized my three goals right now as:</p><ol><li><p>Provide clarity in the evolution of frontier models. </p></li><li><p>Create a vibrant and diverse open (model) ecosystem.</p></li><li><p>To build institutions that make these goals possible.</p></li></ol><p>Within this, Interconnects is at its core a bit different than many of the highly-polished, professional newsletters on this platform &#8211; and this is becoming intentional.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/state-of-the-blog-mid-2026?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/state-of-the-blog-mid-2026?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2>How Interconnects fits into my career goals</h2><p>Interconnects is the tip of the spear of all of my missions in AI. It is meant to start a conversation and to let the reader into the mind of someone at the frontier. This insight makes the writing sometimes a bit raw, sometimes a bit too technical, but it is the map of how I progress my thinking in the ever changing world. </p><p>This style of writing has helped me create very strong relationships with the core group of readers, many of who listen to the voiceovers I do for these posts. The plan is to keep operating and refining the Interconnects experience around those loyal fans. These are to a large part people building the frontier AI ecosystem &#8212; researchers at labs, top investors, policymakers obsessed with the frontier, and students aspiring to have one of those roles.</p><p>I&#8217;m very happy with this sort of raw, high-voice outcome for the blog. It is not something I sought out, but rather accepted as I saw it coming and realized it would be disproportionately successful in a near-future of vast AI slop media. With years of trying to squeeze writing into a busy schedule, the only sort of writing I had time for was that which had a style very closely matching how I think.</p><p>I&#8217;m also very happy to be an independent voice. As a person I don&#8217;t do well with some power structures like having a boss, and I think there are very few people without extreme financial conflicts of interest that are willing and allowed to write. Through a wide job search, few companies were genuinely excited about me continuing writing.</p><p>Over the past few months, I considered taking Interconnects in more of a direction like SemiAnalysis or Stratechery, where it is my full-time gig and number one priority, but it didn&#8217;t seem like the right fit for what I am trying to achieve. I&#8217;m trying to build an open ecosystem and a movement for true open-science at the frontier of AI. These areas are very narrowly populated and trying to influence them with only commentary, analysis, and related research products wouldn&#8217;t work for me.</p><p>These sorts of full-time outcomes are definitely still one of my dreams, and I will do it at some point. The dream of this is also one of the reasons I take conflicts of interest seriously. Though, in this era of AI I can&#8217;t be fully on the outside.</p><p>In this vein, I wanted to disclose two advising agreements I recently signed. I don&#8217;t view them as a compromise of the above independence, as I&#8217;ll happily quit if I feel like I can&#8217;t speak my mind, but as a form of support in accomplishing my missions. </p><p>If I want to make a true open-science ecosystem I have some catching up to do with how the frontier labs approach post-training. The two companies I&#8217;m advising, whose leadership I&#8217;ve become friends with, are Arcee AI and Mercor. Arcee should be fairly obvious as the no-nonsense player building open-weight models. Mercor will make more sense over time, but they&#8217;re a close ally to a lot of my goals in transparent evaluations, open post-training, and neutrality with respect to the leading labs. These advising agreements are based on me wanting to learn more, and I don&#8217;t suspect I will ever engage in the very cursory advising roles that are more of name-stamping.</p><p><em>I keep an up-to-date disclosures statement at the end of the Interconnects about page: <a href="https://www.interconnects.ai/about">https://www.interconnects.ai/about</a></em>.</p><p>Otherwise, my full-time job should still be in the non-profit sector as long as I get the next few months of logistics right.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Some operations &amp; audience notes</h2><p>Interconnects has cultivated an excellent, niche, and largely technical audience with representatives of all the top companies and labs (recently crossed 70K subscribers). I intend to protect this niche audience rather than trying to expand to bigger pastures. I think this success in audience alignment is reflected in my ~900 paid subscribers supporting it with infrequent paywalled content. I appreciate the support greatly, as the money has let me expand Interconnects operations and quality over the last 18 months.</p><p>I created Interconnects AI, LLC last January along with business bank accounts. Since then I&#8217;ve made some money, but I&#8217;ve reinvested it (and more) back into the business and the various AI services I need to try to write these articles. So, at this moment going full-time on Interconnects is a pretty risky financial proposition for me. In fact the Interconnects bank account has hovered around $0 for months (I&#8217;m personally fine having another job). This made me hesitate in going all-in on it, but in reflections I concluded that I would have more impact in AI by building these systems than focusing on commentary. </p><p>Second, as AI services get more expensive (e.g. Fable becoming API only), I&#8217;m going to need to spend more out of pocket to make this happen. I&#8217;m happy to do this in the near term, but I&#8217;m starting to optimize the blog to have more consistent financial growth, so when I want to go all in on writing in a few years I have a safety net.</p><p>I don&#8217;t do special offers, free trials, etc. for Interconnects paid subscribers (mostly to mitigate noise in the <a href="https://www.interconnects.ai/p/discord">Discord</a> community), but if you have the means to <a href="https://www.interconnects.ai/subscribe">support</a> this project it would mean a lot to me as I center my career around it. Joining a lab or a well-paying startup would be a much simpler path for me and my family but it&#8217;s never felt like the right thing to do. </p><p>I have a very arbitrary goal of reaching the 1000 paid subscribers orange checkmark on Substack this summer. So you can help and/or just watch my attempts to make it happen.</p><p>In this vein, I wanted to be direct in sharing how I view a few core operational components of Interconnects, and what you can expect going forward.</p><ol><li><p><strong>All comments will be paywalled</strong>. Whenever I have a popular post without paywalled comments I get a flood of low-quality posts &#8212; many of which are obviously AI generated. This is a detriment of the highly selective audience we&#8217;ve built. If Substack supports a feature like &#8220;only users with a paid subscription somewhere on the platform can engage,&#8221; I&#8217;d implement it. The blog comments, Substack chat, and Discord will be spaces where I perform active curation to maintain a 0% AI slop rate.</p></li><li><p><strong>Slightly more articles will be paywalled</strong>. I want to keep experimenting with what is the right way to do this, but the only metric I can rely on for increasing influence of the blog is revenue. Views, likes, etc. are all vanity metrics which don&#8217;t reliably measure this type of content. Cultivating a highly engaged audience is existential to me in attempting to maintain an AGI-proof expertise.</p></li><li><p><strong>Slightly more in-person events</strong>. With a small community that I respect, I have to opportunity to translate that to excellent real-world experiences. I expect to keep these small, but I want to be more proactive at organizing them so loyal readers know what to expect. The few coming soonest will be for my book launch, which should be in the next month or two. Plus, I know people always want to meet likeminded folks in AI!</p></li></ol><p>Together these should make it easier and more enjoyable to be a loyal fan for Interconnects. I&#8217;m looking forward to continuing convincing my fans that the support is worthwhile.</p><p>Thanks for reading! My career wouldn&#8217;t be possible without all of the support.</p>]]></content:encoded></item><item><title><![CDATA[Frontier post-training recipe review with Finbarr Timbers]]></title><description><![CDATA["Interview" #18]]></description><link>https://www.interconnects.ai/p/frontier-post-training-recipe-review</link><guid isPermaLink="false">https://www.interconnects.ai/p/frontier-post-training-recipe-review</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Tue, 16 Jun 2026 13:29:30 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/201754262/768200397f40335e04d3ab284c0ef0fb.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p>As I&#8217;ve been recapping fundamentals of post-training to wrap up my <a href="https://rlhfbook.com/">RLHF / Post-training book</a> I knew I needed to get <a href="https://finbarr.ca/">Finbarr Timbers</a> back on the podcast to talk about the state of play. Over the last few months we&#8217;ve had many discussions on what we&#8217;d need to do to take an Olmo-style recipe to the frontier, supported by Finbarr&#8217;s extensive reading of recent model technical reports.</p><p>To prepare for this, I put together a summary <a href="https://rlhfbook.com/teach/course/conversation-01/#1">slide deck</a> on the key post-training recipes historically &#8212; the path from InstructGPT to today &#8212; and today &#8212; the key open frontier models. This deck is summarized below as the technical summary, but we do spend 20-35 minutes on it in the podcast, so watching on <a href="https://www.youtube.com/watch?v=sbXEPxIazqY&amp;list=PLL1tdVxB1CpVpEtMHxwuR4uI4Lxjw00_y&amp;index=10">YouTube</a> is likely the best experience for this one.</p><p>I previously <a href="https://www.interconnects.ai/p/finbarr-timbers">interviewed</a> Finbarr in December of 2024, shortly after the release of o1 and T&#252;lu 3 (and before he joined Ai2) on the &#8220;We are so back&#8221; era of RL.</p><p>Chapters:</p><ul><li><p>00:00 Introduction &amp; Olmo reflections</p></li><li><p>06:28 Post-train recipes review (history)</p></li><li><p>23:00 2026&#8217;s model recipes (MiMo Flash, DeepSeek V4, GLM 5, Kimi K2.6, etc.)</p></li><li><p>39:05 Open-ended post-training discussions</p></li><li><p>48:22 Career advice in the LLM race</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/frontier-post-training-recipe-review?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/frontier-post-training-recipe-review?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>Listen on <a href="https://podcasts.apple.com/us/podcast/interconnects-audio/id1719552353">Apple Podcasts</a>, <a href="https://open.spotify.com/show/6XNzfJULeVxR7SneeesDUs">Spotify</a>, and <a href="https://www.interconnects.ai/podcast">where ever you get your podcasts</a>. For other Interconnects interviews, <a href="https://www.interconnects.ai/t/interviews">go here</a>.</p><p>For more educational post-training videos, see the <a href="https://rlhfbook.com/course">course</a> I&#8217;m putting together.</p><div id="youtube2-sbXEPxIazqY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;sbXEPxIazqY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/sbXEPxIazqY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><h2>Technical Summary</h2><p><em>These are notes cleaned up from a slide-deck created with AI assistance &#8212; mostly useful as a discussion topic and reference.</em></p><p>The shape of a post-training recipe has changed more in the last year than in the prior three.</p><ul><li><p>2022&#8211;2023 (InstructGPT): one pipeline &#8212; SFT &#8594; reward model &#8594; RL.</p></li><li><p>2024 (Llama 3, T&#252;lu 3, etc.): open recipes formalize SFT &#8594; DPO &#8594; RL with verifiable rewards. Closed recipes use many stages of RLHF.</p></li><li><p>2025 (DeepSeek R1): reasoning RL (R1) makes large-scale RL the centerpiece.</p></li><li><p>2026 (MiMo Flash V2): recipes fragment into <em>many specialist models</em> that are merged back into one.</p></li></ul><p><strong>The new thing: MOPD</strong></p><p>Multi-teacher On-Policy Distillation (MOPD) is the pattern showing up across the 2026 frontier.</p><ol><li><p>Train N domain-specialist teachers (each: SFT, then RL on the relevant domains).</p></li><li><p>Train one general student by sampling <em>its own</em> trajectories (this is the final post-trained model).</p></li><li><p>On each rollout, minimize reverse-KL to the <em>relevant</em> teacher&#8217;s output distribution, token by token.</p></li></ol><p>Lineage: MiMo Flash v2 introduced it &#8594; DeepSeek V4 &amp; Nemotron 3 Ultra scale it to &gt;10 teachers.</p><p><strong>Why did MOPD emerge?</strong></p><ul><li><p>RL got expensive and conflict-prone. Mixing math, code, and agentic RL in one run eventually trades capabilities off against each other.</p></li><li><p>Specialists are cheap to make / organizationally scalable. SFT-then-RL on a single domain is well understood and parallelizable. As post-training becomes more complex, scaling it across organizations is a big win.</p></li><li><p>On-policy distillation matured. Literature and know-how continued to emerge through the RLVR renaissance.</p></li></ul><p><em>Sources: <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf">DeepSeek V4 &#167;5.1</a>, <a href="https://arxiv.org/abs/2601.02780">MiMo-V2-Flash</a></em></p><div><hr></div><h3>Key historical recipes</h3><p><strong>InstructGPT (Mar. 2022) &#8212; the canonical 3 steps</strong> &#183; <a href="https://arxiv.org/abs/2203.02155">paper</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RXNT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RXNT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png 424w, https://substackcdn.com/image/fetch/$s_!RXNT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png 848w, https://substackcdn.com/image/fetch/$s_!RXNT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png 1272w, https://substackcdn.com/image/fetch/$s_!RXNT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RXNT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png" width="1456" height="909" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:909,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;InstructGPT: SFT on demonstrations &#8594; reward model on comparisons &#8594; PPO&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="InstructGPT: SFT on demonstrations &#8594; reward model on comparisons &#8594; PPO" title="InstructGPT: SFT on demonstrations &#8594; reward model on comparisons &#8594; PPO" srcset="https://substackcdn.com/image/fetch/$s_!RXNT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png 424w, https://substackcdn.com/image/fetch/$s_!RXNT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png 848w, https://substackcdn.com/image/fetch/$s_!RXNT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png 1272w, https://substackcdn.com/image/fetch/$s_!RXNT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0683459-e10e-4c6a-8192-71df98e12836_4400x2748.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p>SFT on human demonstrations</p></li><li><p>Reward model trained on human comparisons</p></li><li><p>PPO against the reward model</p></li></ul><div><hr></div><p><strong>Llama 2 (Jul. 2023) &#8212; multi-stage RLHF</strong> &#183; <a href="https://arxiv.org/abs/2307.09288">paper</a> &#183; <a href="https://www.interconnects.ai/p/llama-2-from-meta">interconnects recap</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IYPf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IYPf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png 424w, https://substackcdn.com/image/fetch/$s_!IYPf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png 848w, https://substackcdn.com/image/fetch/$s_!IYPf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png 1272w, https://substackcdn.com/image/fetch/$s_!IYPf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IYPf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png" width="1456" height="727" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:727,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Llama 2: pretrain &#8594; SFT &#8594; iterative RLHF with rejection sampling and PPO&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Llama 2: pretrain &#8594; SFT &#8594; iterative RLHF with rejection sampling and PPO" title="Llama 2: pretrain &#8594; SFT &#8594; iterative RLHF with rejection sampling and PPO" srcset="https://substackcdn.com/image/fetch/$s_!IYPf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png 424w, https://substackcdn.com/image/fetch/$s_!IYPf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png 848w, https://substackcdn.com/image/fetch/$s_!IYPf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png 1272w, https://substackcdn.com/image/fetch/$s_!IYPf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e89fca4-38bf-4dd9-9bfb-f9dc9bfcedbc_5215x2605.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p>SFT, then iterative RLHF over multiple rounds</p></li><li><p>Each round: rejection sampling &#8594; PPO</p></li><li><p>Two reward models &#8212; separate helpfulness and safety</p></li></ul><div><hr></div><p><strong>Llama 3 (Jul. 2024) &#8212; a complex multi-stage recipe with simpler optimizers</strong> &#183; <a href="https://arxiv.org/abs/2407.21783">paper</a> &#183; <a href="https://www.interconnects.ai/p/llama-405b-open-frontier-model">interconnects recap</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PjC2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PjC2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png 424w, https://substackcdn.com/image/fetch/$s_!PjC2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png 848w, https://substackcdn.com/image/fetch/$s_!PjC2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png 1272w, https://substackcdn.com/image/fetch/$s_!PjC2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PjC2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png" width="1456" height="536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:536,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Llama 3 post-training: reward model &#8594; rejection sampling &#8594; SFT &#8594; DPO, iterated over rounds with best models feeding the next&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Llama 3 post-training: reward model &#8594; rejection sampling &#8594; SFT &#8594; DPO, iterated over rounds with best models feeding the next" title="Llama 3 post-training: reward model &#8594; rejection sampling &#8594; SFT &#8594; DPO, iterated over rounds with best models feeding the next" srcset="https://substackcdn.com/image/fetch/$s_!PjC2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png 424w, https://substackcdn.com/image/fetch/$s_!PjC2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png 848w, https://substackcdn.com/image/fetch/$s_!PjC2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png 1272w, https://substackcdn.com/image/fetch/$s_!PjC2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F284b45b2-9450-40c8-90aa-c72f252a61f9_9429x3473.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p>Per round: reward model &#8594; sample K per prompt &#8594; rejection sampling &#8594; SFT &#8594; DPO</p></li><li><p>No online RL &#8212; the RM only filters; run over 6 rounds, best models seed the next</p></li></ul><div><hr></div><p><strong>T&#252;lu 3 (Nov. 2024) &#8212; simple three-stage post-training</strong> &#183; <a href="https://arxiv.org/abs/2411.15124">paper</a> &#183; <a href="https://www.interconnects.ai/p/tulu-3">interconnects recap</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0oG2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0oG2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png 424w, https://substackcdn.com/image/fetch/$s_!0oG2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png 848w, https://substackcdn.com/image/fetch/$s_!0oG2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png 1272w, https://substackcdn.com/image/fetch/$s_!0oG2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0oG2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png" width="1456" height="508" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/edc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:508,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;T&#252;lu 3: curate prompts &#8594; SFT &#8594; DPO &#8594; RLVR with a held-out eval suite&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="T&#252;lu 3: curate prompts &#8594; SFT &#8594; DPO &#8594; RLVR with a held-out eval suite" title="T&#252;lu 3: curate prompts &#8594; SFT &#8594; DPO &#8594; RLVR with a held-out eval suite" srcset="https://substackcdn.com/image/fetch/$s_!0oG2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png 424w, https://substackcdn.com/image/fetch/$s_!0oG2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png 848w, https://substackcdn.com/image/fetch/$s_!0oG2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png 1272w, https://substackcdn.com/image/fetch/$s_!0oG2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc6b8f1-494b-418e-ae46-afcd4c650a6a_13827x4822.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Curated prompts &#8594; SFT &#8594; DPO &#8594; RLVR (RL with verifiable rewards &#8212; the acronym was coined in this paper).</p><div><hr></div><p><strong>OLMo 3 (Dec. 2025) &#8212; a reasoning update to the T&#252;lu 3 recipe</strong> &#183; <a href="https://arxiv.org/abs/2512.13961">paper</a> &#183; <a href="https://www.interconnects.ai/p/olmo-3-americas-truly-open-reasoning">interconnects recap</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Nqpl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Nqpl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png 424w, https://substackcdn.com/image/fetch/$s_!Nqpl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png 848w, https://substackcdn.com/image/fetch/$s_!Nqpl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png 1272w, https://substackcdn.com/image/fetch/$s_!Nqpl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Nqpl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png" width="1456" height="370" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:370,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;OLMo 3 model flow: Pretraining &#8594; Midtraining &#8594; Long context, then Think / Instruct / RL-Zero branches each SFT &#8594; DPO &#8594; RLVR&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="OLMo 3 model flow: Pretraining &#8594; Midtraining &#8594; Long context, then Think / Instruct / RL-Zero branches each SFT &#8594; DPO &#8594; RLVR" title="OLMo 3 model flow: Pretraining &#8594; Midtraining &#8594; Long context, then Think / Instruct / RL-Zero branches each SFT &#8594; DPO &#8594; RLVR" srcset="https://substackcdn.com/image/fetch/$s_!Nqpl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png 424w, https://substackcdn.com/image/fetch/$s_!Nqpl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png 848w, https://substackcdn.com/image/fetch/$s_!Nqpl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png 1272w, https://substackcdn.com/image/fetch/$s_!Nqpl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7da11f8-95c0-4d19-8481-81311d025715_9718x2468.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p><strong>DeepSeek R1 (Jan. 2025) &#8212; RL as the centerpiece</strong> &#183; <a href="https://arxiv.org/abs/2501.12948">paper</a> &#183; <a href="https://www.interconnects.ai/p/deepseek-r1-recipe-for-o1">interconnects recap</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BDov!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BDov!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png 424w, https://substackcdn.com/image/fetch/$s_!BDov!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png 848w, https://substackcdn.com/image/fetch/$s_!BDov!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png 1272w, https://substackcdn.com/image/fetch/$s_!BDov!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BDov!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png" width="1456" height="663" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:663,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;DeepSeek-R1 multi-stage pipeline: R1-Zero, then cold-start SFT &#8594; reasoning RL &#8594; rejection-sampling SFT &#8594; final RL&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="DeepSeek-R1 multi-stage pipeline: R1-Zero, then cold-start SFT &#8594; reasoning RL &#8594; rejection-sampling SFT &#8594; final RL" title="DeepSeek-R1 multi-stage pipeline: R1-Zero, then cold-start SFT &#8594; reasoning RL &#8594; rejection-sampling SFT &#8594; final RL" srcset="https://substackcdn.com/image/fetch/$s_!BDov!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png 424w, https://substackcdn.com/image/fetch/$s_!BDov!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png 848w, https://substackcdn.com/image/fetch/$s_!BDov!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png 1272w, https://substackcdn.com/image/fetch/$s_!BDov!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97083bed-0200-4ef7-88ef-15e4742c7efe_4870x2218.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The recipe:</p><ul><li><p>R1-Zero &#8212; pure RL (GRPO) on the base, <em>no SFT</em>; used to seed reasoning behaviors for the full run, not a separate product</p></li><li><p>R1 &#8212; cold-start SFT &#8594; reasoning RL &#8594; rejection-sampling SFT &#8594; final RL &#8594; distill to dense</p></li><li><p>A big change in recipes: Large-scale RLVR as the primary driver, SFT to distill and refine RL behaviors</p></li></ul><div><hr></div><p><strong>DeepSeek evolution after V3</strong></p><ul><li><p><strong><a href="https://arxiv.org/abs/2412.19437">V3</a></strong> &#183; Dec &#8216;24 &#8212; SFT + GRPO RL.</p></li><li><p><strong><a href="https://arxiv.org/abs/2501.12948">R1</a></strong> &#183; Jan &#8216;25 &#8212; multi-stage RL; reasoning <em>emerges</em>.</p></li><li><p><strong><a href="https://huggingface.co/deepseek-ai/DeepSeek-V3.1">V3.1</a></strong> &#183; Aug &#8216;25 &#8212; hybrid think / non-think in one model.</p></li><li><p><strong><a href="https://arxiv.org/abs/2512.02556">V3.2</a></strong> &#183; Dec &#8216;25 &#8212; 6 specialists via RL &#8594; SFT distillation &#8594; one mixed GRPO.</p></li><li><p><strong><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf">V4</a></strong> &#183; Apr &#8216;26 &#8212; 10+ domain experts &#8594; MOPD.</p></li></ul><div><hr></div><h3>2026 style recipes!</h3><p><strong>MiMo Flash v2 (Jan. 2026) &#8212; where MOPD started</strong> &#183; <a href="https://arxiv.org/abs/2601.02780">paper</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YlO_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YlO_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png 424w, https://substackcdn.com/image/fetch/$s_!YlO_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png 848w, https://substackcdn.com/image/fetch/$s_!YlO_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png 1272w, https://substackcdn.com/image/fetch/$s_!YlO_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YlO_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png" width="1456" height="577" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:577,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;MiMo Flash v2 post-training: SFT &#8594; domain teachers &#8594; multi-teacher on-policy distillation&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="MiMo Flash v2 post-training: SFT &#8594; domain teachers &#8594; multi-teacher on-policy distillation" title="MiMo Flash v2 post-training: SFT &#8594; domain teachers &#8594; multi-teacher on-policy distillation" srcset="https://substackcdn.com/image/fetch/$s_!YlO_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png 424w, https://substackcdn.com/image/fetch/$s_!YlO_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png 848w, https://substackcdn.com/image/fetch/$s_!YlO_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png 1272w, https://substackcdn.com/image/fetch/$s_!YlO_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95013bf-1d1c-4d4b-9a5f-0c9f67c32d48_3709x1469.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Stages: Stage 1 SFT &#8594; Stage 2 train ~6 domain-specialist teachers (with older style post-training recipes) &#8594; Stage 3 MOPD into a single student.</p><p>First clean articulation of multi-teacher on-policy distillation as the consolidation step &#8212; replaces a single monolithic RL stage with distill-from-specialists.</p><div><hr></div><p><strong>Nemotron 3 Ultra (Jun. 2026) &#8212; two rounds, many teachers</strong> &#183; <a href="https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf">paper</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iU6A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iU6A!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png 424w, https://substackcdn.com/image/fetch/$s_!iU6A!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png 848w, https://substackcdn.com/image/fetch/$s_!iU6A!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png 1272w, https://substackcdn.com/image/fetch/$s_!iU6A!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iU6A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png" width="1456" height="743" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:743,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Nemotron 3 Ultra: two-iteration multi-teacher on-policy distillation&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Nemotron 3 Ultra: two-iteration multi-teacher on-policy distillation" title="Nemotron 3 Ultra: two-iteration multi-teacher on-policy distillation" srcset="https://substackcdn.com/image/fetch/$s_!iU6A!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png 424w, https://substackcdn.com/image/fetch/$s_!iU6A!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png 848w, https://substackcdn.com/image/fetch/$s_!iU6A!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png 1272w, https://substackcdn.com/image/fetch/$s_!iU6A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9fd0c70-d28a-4591-a6b3-e372df128185_2443x1247.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Stages: SFT &#8594; multi-teacher on-policy distillation, run over two iterations, with &gt;10 teachers spanning reasoning, code, math, and agentic domains.</p><p>Novel: multi-round MOPD across different domains &#8212; distill, then re-distill from refreshed teachers.</p><div><hr></div><p><strong>MAI-Thinking-1 (Jun. 2026) &#8212; closer to R1 than V4</strong> &#183; <a href="https://microsoft.ai/news/introducing-mai-thinking-1/">announcement</a></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X8PS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X8PS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png 424w, https://substackcdn.com/image/fetch/$s_!X8PS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png 848w, https://substackcdn.com/image/fetch/$s_!X8PS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png 1272w, https://substackcdn.com/image/fetch/$s_!X8PS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X8PS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png" width="1456" height="280" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:280,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;MAI-Thinking-1: specialist RL climbs &#8594; trace-distillation SFT &#8594; consolidate &#8594; final climb&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="MAI-Thinking-1: specialist RL climbs &#8594; trace-distillation SFT &#8594; consolidate &#8594; final climb" title="MAI-Thinking-1: specialist RL climbs &#8594; trace-distillation SFT &#8594; consolidate &#8594; final climb" srcset="https://substackcdn.com/image/fetch/$s_!X8PS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png 424w, https://substackcdn.com/image/fetch/$s_!X8PS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png 848w, https://substackcdn.com/image/fetch/$s_!X8PS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png 1272w, https://substackcdn.com/image/fetch/$s_!X8PS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a41a711-5a0a-449e-8a0f-375b48021692_1790x344.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Stages: mid-trained base &#8594; 3 specialist RL &#8220;climbs&#8221; (e.g. STEM) &#8594; trace-distillation SFT to consolidate the climbs &#8594; a final RL climb &#8594; MAI-Thinking-1.</p><p>Closer to DeepSeek R1 than to V4 &#8212; multi-stage RL with trace-distillation SFT to consolidate, <em>not</em> on-policy MOPD. Not the only lab without MOPD!</p><div><hr></div><p><strong>Kimi K2.5 (Jan. 2026) &#8212; agentic, multimodal</strong> &#183; <a href="https://github.com/MoonshotAI/Kimi-K2.5/blob/master/tech_report.pdf">paper</a> &#183; <a href="https://www.kimi.com/blog/kimi-k2-5.html">blog</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1VDp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1VDp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png 424w, https://substackcdn.com/image/fetch/$s_!1VDp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png 848w, https://substackcdn.com/image/fetch/$s_!1VDp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png 1272w, https://substackcdn.com/image/fetch/$s_!1VDp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1VDp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png" width="1456" height="779" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:779,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Kimi K2.5 Agent Swarm: self-directed parallel agent orchestration&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Kimi K2.5 Agent Swarm: self-directed parallel agent orchestration" title="Kimi K2.5 Agent Swarm: self-directed parallel agent orchestration" srcset="https://substackcdn.com/image/fetch/$s_!1VDp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png 424w, https://substackcdn.com/image/fetch/$s_!1VDp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png 848w, https://substackcdn.com/image/fetch/$s_!1VDp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png 1272w, https://substackcdn.com/image/fetch/$s_!1VDp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80512ce8-98d2-47b3-b2cc-b5b264e36416_1926x1031.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Stages: text-only SFT &#8594; joint text&#8211;vision RL across coding, vision, reasoning, agentic tasks. (No mention of MOPD.)</p><div><hr></div><p><strong>GLM-5 (Feb. 2026) &#8212; staged RL by capability</strong> &#183; <a href="https://arxiv.org/abs/2602.15763">paper</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3dzI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3dzI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png 424w, https://substackcdn.com/image/fetch/$s_!3dzI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png 848w, https://substackcdn.com/image/fetch/$s_!3dzI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png 1272w, https://substackcdn.com/image/fetch/$s_!3dzI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3dzI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png" width="1456" height="833" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:833,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;GLM-5 pipeline: Base &#8594; SFT &#8594; Reasoning RL &#8594; Agentic RL &#8594; General RL with cross-stage distillation&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="GLM-5 pipeline: Base &#8594; SFT &#8594; Reasoning RL &#8594; Agentic RL &#8594; General RL with cross-stage distillation" title="GLM-5 pipeline: Base &#8594; SFT &#8594; Reasoning RL &#8594; Agentic RL &#8594; General RL with cross-stage distillation" srcset="https://substackcdn.com/image/fetch/$s_!3dzI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png 424w, https://substackcdn.com/image/fetch/$s_!3dzI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png 848w, https://substackcdn.com/image/fetch/$s_!3dzI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png 1272w, https://substackcdn.com/image/fetch/$s_!3dzI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d574f54-c68b-4c76-8720-92c80bb8b96d_5129x2933.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Stages: Base &#8594; SFT &#8594; Reasoning RL &#8594; Agentic RL &#8594; General RL.</p><div><hr></div><h2>Transcript</h2><p>00:00:00 Nathan Lambert: Hello, we are back on a Interconnects conversation. I don&#8217;t really say I do interviews. People criticize me &#8216;cause I interrupt the guests too much. &#8216;Cause I&#8217;m not a good interviewer, but I&#8217;m here to entertain people. Um, this is also fun for me because I&#8217;m trying to make, like, a post-training course, and it kind of fits as, uh, in the advanced end of this.</p><p>So it&#8217;s kind of a crossover between Interconnects content and other stuff that I&#8217;ve been spending my time on this summer. So I&#8217;m happy to welcome Finbarr back. I think... Are you the first return guest? I haven&#8217;t checked.</p><p>00:00:37 Finbarr Timbers: Oh, wow.</p><p>00:00:37 Nathan Lambert: Um, Finbarr and I worked on this sort of post-training recipe stuff for a while at AI2. Um, I left recently. This is one of Finbarr&#8217;s last days at AI2. It&#8217;s already been announced. It&#8217;s not a spoiler here. So we&#8217;re gonna kind of reflect on some things on building post-training recipes for OLMO. Um, then we have a little, like, review slide deck and notes on the kind of state and evolution of frontier post-training recipes over time, which is pretty interesting because there&#8217;s, what is it, like two to four kind of canonical recipes that there has been.</p><p>So it&#8217;s kind of interesting when you see the field converge on something new, which it&#8217;s doing right now with multi-teacher on policy distillation. For some reason, that&#8217;s a bit of a mouthful. It is a long acronym. And then we&#8217;ll just kind of end with various discussion points on post-training and what we&#8217;re up to. So, happy to give you the floor if you have any hot takes you wanna start with to get people to, draw people in. Otherwise, I think, uh, I&#8217;m excited to kind of reflect on this, &#8216;cause I know you&#8217;ve been reading a ton of papers recently and kind of prep, laying some of this groundwork.</p><p>00:01:43 Finbarr Timbers: Well, yeah. I mean, today is my last day at AI2, so it- it&#8217;s ki- it feels very appropriate to be, to be talking to you as you&#8217;re the one who recruited me to AI2. So, uh, yeah, that&#8217;s pretty special, and it&#8217;s great to be, uh, yeah, the, the first repeat guest. I feel honored, uh, to be back on. So yeah, thanks, uh, for having me.</p><p>00:02:03 Nathan Lambert: Yeah. Do we wanna start with OLMO? I think that-</p><p>00:02:05 Finbarr Timbers: Sure</p><p>00:02:06 Nathan Lambert: ... people... I think I, uh, need to do this carefully, but I&#8217;ve talked about OLMO-3&#8217;s post-training many times to people. I haven&#8217;t done this in a very direct way on the podcast, but I would say that post-training OLMO-3 to make this reasoning model was a major accomplishment for many individuals to do this. But also, the complexity of what we were doing was pushing against the limits of AI2&#8217;s organizational capacity, and a lot of modern post-training is, like, your ability to wrangle compute data into a work stream.</p><p>And in order to do that in a complicated way, you really are wrangling an org chart. And that&#8217;s like part of why it&#8217;s like OLMO-3 was, by its nature, pretty late as a reasoning model. It was, like, a pretty rigid reasoning model, and that&#8217;s, like, partially reflected in the recipe being pretty simple. But then when you, like, compare it to all these new recipes with tool use and multi-teacher distillation and all of this, it&#8217;s just like a, a, a fork in the road where it&#8217;s like you could do this very simple thing and make a strong recipe, but it is not representative of what all the frontier labs are doing.</p><p>And I think that that kind of fork in being able to say that things are similar happened kind of after Tulou-3, where Tulou-3, I think, was also much simpler with this three-stage SFT-DPO RL recipe. But that simpler recipe was probably closer in outcome to what the labs are doing, but now doing that sort of three-stage recipe for a reasoning model, and especially a tool use, like, agent model, just doesn&#8217;t really apply. And that&#8217;s the point. That&#8217;s why I think the point of this podcast is to be like, what are the, what are the way, what are they doing to make these, like, true frontier models, and then shed some light on how it contrasts to the more a- like, open academic ones.</p><p>00:03:56 Finbarr Timbers: Well, actually, I think that&#8217;s interesting. What was the proce- so, you know, I, I only, um, came around for OLMO-3. I wasn&#8217;t around for the earlier, um, versions. What was the process like to go from Tulou-3 to OLMO-2? Because, like, y- just looking on, on Archive, um, I think Tulou-3 came out in November of &#8216;24, and then OLMO-2 came out in December of, of &#8216;24.</p><p>00:04:22 Nathan Lambert: We just applied the recipe.</p><p>00:04:24 Finbarr Timbers: Yeah. I, I mean, so, so I think that actually, like, yeah, and then, you know, um, DeepSeeker-1 came out in January, end of January &#8216;25, and, you know, OLMO-3 was then released in October. Was it October or November of &#8216;25? Like, I think-</p><p>00:04:39 Nathan Lambert: I think November.</p><p>00:04:41 Finbarr Timbers: Yeah, November. Yeah, right. It was November. So it&#8217;s-</p><p>00:04:43 Nathan Lambert: It was like do or die with Thanksgiving.</p><p>00:04:45 Finbarr Timbers: I remember that. Uh, yeah, &#8216;cause Canadian Thanksgiving had, had already happened-</p><p>00:04:50 Nathan Lambert: Yeah</p><p>00:04:50 Finbarr Timbers: ... which, yeah, I was happy. Um, but, uh, like, like I think it was, sure, maybe it was late, but I think it was only late by a few months. Like, it&#8217;s, it&#8217;s actually, like, you know, if I think of my past experience with model turnaround times, like a nine-month model turnaround, you know, from R1 coming out, like that&#8217;s actually, that&#8217;s not bad. I think, you know, something like six months would&#8217;ve been nicer, but-</p><p>00:05:12 Nathan Lambert: I, I think it&#8217;s slow &#8216;cause we didn&#8217;t re- it would be fast if we had rebuilt the R1 recipe. But what we did was we, like, ported reasoning into our existing recipe-</p><p>00:05:21 Finbarr Timbers: Yeah. Okay</p><p>00:05:22 Nathan Lambert: ... which is a simpler task, but has, like, a lower ceiling, in my opinion. Where it&#8217;s like the DeepSeek and the newer style recipes, which we&#8217;ll talk about, I think they just have a much higher ceiling in how much you can keep hill climbing them. Or they&#8217;re just, like, more prescri- more pedagogical of what the frontier is doing. Like, for the size models that OLMO was, which was like 7 to 30B, I&#8217;m not sure that doing this DeepSeek style RL first recipe is actually useful.</p><p>00:05:52 Finbarr Timbers: Uh, well, I, yeah, I think that&#8217;s a good point. And I mean, I think that&#8217;s really reflected in what we see in the research where you s- you know, you obviously you see the big, uh, the step change and you know how quickly things are improving When, you know, R1 comes out. So, like, I think that a great point, and it really does seem to saturate, or to, to not saturate, sorry, with, with compute. Um-</p><p>00:06:11 Nathan Lambert: Yeah. Um, shall we just do the slide deck? We&#8217;re throwing around, like, recipe-</p><p>00:06:15 Finbarr Timbers: Sure. Yeah, let&#8217;s do it</p><p>00:06:16 Nathan Lambert: ... names. Like, I feel like it might be useful to just do it because a lot of people probably want to follow but don&#8217;t exactly know. I&#8217;m, I&#8217;m gonna share, I&#8217;m gonna share a screen. So people listening, it might be useful to either, you can pull this slide deck up on your phone and click through it. It&#8217;s not super information dense, but you can also just watch it on YouTube. All of this will be linked.</p><p>Generally, this is just like a quick survey on how frontier recipes have evolved. We&#8217;ll go through the history quickly and then talk about what is currently happening and kind of probably interleave the old mode discussion we were having. Uh, okay. There&#8217;s a bunch of canonical recipes we&#8217;ll talk about. This is where I got the two to four number. I think the recipes are like InstructGPT, which is what coined the initial RLHF with this like three-stage idea, which took a while to get people to move on from, which was like SFT reward model and RL.</p><p>And I see as like Llama 3 and 2.3 as kind of practical implementations of that with, with other tricks of the trade. So those two could potentially be merged together. It&#8217;s just like kind of pre- and post-ChatGPT moment. And then the two most recent canonical recipes that we&#8217;ll cover in this I would say are like DeepSeek-R1, which is the shift to doing like reasoning focused and bigger RL stages than this kind of SFT focus from before, and then NeMo Flash and some of the new models from 2026 which add this distillation element.</p><p>00:07:42 Finbarr Timbers: Well, and, and I think it&#8217;s worth pointing out too that it&#8217;s not just NeMo Flash, like it was kind of a consistent theme. Like you saw this with DeepSeek, th-they referenced it in, uh, the V3 paper and then it&#8217;s, you know, it&#8217;s Qemi K 2.5, it&#8217;s GLM 5. Like it&#8217;s all of these papers, you know, start talking about this specialist, um, RL stage.</p><p>00:08:03 Nathan Lambert: Yeah. I think there&#8217;s a debate on how we draw it and whether or not distillation is... If you&#8217;re, if you have distillation as a technique, as a key milestone, then they were, the Xiaomi was the first and, but it&#8217;s kind of a march over time where you kind of see them change, and we&#8217;ll, we&#8217;ll go through this. I don&#8217;t, I don&#8217;t need to interrupt.</p><p>00:08:23 Finbarr Timbers: When you say distillation, I do think it&#8217;s important to distinguish between the straight up like, you know, distillation of the leading closed models and, you know, distillation of these domain specific models where, you know, I, I, I suspect that the, you know, the, the Chinese labs are doing both.</p><p>00:08:41 Nathan Lambert: Yeah.</p><p>00:08:41 Finbarr Timbers: But, you know, a lot of what they&#8217;re do, you know, but a, a lot of what they&#8217;re doing is this, um, training these domain specific models like, you know, a math model, a coding model, uh, you know, logic model, whatever, and then distilling those models back in and not just distilling from... So when we&#8217;re talking about distillation, it&#8217;s not just distilling from the leading closed models.</p><p>00:09:01 Nathan Lambert: Yeah. It&#8217;s a pain. I agree. The distillation term is horribly overloaded. Um, there&#8217;s a review slide. Do we need to review multi-teacher on policy distillation? It might be too complicated to need to do it. We could come back to it. I think I kind of want to just go through the actual models, and then we could use the supporting slides as needed. Um, this famous InstructGPT three-step thing, I think many people have heard of it, but this is what constituted post-training at the time of ChatGPT coming out, so it&#8217;s kind of important grounding of this human supervised SFT data, mostly human supervised preference rankings to make a reward model and then do RL on that, and the model gets better.</p><p>And it&#8217;s pretty interesting how all of these have been kind of phased out, at least in terms of what we know openly, where they&#8217;re, we don&#8217;t use that much human demonstration data for SFT. There&#8217;s likely some human preference data still in the loop, but I would guess that synthetic has a much bigger role, and there are reward models, but they&#8217;re like not the cl- key RL target anymore. So in four years, most, almost all the canonical pieces have been moved on. And like this evolution is kind of within there. I think the early models after InstructGPT, like Llama 2, um, even Llama 3, these are pretty similar, which is like you&#8217;re starting to break down this recipe with different tools like projection sampling, DPO, some increased iterations. I think increased iterations is just that there was more incentive to squeeze more out of the models, and they just like broke things down more, where InstructGPT seemed like a bit more open-ended research where this kind of cleanness was fine. So-</p><p>00:10:48 Finbarr Timbers: Well, I think that&#8217;s interesting, uh, with respect to how much everything has scaled, uh, right? Because, you know, InstructGPT was before ChatGPT was, was released, and so, you know, it&#8217;s something, like just the complexity of what was done is that which a small team or even a single team could do. But then when you start looking at, you know, Llama 3, like it just starts to be a more complicated process and, you know, where you start to have a lot more, you know, specialized data and there&#8217;s, you know, a lot more, you know, room for scale and for kind of money and complexity be poured in.</p><p>00:11:25 Nathan Lambert: Yeah. It&#8217;s like, uh, both for-profit and nonprofit efforts to do post-training want me to advise them, and I&#8217;m like, &#8220;I don&#8217;t really know how I&#8217;m gonna give you advice unless I&#8217;m spending twenty hours a week look, understanding the details of your recipe,&#8221; &#8216;cause it&#8217;s like, well, I can&#8217;t really give you a one sentence thing of do X without understanding all the complexities of the model and the post-training process that go into it. Which makes it, like makes it hard from kind of like a transparency point of view. Even if it&#8217;s fully detailed, it&#8217;s definitely still hard to modify and study.</p><p>00:12:00 Finbarr Timbers: Absolutely.</p><p>00:12:02 Nathan Lambert: So then like two through three in AI2, a lot of this was we&#8217;re trying to beat the results of this Llama 3 post-training, which is pretty complicated, but we don&#8217;t have the ability to scale the organization as far. So I, I, I think that&#8217;s a big reason why the actual workflow is a lot simpler, where we have three clear stages that are doing slightly different things, and they build on each other. And that&#8217;s like... It&#8217;s never stated very explicitly in these papers on like how the org chart impacts the recipe, but I would, I, I think it&#8217;s a very strong signal within the, at least the delta between the fully open work and the kind of partially open work that you get from industry.</p><p>00:12:43 Finbarr Timbers: Yeah, absolutely. And, and I think especially as we&#8217;ll see with the domain-specific models, like that&#8217;s like really clear, like something where you could really easily scale up your org chart to-</p><p>00:12:54 Nathan Lambert: Yeah</p><p>00:12:54 Finbarr Timbers: ... build that up.</p><p>00:12:56 Nathan Lambert: Yeah. And I threw Olmo 3 in after this, after the two through three slide, mostly just to show that the recipe was so similar to two through three, and the org chart hadn&#8217;t really changed. Like we didn&#8217;t have more ability to scale, and like there was a, a little bit of separation between the model types, between like the think and the instruct models. But like without a major reinvent- like a major org change, it was just kind of stuck in this and do the best you can with it.</p><p>00:13:22 Finbarr Timbers: Yeah. Absolutely.</p><p>00:13:23 Nathan Lambert: Be- because like the real big change was this with DeepSeeker-one. They, I had never seen this plot before, but they had this plot, maybe they added it for the nature version of the paper, where they kind of show their recipe, where they like take the base model, they do RL zero, and then they sample from the RL zero to like filter prompts, and then they use that as SFT. This is like going through this. They use that as SFT for the next version of the model to create like a development internal RL DeepSeek-R1, and then they do this like repeated sampling to train multiple RL versions and kind of distill, distill in the sense of, of clarify and refine the reasoning behavior of the model before going through the final pipeline, which again is a mix of, um, reasoning and non-reasoning SFT into a bigger RL run. And-</p><p>00:14:11 Finbarr Timbers: Well, and I think this is really interesting because it starts to show, I mean, first of all, the, the complexity here. We&#8217;re starting to use, um, yeah, like synthetic data as this primary input here, but it&#8217;s not just like, you know, it&#8217;s trying to elicit, you know, specific behaviors, and it&#8217;s this kind of like industrial process, um, instead of like this, you know, it&#8217;s not as much of an elegant research recipe. It&#8217;s more like, you know, we train a model, and then we use it as best we can, and we keep iterating. Um, but I think the other thing that&#8217;s interesting is, is we&#8217;re starting to see here the SFT serving as the cold start. First of all, where, where that&#8217;s, you know, I think before SFT was more of like a generally useful stage, whereas here its, its primary purpose is this, this cold start for RL.</p><p>And then the other interesting bit is, you know, DPO, uh, starts to disappear at this point from the leading recipes. I mean, Olmo 3 still does it, but you know, basically everyone else does away with it and just, you know, has the preferences included, um, as in, as a reward model or, you know, at so- at some way, um, in the reward bit of the RL stage. And so that&#8217;s a really interesting change, where the, the supervised part of post-training is just, you know, massively deprioritized.</p><p>00:15:27 Nathan Lambert: Yeah. So my hypothesis for the dropping of DPO on these models is that, uh, as, as you&#8217;re doing like a cleaner recipe, essentially the need falls away. Versus if you look at Olmo, which is taking tons of potential gains by refining your model on outputs of strong open weight models, like largely Qwen and DeepSeek is the training data for the SFT of Olmo 3. Uh, and like the delta between that SFT data and the base model is still pretty big in the probability distributions. So DPO kind of helps further refine and clean up that distribution in a way that kind of has very rough edges. And but when you have a more refined, like industrial process on post-training, th-that will, that potential benefit will be harder to gain. Something interesting that I didn&#8217;t fully con-confirm before this is, for example, NVIDIA used to also be on this DPO train with their smaller Nemotron models.</p><p>And, and I would guess that potentially like D- Nemotron Ultra would not. But it&#8217;s, and, and that&#8217;s because they&#8217;re at much further down this development tree and using on pol- like these more on policy methods for creating the SFT data. And their model, I would guess, will become kind of more robust out of distribution and like have weird, less weird rough edges before because of it. So that&#8217;s kind of my hypothesis on DPO, and people that use DPO will be looked down upon. But it&#8217;s like if you&#8217;re trying to bootstrap a recipe off the ground and just take gains where you can, I still think it&#8217;ll work for a lot of people in a kind of compute efficiency standpoint.</p><p>00:17:05 Finbarr Timbers: Yeah. I mean, I think generally, uh, there&#8217;s something interesting with the, the preference tuning that, yeah, like maybe, um, it isn&#8217;t being given the proper, um, respect that it deserves. &#8216;Cause o-one of the interesting bits about the Nemotron 3 super paper was that they saw pr- they, they do a, a traditional RLHF stage in their RL, which has also, you know, fallen with fashion and development, and they see pretty massive gains with it. So I think some of these changes are more, you know, driven by what&#8217;s in fashion rather than perhaps like a fully rigorous, you know, set of ablations.</p><p>00:17:41 Nathan Lambert: It&#8217;s very remarkable to me that the preferences loss function can do so much for these models. Like the models have so much potential there, and it&#8217;s just, it&#8217;s really a contrastive loss on pretty granular feedback. And they learn all sorts of things. Like they&#8217;ll, they&#8217;ll get better at math and code, or their reasoning strategies will be refined. And so I, I... That&#8217;s remarkable to me. I think there will still be funny research on like using preference- Base losses with verifiable outputs. Like, I, I think all this would work. Like DPO on verifiable rewards and stuff like this, it&#8217;s just kind of intellectually less appealing.</p><p>00:18:19 Finbarr Timbers: Yeah. Well, I think that&#8217;s, uh, you know, that&#8217;s where I thought that the, uh, delta learning, um, hypothesis style, uh, DPO, like what Olmo-3 did, where you, um, where the, the preference, you create these synthetic preferences by having like strong, by like bigger and smaller models of the same family, like is where you get your preferences from. I thought that was a really interesting signal because it, it seems really analogous to some of the work, some of the guidance stuff that we see in diffusion models, like how you have the classifier-free guidance, which has something similar, and there, there were very similar results there, which showed that you could have the--</p><p>But like one signal they used was further along in training versus earlier in training models as like, uh, a source of, of signal that you could guide along. And, and that worked quite well. And so I suspect that these signals, um, for, for preferences in that way, like that they could actually be more robust, but because, you know, some of the largest labs don&#8217;t have to do that, perhaps we&#8217;re not citing them as much.</p><p>00:19:18 Nathan Lambert: Yeah. Or they don&#8217;t tell us. Like, to continue this, it&#8217;s kind of cool to look at-- So the DeepSeek models have kind of gone through this, what I would call like l- closer to Llama recipes to DeepSeek-R1, which is d- like most definitively the canonical recipe for reasoning models, and then continue to change closer to this multi-teacher format. So if you look at the VC-3.3 paper, um, before R1, they do something remarkably similar to two to three type thing, where they have a mix of SFT and then they use it ver-- like this RL on verifiable rewards. They didn&#8217;t call it that, or their paper wasn&#8217;t out at the time. And so they did this before R1 came out, which was just kind of a less reasoning-focused models and used the same tools but with a different ratio of implementation weight.</p><p>00:20:07 Finbarr Timbers: And, and what&#8217;s interesting is that this comes out basically at the same time as two to three, and it&#8217;s a very similar two to three and Olmo-2. It&#8217;s a very similar recipe, just done with more complete.</p><p>00:20:16 Nathan Lambert: Yeah. Yeah. And then we have this R1, which we&#8217;ve just talked about at length in January, which is a month later. They have a few more releases through this. They have some updates to their V3 and R1 models, which have dates, which are largely the same recipe. And then the next documented change in their recipe was V3.1, which is when they merged this thinking and non-thinking into one model, which everybody that does this says, has said that it has been hell to train in. But you kind of need it from a serving perspective, and it&#8217;s obvious that long term, at least obvi- it&#8217;s obvious to me that long term all the models will be reasoning models, and you&#8217;ll just have reasoning models that are very efficient based on the gains that are there.</p><p>So this is kind of a needed change that they made. And then in December of 2025, they released V3.2, which is when there&#8217;s kind of meaningful changes to their recipe, and they&#8217;re talking about this expert creation with separate mini recipes, and then using that within their kind of R1 data process to do SFT data and then like a big RL run at the end with GRPO. So it took about a year for this, uh, like kind of evolution of the R1 style recipe to land in their models. And I think this, this is like a very big complexity step that isn&#8217;t represented in something like Olmo-3, and it&#8217;s kind of where you can see a fork in the recipes over time as like they, it, they become way more industrial and scaled at these frontier labs.</p><p>00:21:46 Finbarr Timbers: Yeah. And I think, you know, another one good thing here, just from a historical note, is that I think it was with the O3-24 release where they updated the original V3 paper. So, you know, V3 comes out before R1, then R1 comes out, and then after R1 comes out, they actually go back and update the V3 paper, maybe getting ready for the nature submission or, or, or something.</p><p>00:22:07 Nathan Lambert: Yeah.</p><p>00:22:07 Finbarr Timbers: Um, and they make a reference there to say like, &#8220;Oh, you know, something you could do is you could train these domain specialist models and then combine them.&#8221; Uh, and then, you know, that later becomes kind of what, you know, the more of a priority as they talk about in V3.2.</p><p>00:22:21 Nathan Lambert: That&#8217;s a fun note. Yeah. And then most recently in April 26th is this V4 model, which has even more experts. They add this new loss function for multi-teacher on policy distillation, which I said follow Jiaoli. And this is kind of a microcosm of the arc that the whole industry went through, at least the people who share what their post-training details are, of realizing how core RL is, changing the recipe around scaled RL, and then figuring out how to kind of scale to more domains in the scaled RL format without just like grinding to a halt in operational complexity.</p><p>00:22:58 Finbarr Timbers: Yeah.</p><p>00:23:00 Nathan Lambert: So then kind of the next stage of this is these, what I call twenty twenty-six style recipes, which are all these models that are doing this multi-teacher, um, infusion of knowledge. And then some of them are using on-policy distillation and some are not. It&#8217;ll be one of the key things to see is like how crucial is this on-policy distillation to really keeping up at the frontier. So the paper that kind of, that named this term was the MimoFlash V2 paper. I think the model was released in December and the paper in January, which a lot of things will look similar to this, um, kind of RL, large RL style recipe. But with this large RL run is more, is where the on-policy distillation comes in. So for, I c- this is probably a better time to explain. I have this great, great little feature.</p><p>So this is like the summary of what multi-teacher on policy distillation is. Generally, it fits within an RL framework where you have the model you are training, the, like the general model, sample its own trajectories, and then you route the trajectories to various expert models you have trained. And each kind of sample is trained with this distillation KL loss to match the tokens of that expert. And People have, multiple models have shown that this type of supervision is really useful for the models. You could combine it with other RL losses, such as verifiable rewards, which for example, Sasha Rush gave a good mini spiel on that and how they use that with Composer, which is a, a video that I really recommend people watching as well. But the, the key of it is that it is a different loss function, but it plays very nicely in the RL frameworks that people are already using. So they use these teachers-</p><p>00:24:45 Finbarr Timbers: Just RL, like it&#8217;s, it&#8217;s, like if you-</p><p>00:24:47 Nathan Lambert: Yeah</p><p>00:24:47 Finbarr Timbers: ... actually implement it, you know, I&#8217;m talking with some of the people at AI2 about implementing it now. And it&#8217;s like you take your RL setup, and then you just, you know, you, you have some very, your, uh, set of tweaks on the, the learner to actually implement this. So it&#8217;s quite straightforward.</p><p>00:25:02 Nathan Lambert: Yeah, so this is a fancy diagram that makes it more complicated than it needs to be, but it also a very nice diagram, which shows the various, um, domain teachers that they have, search agent, code agent, math, reasoning, safety, and how they put these together. And the, the experts are used both for SFT data and then this final supervision. And the recipe for the experts would look something like this DeepSeek recipe, which is complicated on its own, which is like make a very good reasoning model that is good at one thing.</p><p>00:25:29 Finbarr Timbers: Well, and I think it is complicated, but it&#8217;s also like if you, if you think about being the actual researcher like working on it, it&#8217;s like, you know, you have a base model, and then you have an RL set up, and you know, you&#8217;re just constantly updating both and then rerunning RL. So, you know, the, the most complicated like, uh, part of it is just, you know, writing down the history and tracing everything. But it&#8217;s kind of like a very natural, organic way, uh, for the r- the RL to evolve through, you know, iterative experimentation.</p><p>00:25:57 Nathan Lambert: Yeah. So like once you have a recipe, you&#8217;re progressively tinkering with each part, and it&#8217;s, it&#8217;s fairly stable, but it&#8217;s hard to rebuild from scratch. So like we&#8217;ll see how, see how long the recipe shape lasts, but it&#8217;ll probably be order of years. Um, another big one in this like also shared a lot of details on this on policy distillation approach was Nemotron-3 Ultra, which is obviously exciting to me to have a, like a US-made model that is very strong performance, and NVIDIA released a lot of datasets with it.</p><p>But they, they also talked about a lot of their very n- n- like implementation details of what was hard with on policy distillation. I, like I have notes somewhere on this. They do this thing where they have two rounds of on policy distillation, as they found it to be better to integrate some teachers one after another. And the paper has a lot more details. I&#8217;ve, I, I don&#8217;t wanna go scroll through the paper, but we could also do this. Did you have any o- other impressions? Like I have the, we have this other doc we can pull up that-</p><p>00:27:01 Finbarr Timbers: Oh</p><p>00:27:01 Nathan Lambert: ... also you might have had other details on it.</p><p>00:27:03 Finbarr Timbers: Yeah. Well, I think something else, um, that, that is worth, um, you know, contrasting the, the paper to is the Nemotron-3 super paper. &#8216;Cause in the Nemotron-3 super paper, they had a similar complicated recipe, but they did multiple rounds of RL. Like there they had three rounds of RLVR, followed by a round of, um, software engineering RL, and then followed by an RLHF stage. So it was, it, it was really interesting to see them go from doing that, like, you know, one of the most complicated, um, RL setups or in terms of, you know, successive stages, uh, that I&#8217;ve seen. To then, you know, you know this setup where it&#8217;s still complicated, but it&#8217;s a lot, um, you know, it&#8217;s a lot con- conceptually a lot simpler.</p><p>00:27:54 Nathan Lambert: Yeah. I, I pocket the paper up. It&#8217;s gonna be hard for me to... I, like I had highlighted a few details. The, the interesting parts are kind of around the, um, various NVIDIA details on all the teachers. There&#8217;s just so many details in their paper on training-</p><p>00:28:10 Finbarr Timbers: Yeah</p><p>00:28:10 Nathan Lambert: ... all the teachers. I think, okay, so I have some of it. I have some of this up. It&#8217;s like I have an interesting quote that&#8217;s like, &#8220;One key finding from our trials of doing on policy, multi-teacher on policy distillation is that teacher models trained with substantially different training pipelines cannot be effectively combined through a straightforward on policy distillation merge, resulting in suboptimal performance.&#8221; So it&#8217;s like they&#8217;d have to do some cross teacher alignment, um, to make sure that they&#8217;re actually similar, which I feel like could become a whole, uh, organizational nightmare. It&#8217;s like they say, &#8220;We hypothesize that when the teacher and student are trained on different SFT data, they acquire different reasoning behaviors and induce different output distributions. This distribution mismatch can cause student-generated trajectories to be out of distribution for the teacher, result- reducing the quality and reliability of the supervision- supervision signals provided by the teacher.&#8221;</p><p>00:29:00 Finbarr Timbers: Yeah, that&#8217;s interesting actually because there was a paper, uh, I, I can&#8217;t remember the name of it, but there was a paper that I read, um, recently which claimed that what you need to do is constantly... So, so you know, you know, one thing you could do, which was kind of the, the obvious thing to do, is you, you take your base model, right? You do, um, whatever general SFT that you&#8217;re doing, and then you take, you do, you know, a bunch of RL, you train domain-specific agents, you train them, you know, all the way until they&#8217;ve converged or until you&#8217;ve run out of money.</p><p>Uh, and then you take these final experts, and then you do some sort of, you know, on policy distillation to combine them into your, your final model. Um, but with the paper, and I&#8217;ll, I&#8217;ll try to find it and then give it to you, um, see if we can share it. What they claimed was that you need to, um, instead of using the converged model, you need to do it in like successive stages with like the in-progress model. So if, you know, you train your RL for like a thousand steps, you need to, you can&#8217;t use the, you know, the thousand step checkpoint to, for the on policy distillation. You have to do it in stages, and first use the, you know, two hundred and fifty step checkpoint and the five hundred checkpoint and, you know, gradually bring that base model like up to speed or else there&#8217;s gonna be too much divergence, and the, the KL divergence will just be like too, um, too distinct-</p><p>00:30:17 Nathan Lambert: Yeah</p><p>00:30:18 Finbarr Timbers: ... to learn from.</p><p>00:30:19 Nathan Lambert: Yeah. So essentially the last state-- sentence in this paragraph I had read most of is literally like, &#8220;We encountered this issue in practice because the teacher and student models were developed in parallel.&#8221;</p><p>00:30:29 Finbarr Timbers: Yeah.</p><p>00:30:29 Nathan Lambert: It&#8217;s like they&#8217;re like, &#8220;This is a problem because of it&#8217;s, like, hard to do everything at once.&#8221; Which is w- this is the type of thing where having research in it would be so great, and I think NVIDIA could release some of the teachers so that people could just like-</p><p>00:30:45 Finbarr Timbers: Yeah. That&#8217;d be great</p><p>00:30:45 Nathan Lambert: ... if you have the teachers and you have the intermediate model stage, you could do the problem of, like, just studying multi-teacher on policy distillation from the starting point and understanding the training dynamics.</p><p>00:30:57 Finbarr Timbers: Yeah.</p><p>00:30:57 Nathan Lambert: Which is the type of thing we would want to do at Oldo. We just haven&#8217;t scaled our recipe to this point yet.</p><p>00:31:03 Finbarr Timbers: Yeah, absolutely.</p><p>00:31:04 Nathan Lambert: So I will keep encouraging NVIDIA to do this.</p><p>00:31:07 Finbarr Timbers: That&#8217;d be great. NVIDIA-</p><p>00:31:08 Nathan Lambert: I think, uh-</p><p>00:31:08 Finbarr Timbers: ... listen.</p><p>00:31:10 Nathan Lambert: They, they listen. The other side of things is a bunch of models released in 2026 that do not do this multi-teacher on policy distillation, and they also don&#8217;t do nearly as many teachers. So I would say that this, like, Microsoft model, which I don&#8217;t say this as a diss, it&#8217;s, like, hard to get a new team off the ground, is they went for a simpler approach to try to get a solid model, and it has three more general experts combined w- via SFT and then, like, a longer RL run. So it looks a lot more like DeepSeeker one, but I suspect that what they will do next is make finer grain teachers and see if they need to switch to on policy distillation.</p><p>00:31:48 Finbarr Timbers: Yeah. And I think, you know, in one of our, um, group chats, you described the MAI thinking model as a conservative recipe. A-and I think that&#8217;s a really good description of it. Like they, you know, the, the team came up with this conservative recipe, and then I think that they did a really great job of actually executing on it. &#8216;Cause I think, you know, if you try to make too many changes at once, it&#8217;s really easy for the recipe to collapse under its own complexity, and I&#8217;ve seen that a bunch of times, you know, across my career.</p><p>Try to make too many changes and, you know, it all goes poorly. So I thought that was, um, a really good choice on their part. I, I also think that, uh, it&#8217;s not super clear to me, may-maybe you&#8217;ve seen some papers on this that I haven&#8217;t seen, but it&#8217;s not super clear to me how well the trace distillation SFT does or, you know, h- how much better on pols- online policy distillation is versus the trace distillation SFT.</p><p>00:32:41 Nathan Lambert: Yeah. It&#8217;s like what&#8217;s, what is the relative magnitude in the final performance?</p><p>00:32:45 Finbarr Timbers: Yeah.</p><p>00:32:45 Nathan Lambert: So the Nemotron Ultra paper has a table on how far the on policy distillation goes relative to the teacher, and they also have the starting point. So I guess that&#8217;s a potential way to do this. Here, I could, I could just pull this up. Let me switch.</p><p>00:33:00 Finbarr Timbers: Oh, sure.</p><p>00:33:04 Nathan Lambert: So I, I had this open, but in a different tab. Okay. Here&#8217;s, here&#8217;s this paper. This is page twenty-seven is which the paragraph I just read, and then it also has this kind of-</p><p>00:33:17 Finbarr Timbers: Oh, fascinating</p><p>00:33:18 Nathan Lambert: ... is it a great table. I spent a while looking at this earlier. So essentially, it&#8217;s like where they get after SFT-</p><p>00:33:24 Finbarr Timbers: Wow</p><p>00:33:24 Nathan Lambert: ... on each of the benchmarks on the general model. And then I think... Okay, so the sort of gains over the RLVR student recovery of the specialty student. So I need to make sure... Okay, so it denotes the initial student checkpoint, where RLVR denotes the s- initial student checkpoint, and then the multi-teacher on policy distillation. So I&#8217;m not sure what this SFT column can figure out, but you could see the kind of like where the teacher is relative to on policy distillation. I think this is like the closest information we have on the relative performance gains.</p><p>00:33:59 Finbarr Timbers: Yeah. That&#8217;s fascinating because the DeepSeek, I forget which one, maybe it was V3.2 paper claims or, or maybe it was, um, R1 actually claims that you can domain-specific... That, that, you know, doing the general stage, uh, captures the performance, uh, of it. But, you know, that, that doesn&#8217;t really seem to be... A-a-and yeah, a-a-and then so, you know, doing the domain-specific distilling in, and then doing a general stage on top of that captures the original performance. But that doesn&#8217;t seem to be the case here. Like, you know, the, the gap maybe isn&#8217;t huge, but there is still, most of the time, there&#8217;s a pretty big... There, there&#8217;s like, you know, a significant gap, even if it&#8217;s not huge. So that&#8217;s really interesting.</p><p>00:34:42 Nathan Lambert: Yeah. I wish this table and text was clearer. It&#8217;s like I literally can&#8217;t fully parse it. It&#8217;s like RLVR denotes the initial student checkpoint, and then OPD denotes the checkpoint after first and second iterations. It&#8217;s like, what is the checkpoint that was used at the start of on policy distillation?</p><p>00:35:01 Finbarr Timbers: I think it was the RLVR one, so that they do a general SFT stage, and then they do an RLVR stage that covers the non-teacher, the, the areas that where they don&#8217;t have specialized models. Then they do MOPD.</p><p>00:35:15 Nathan Lambert: Yeah. And then that makes sense with this recovery rate, which is like final model minus RLVR, which would be like the gains for the OPD relative to the teacher minus RLVR, which would be like what gains you needed to still cover.</p><p>00:35:31 Finbarr Timbers: Yeah.</p><p>00:35:32 Nathan Lambert: And like what, what gains the teacher could potentially give you. So more research like this. Happy to see some of it a- out there. I&#8217;m gonna switch back.</p><p>00:35:43 Finbarr Timbers: Yeah. Something I found interesting about the, um, the, uh, both the Nemotron papers and then the MAI thinking paper is that they don&#8217;t talk as much about some of the more detailed, um, post-training decisions that have shown some pretty strong gains in, um, some of the other papers. Like I, I believe it was GLM five where they talk about doing a difficulty curriculum and a difficulty filtering stage.</p><p>00:36:11 Nathan Lambert: Yeah.</p><p>00:36:12 Finbarr Timbers: And that&#8217;s just not something that&#8217;s really talked about in these other papers. They&#8217;re saying they, they don&#8217;t, you know, uh, I think it was QEM 2.5 used a temperature. It&#8217;s kind of funny. So QEM K 2.5 and GLM five both have temperature schedules, uh, and they both claim the exact opposite thing. So one of them says you have to start with a high temperature and go low. The other one says you have to have a low temperature and go high. And, uh, y- I don&#8217;t know. And then so, you know, you don&#8217;t see that discussion, uh, I, I don&#8217;t think in Some of the other papers, which is kind of interesting</p><p>00:36:40 Nathan Lambert: Yeah. I, I still think the Chinese labs are much more willing to share, like really, really nitty-gritty tech details. The NVIDIA paper is like mostly a list of like methods to create a teacher or like-</p><p>00:36:51 Finbarr Timbers: Yeah</p><p>00:36:51 Nathan Lambert: ... domain-specific teachers, which is useful, but I think like I was less... It&#8217;s like less of a fun read. They&#8217;re like, there&#8217;s 15 pages of different domains, so I&#8217;m like, &#8220;Okay, I don&#8217;t, like I don&#8217;t need this.&#8221; Yeah, like KBK 2.5 and, uh, GLM 5 actually have like more similar recipes, which are also on the simpler side, which is like you create this SFT stage, and then you do RL. The RL might be staged. Um, there&#8217;s not this on-policy distillation. There&#8217;s a bit less talk on how many experts they have and what their domains of expert-s are. I think it, it&#8217;s obvious, like you have to take all this with a grain of salt, and it&#8217;s like what, how they decided to present the information is like a big factor in this. And then like they might actually be closer in reality and then it just wasn&#8217;t described in a certain way.</p><p>00:37:44 Finbarr Timbers: I, I think another interesting bit is that you see the Chinese labs, uh, all seem to be converging towards sparse attention, whereas, uh, we don&#8217;t see the, you know, where was the American labs, at least NVIDIA and, you know, AI2 seem to be more converging towards hybrid attention. Uh, like N- uh, the NVIDIA Ne- Nemotron Ultra used the Mamba, um, attention, whereas, you know, we see, you know, DeepSeek sparse attention and then the Mimo, eh, MSA, whatever that stands for, Mimo Sparse Attention. So I, I think that&#8217;s, uh, an interesting divergence.</p><p>00:38:20 Nathan Lambert: Yeah. I am not the person to ask, but I agree.</p><p>00:38:23 Finbarr Timbers: [laughs]</p><p>00:38:23 Nathan Lambert: It&#8217;s like I... Like I, I often get asked of like, this is to, to... Don&#8217;t, we&#8217;ll avoid the full rabbit hole, but I often get asked like, &#8220;Are the Chinese labs more efficient?&#8221; And I&#8217;m like, &#8220;I don&#8217;t really know how I&#8217;m gonna give you advice unless I&#8217;m spending twenty hours a week look, understanding the details of your recipe,&#8221; &#8216;cause it&#8217;s like, well, I can&#8217;t really give you a one sentence thing of do X without understanding all the complexities of the model and the post-training process that go into it. Which makes it, like makes it hard from kind of like a transparency point of view. Even if it&#8217;s fully detailed, it&#8217;s definitely still hard to modify and study.</p><p>00:38:42 Finbarr Timbers: Yeah</p><p>00:38:42 Nathan Lambert: ... like if you make a GPT model 1% more efficient, you&#8217;re making like fat stacks of profit. Like, I think that&#8217;s like a more effective market mechanism, but-</p><p>00:38:53 Finbarr Timbers: And then-</p><p>00:38:53 Nathan Lambert: The Chinese lab-</p><p>00:38:54 Finbarr Timbers: You know-</p><p>00:38:54 Nathan Lambert: Yeah</p><p>00:38:55 Finbarr Timbers: ... if you make, you know, serving ChatGPT more efficient, Sam Altman can say, &#8220;Hey, here&#8217;s a bunch of stock.&#8221; Like, so yeah.</p><p>00:39:02 Nathan Lambert: Yeah. But, uh-</p><p>00:39:03 Finbarr Timbers: Um</p><p>00:39:03 Nathan Lambert: ... they do great, like the Chinese labs do great research.</p><p>00:39:05 Finbarr Timbers: Absolutely.</p><p>00:39:05 Nathan Lambert: I just think it&#8217;s kind of a bit different. Okay, we can move into more open-ended stuff here.</p><p>00:39:12 Finbarr Timbers: Sure.</p><p>00:39:12 Nathan Lambert: I think that we have like a bunch of docu... We have th- a bunch of things in a document here. I&#8217;m sure more will come up. How do you think about open models and kind &#8216;cause i- it just doesn&#8217;t strike me that there&#8217;s this, like, you know, I think that there&#8217;s a large business to providing... Well, actually that&#8217;s not even super clear. There&#8217;s, you know, we&#8217;ve seen a number of companies providing, you know, RL fine-tuning services, you know, RL as a service. We&#8217;ve seen a lot of companies try to provide fine-tuning as a service, and, you know, none of them have really taken off. Like, I think OpenAI has started to shut down, I think they shut down their RL fine-tuning. I think they might be shutting down their fine-tuning. May be wrong about that.</p><p>00:45:51 Nathan Lambert: Well, it&#8217;s like Cursor used Fireworks for their actual training run, and I&#8217;m like, I don&#8217;t really know all the details of this, but Cursor does something for fat- I think like fast weight tran- or Fireworks does-</p><p>00:46:01 Finbarr Timbers: Yeah</p><p>00:46:01 Nathan Lambert: ... a fast weight transfer and other things to make it so that they can scale their RL inference compute very nicely. So that&#8217;s one type of it. I don&#8217;t know how big of a long tail that business is, but also I think Tinker is a better business than most people expected. It makes some real amount of money. It&#8217;s like in the hierarchy, I think selling compute, not the best business.</p><p>00:46:23 Finbarr Timbers: Yeah.</p><p>00:46:23 Nathan Lambert: Selling inference, great business. And Tinker-like APIs, if you can&#8217;t transition it into selling tokens, is somewhere in between the two, where they could take some amount of margin that&#8217;ll be slightly higher than just selling the compute. And they obviously get a margin by having, like, they get compute at a cheaper rate than their customers-</p><p>00:46:43 Finbarr Timbers: Yeah</p><p>00:46:43 Nathan Lambert: ... and that&#8217;s like part of the margin they&#8217;re taking. But I don&#8217;t see it being as nice as inference, so it&#8217;s kind of existential for them to make it so that these fine-tuning APIs feed into a inference business pretty nicely.</p><p>00:46:56 Finbarr Timbers: Yeah.</p><p>00:46:56 Nathan Lambert: Because then you can be somewhat locked in on you train the model on our infrastructure. You actually can own the model weights, but the training dynamics to inference mismatch is perfect because you trained exactly on our inference engine, and are gonna get what you want out of it.</p><p>00:47:11 Finbarr Timbers: Yeah. And it also helps a lot with utilization because you can then, you know, utilize it. You, you can share that utilization across a lot of clients. So I think it makes a lot of sense. I think it&#8217;s probably a better model for a lot of, um, users. Like, I think of academic users, like it probably makes way more sense to do this. Or, you know, for that matter, if you&#8217;re, you know, as, uh, uh, starting a new, um, ar- you know, post-training lab now, as you know, I, I know a few people, um, who are. Like, I think that&#8217;s where it, it probably makes a lot of sense to start with something like the Tinker API, and then, you know, at some point if you wanna try and capture that margin, maybe then you try to do something more custom. But if you, if you can use something like that, like that&#8217;s great, and the economics are just, you know, fundamentally more sustainable. I or, you know, they&#8217;re better for you rather than trying to, you know, g- go to CoreWeave or whoever and say, or Serv scale and say, &#8220;Hey, I need, you know, 10,000 networked, uh, DB200s,&#8221; you know? That&#8217;s just a very expensive, um, thing to do, especially if you can&#8217;t keep it running all the time.</p><p>00:48:14 Nathan Lambert: Yeah. Do you have a, do you have any more hot takes on post-training before I ask you some more general things?</p><p>00:48:22 Finbarr Timbers: Uh, well, something I&#8217;m, I&#8217;m generally interested in and, you know, I, I&#8217;m the wrong person to, to speak to about it. I&#8217;d love to talk to someone who&#8217;s maybe a, a, a capital allocator, like who&#8217;s, you know, deciding or a compute allocator who&#8217;s deciding where to put, uh, compute or, you know, where to hire team members. Um, because I&#8217;m kind of curious how Uh, the high level decisions are made allocating resources between pre-training and post-training. Uh, &#8216;cause, you know, what I kind of have seen as, as a general trend is, is that you see a lot of papers where there&#8217;s, you know, more focus put on one or the other. Uh, like I think... So, so yeah, so that&#8217;s something kind of interesting to me is how people who are, you know, making this decision, how, how they&#8217;re making that decision and how they&#8217;re thinking about it.</p><p>00:49:10 Nathan Lambert: Yeah. It&#8217;s like the hardest decision to get out of labs. I&#8217;ve like, I used to spend time trying to get them to share more, but I, I think it&#8217;s like such a sensitive decision to where they see progress coming. Like they&#8217;re making that decision ba- allocating compute based on where they think the most progress is and what the like return on investment is. So if you go to Anthropic and they&#8217;re like, &#8220;Here&#8217;s where our percent, here&#8217;s our distributions,&#8221; it&#8217;s like, okay, that&#8217;s where labs see their bets and/or where they see they are weak.</p><p>And it&#8217;s like you invest more compute in the pro- to make progress in the area that you are interested in, which I always think makes a lot of the open research kind of boring right now, is like the people that get compute are just way more likely to succeed as academics and researchers, which is a horrible equilibrium for the world, but kind of realistically true. I, I, I don&#8217;t know how to make a lot of that. I wanted to ask you how you feel about the craze that people have to cash in on making money and join a lab before the ladder gets pulled up, and what people should be optimizing for in their careers in face of meaningful opportunity costs.</p><p>00:50:18 Finbarr Timbers: Yeah. I think it&#8217;s, well, that&#8217;s actually very, very timely. Uh, but yeah, no, I, I think that that&#8217;s, um, really important to, to talk about. I mean, I think it&#8217;s always worth focusing on whether what you&#8217;re doing and spending time on is gonna be generally valuable or if it, if it&#8217;s like a really short-term exploitation type thing in, in the, you know, RL like explore versus exploit setup. I, I mean, something that I&#8217;ve seen throughout my career has been often the places that pay the most, um, are also the places where you&#8217;re doing the most interesting work, right? Like, you know, if, if you&#8217;re gonna go work at OpenAI, OpenAI or, you know, Anthropic or the Frontier Lab, like they pay a lot of money. They also have a lot of resources, so you&#8217;re gonna make a lot of money and learn a lot.</p><p>Um, uh, so I think it&#8217;s worth trying to decide i- is that the, is the opportunity that you&#8217;re doing that or is the, is the opportunity like, you know, in 2021 or 2022 or whatever, where you might say, you know, I was at DeepMind at the time and it&#8217;s like, okay, do I work at DeepMind, which paid a lot less than like crypto? Should I go just, you know, work in crypto and try to, you know, mint NFTs or whatever? I think that would&#8217;ve been a mistake, but, you know, trying to figure out, um, if you&#8217;re gonna be able to do interesting work is really important and also, you know, try to figure out if you&#8217;re going to be able to, you know, push forward science. You know, if, if what you&#8217;re doing is more just saying, going to, you know, data vendors and saying, you know, &#8220;Okay, you know, we, I need a bunch of data to do whatever.&#8221; And then, you know, they, they give you a bunch of data, you train a model, you say it&#8217;s good or bad or whatever.</p><p>You know, I don&#8217;t think that&#8217;s as interesting and, and I don&#8217;t think you&#8217;re gonna learn a lot even though that&#8217;s, you know, work that would probably drive model progress for it. I think if you&#8217;re able to, you know, make, focus more on the science and make more scientific conclusions, I think that can be, you know, a lot better for your long-term career. And I think that&#8217;s where places like AI2 and the other, um, academic research labs, you know, Marin is doing a really great job of this. Um, I think that&#8217;s where you can have a lot of impact in that they don&#8217;t have the budget to go and buy a lot of data, and so that leverage just really isn&#8217;t, um, open to them to pull. And so they have to focus on science and driving innovation, and that&#8217;s where you can see things like the Almix, uh, paper, which I thought was a really excellent, uh, sc- you know, scientific paper, but also, you know, meaningfully, I think, advanced, uh, the state of the art.</p><p>00:52:32 Nathan Lambert: Yeah. No, mostly this is grounded in visiting the Bay Area, and every time I go I&#8217;m like, &#8220;Holy shit, what is going on here?&#8221; Like all these very junior people are like have way too much dread about their, uh, opportunity cost and both of us aren&#8217;t based in the Bay Area, so I feel-</p><p>00:52:46 Finbarr Timbers: No</p><p>00:52:46 Nathan Lambert: ... somewhat removed from it, which gives me a little bit more time to pause and be like, what exactly is the right thing to optimize for? I per- I-- it&#8217;s easy for me to say as somebody that&#8217;s established, but I think there&#8217;s opportunity for a lot of people to just, if they have conviction on something, to try to go and do it and not just follow everybody that goes down the funnel of joining one of the established labs or the Neo labs where I don&#8217;t hear from many people that join as a junior person at these places and end up with very high responsibility. Like they&#8217;re contributing to something that matters or they&#8217;re around a cool group of people, but I don&#8217;t hear from that many people that are like, &#8220;Wow, I am doing the highest leverage stuff and the most interesting things.&#8221;</p><p>00:53:30 Finbarr Timbers: Well, I think that, you know, it&#8217;s kind of funny for, for me to say this as I, my career has been more on, on the opportunistic, uh, side of things. Um, but you know, twice now, uh, I&#8217;ve been at organizations where, um, I, I&#8217;ve been working... So, you know, at, at DeepMind, uh, I, I was part of the Alberta office where DeepMind had, you know, aqua hired the, uh, computer poker research group from the University of Alberta. And so, you know, this was a group of people who were really invested in, uh, computational game theory and g- you know, poker playing, um, algorithms. And they were all in on that and, you know, they, they were all in on that to the point that, you know, they were one of the two leading, uh, labs in the field and, um, were, you know, b-because they were so strong at this, they were then, you know,</p><p>DeepMind came and, you know, acquihired them and, and they all joined and they, you know, did quite well from that, um, acquisition there. And then, you know, you know, I joined later because I was, uh, you know, interested in, in working with them and doing game theory and stuff. But you know, it was this group of people who had this conviction that what they were doing was really important and, you know, it worked out quite well for them. And then, you know, the same thing at AI2, where at AI2, you know, there was all of these people who were really interested in, uh, NLP research, you know, even before language models. Like we see people like, you know, like Kyle a-and Dirk I think were both at AI2 for like almost a, a decade.</p><p>Like they had these really long tenures, um, and then they did really well and then, you know, they&#8217;ve, they&#8217;ve since had some, you know, strong, um, opportunities, uh, coming out of that with, with, um, yeah, some of the opportunities that have been available to them. And I, and I think that the consistent theme there has been that, you know, if you have high conviction that what you&#8217;re doing is important and interesting, then like it, it&#8217;s not a mistake to follow that and to, you know, try to become really strong, um, in that area.</p><p>00:55:15 Nathan Lambert: Yeah. I mostly think it&#8217;s good for the world to have a di- more diverse set of approaches.</p><p>00:55:19 Finbarr Timbers: Yeah.</p><p>00:55:19 Nathan Lambert: It&#8217;ll be interesting to see what the deal labs actually produce if, if they can manage to do things that are diverse. My personal idea is that they&#8217;re so big now that most of them need to end up doing something that is somewhat similar, which is-</p><p>00:55:33 Finbarr Timbers: Yeah</p><p>00:55:34 Nathan Lambert: ... hard, but like they need to keep risking the comp- they effectively need to risk their $20 billion valuations to do something interesting that&#8217;s not just gonna be like squashed by an OpenAI or Anthropic side project.</p><p>00:55:48 Finbarr Timbers: Yeah, absolutely. And I think it&#8217;s tough because when you&#8217;re raising, when you&#8217;re, you know, you have these huge seed rounds and you&#8217;re raising, you know, 200 million or, you know, a billion dollars or whatever, then it&#8217;s like you have to pretty quickly show results to be able to-</p><p>00:56:01 Nathan Lambert: Yeah</p><p>00:56:01 Finbarr Timbers: ... you know, grow off of that.</p><p>00:56:04 Nathan Lambert: Yeah. So a to-be continued conversation.</p><p>00:56:11 Nathan Lambert: Any last words? I don&#8217;t, I don&#8217;t need to stretch it on if we don&#8217;t have anything to add to our conversation.</p><p>00:56:16 Finbarr Timbers: No, I, I think this was pretty good. I think it was really great, uh, getting a chance to catch up and talk about some of this stuff. You know, I, I&#8217;ve been reading all of these papers and thinking about all the different recipes, so it&#8217;s great to get to, um, to chat about it and put it out into the ether. So yeah, thanks for having me on.</p><p>00:56:31 Nathan Lambert: Yeah, thanks for coming back. We&#8217;ll talk soon.</p><p>00:56:33 Finbarr Timbers: Sounds good.</p>]]></content:encoded></item><item><title><![CDATA[Welcome to the AGI era of AI governance]]></title><description><![CDATA[It's a one-way door and we weren't ready for it.]]></description><link>https://www.interconnects.ai/p/welcome-to-the-agi-era-of-ai-governance</link><guid isPermaLink="false">https://www.interconnects.ai/p/welcome-to-the-agi-era-of-ai-governance</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Sun, 14 Jun 2026 17:43:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/74e1df37-06b8-444e-a5a4-fbc1ead00360_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The executive branch of the United States forcing Anthropic to turn off access &#8212; both internally and externally &#8212; to their latest <a href="https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety">Claude 5 Mythos/Fable</a> models is the starting gun of a new era in AI governance. This is the era defining AI agents that effectively complement human workers, unlocking new ways of working and new domains of applying existing tools. The models will keep getter better at a rapid pace, forcing more governance challenges.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/welcome-to-the-agi-era-of-ai-governance?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/welcome-to-the-agi-era-of-ai-governance?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>An unstable equilibrium at the frontier</h3><p>On Friday, just after markets closed, the government reached out to Anthropic to <a href="https://www.anthropic.com/news/fable-mythos-access">force them to suspend access</a> to their model to any foreign national or user abroad. When writing this, the saga is still unfolding. It is likely that Anthropic and the Government reach an agreement to re-release the model, but it is a messy indicator of new types of AI governance. The White House was <a href="https://www.theverge.com/ai-artificial-intelligence/949601/amazon-anthropic-fablemythos-government-ban">tipped off</a> to the risk by Anthropic&#8217;s largest financial and technology partner, Amazon. </p><p>These dynamics are very complex, and while the picture is incomplete I will only offer commentary on what will likely be a lasting opinion of the saga:</p><ol><li><p><strong>An export ban on any model weights is going to be a lasting, negative policy for the U.S.</strong> This goes for both open and closed models, despite the open models being more likely to get <em>import</em> banned soon.</p></li><li><p><strong>There was a legitimate cybersecurity concern regarding a jailbreak of the Fable model</strong>, even if it was very narrow. I do not personally think this should&#8217;ve set off this series of events, as no model is perfectly immune to jailbreaks and we need to prepare our infrastructure for a world where every global entity will have access to models of this capability level in the coming years.</p></li><li><p><strong>Anthropic&#8217;s constant fear-mongering over the last few years has accelerated this moment.</strong> Without the constant messaging comparing AI to nuclear weapons, etc., I suspect this style of governance would not have manifested for another 6 to 12 months (of AI progress). There is some amount of Anthropic reaping what it has sowed.</p></li><li><p><strong>People should not be taking a victory lap at our nation&#8217;s leading AI company being repeatedly subject to potentially political attacks</strong> (yes, I&#8217;m looking at people in the open-source community)<strong>.</strong> This is economically highly unstable and undermines the economic stability of the American system. Sufficient instability could cause an economic recession and pop the AI bubble.</p></li><li><p><strong>The White House&#8217;s actions to ban the model are heavy handed and are influenced by a <a href="https://x.com/PeteHegseth/status/2065897156226015690">strong political inclination </a></strong><em><strong><a href="https://x.com/PeteHegseth/status/2065897156226015690">against </a></strong></em><strong><a href="https://x.com/PeteHegseth/status/2065897156226015690">Anthropic</a>.</strong> There&#8217;s a deep contradiction in the government&#8217;s demands and goals, as there is no domestic AI industry if foreign nationals cannot build with frontier AI in the U.S. </p></li><li><p><strong>It is unclear to me as to why Amazon needed to take the information they had directly to the White House</strong>, what the lead-up to this release and ban were (lots of accounts online of pre-release discussions), and how crisis management between leading technology companies and the White House would normally play out. This points to weird dynamics with how Anthropic interacts with the government, which could be a reaction to being politically singled out. <br><br>Whatever the cause, it is a bad dynamic and casts a shadow over how Anthropic engages here. If a private company tries to control the government the government will push back stronger. I do not think Dario was actually at a wellness retreat (as said by a White House representative) but I do think he took longer to answer the call than is acceptable with who was on the other end of the line.<br><br>It&#8217;s hard for me to write anything cogent about these back and forth dynamics but they are very important to the truth of what is happening.</p></li></ol><p>This picture is very complex. It points to a near-term world where model releases are judged on vibes by an executive branch with minimal technical talent, gating releases behind a blur of politically judged technical assessments. </p><p>This is a government that has internalized that we are in the AGI era. They were not ready for it and feel like they must act fast to regain lost time. </p><p>This is a government that took office when we were still in the ChatGPT era of AI governance &#8212; models that just answer questions. Their original <a href="https://www.interconnects.ai/p/the-white-houses-plan-for-open-models">AI Action Plan</a> reflected this era, highlighting how a lot of previous safety discussions were more smoke and mirrors than concern, and how badly we needed open-source models to catch up. On balance, their opinions and prescriptions here were mostly agreeable and useful for the industry. </p><p>It is important to remember that even in this ChatGPT era of governance, the AI companies often told us that the models were a risk and could not be released without great care. Many spokespeople at leading labs didn&#8217;t give themselves enough headroom in language escalation for when the models took a meaningful jump in performance. </p><p>As we shift from <a href="https://stratechery.com/2026/the-inference-shift/">answer inference to agentic inference</a>, I&#8217;m still on the side that the risks have been exaggerated, but the margin between the real risks of the latest AI systems and the language used to describe them has shrunk. This has spooked power structures outside of the leading AI labs.</p><p>Entering the AGI era of AI governance is when a lot more sovereigns around the world start to take AI seriously. <a href="https://europe2031.ai/">Europe</a>, the Middle East, and likely also China are realizing they could be left without frontier AI. This swing is something that the open-source community takes as a win, seeing the long-term trajectory where businesses will want to control their own intelligence when the government can arbitrarily rule against the leading platforms.</p><p>The key to understanding this era is that these events between Anthropic and the federal government, which seem like two of the biggest events in the history of AI policy, are just the starting gun to what is to come. This is the new normal, as existing power structures feel the need to assert their control over the rapidly rising technology. This will be messy, fraught, and at times even perilous. </p><p>When following these events I&#8217;m filled with a sense of dread, staring into the void, because I know more is to come. It&#8217;ll take resolve to follow de-escalatory paths that help us achieve the positive visions of broadly diffused, safe, and cheap super-powerful AI around the world.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/subscribe?"><span>Subscribe now</span></a></p><h3>Sovereign AI, open-source needs, and battles soon to come</h3><p>The open-source AI advocates wildly celebrating this sequence of events aren&#8217;t remotely ready for when the spotlight of rapid-response AI policy turns its gaze upon them. It is very likely that a similarly aggressive action against an open model will come, but we don&#8217;t know if this is in 3 months or 2 years.</p>
      <p>
          <a href="https://www.interconnects.ai/p/welcome-to-the-agi-era-of-ai-governance">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Claude Fable 5 and new AI safety fables]]></title><description><![CDATA[One step further into the power politics of frontier AI systems.]]></description><link>https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety</link><guid isPermaLink="false">https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Tue, 09 Jun 2026 22:59:35 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b9a6c144-02b3-4a08-947c-c517667d2249_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>Edit Jun. 11: Anthropic <a href="https://x.com/ClaudeDevs/status/2064949876463645026">changed</a> their silent model manipulation of AI research queries to also use a classifier like the other safety domains. This addresses a key concern I had in the mistreatment of &#8220;safety&#8221; in the release, and props to Anthropic for a quick change, but it does not fully address the trust that has been broken. I shared more reflections <a href="https://natolambert.substack.com/p/anthropic-walks-back-silently-nerfing">here</a>.</h5><p>Today, Anthropic <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">released</a> their Claude Fable 5 model to consumer and enterprise audiences. This is the general-access variant of their Mythos-class models. With it, Anthropic rolled out a series of safety measures &#8212; some explicitly called out to users and some modifying the model without telling the user. It should be less surprising than it is that the next major step in AI capabilities came with heavier-handed safety measures indicating Anthropic&#8217;s intention to protect, or entrench, their current lead.</p><p>The unevenly applied safety policies that Anthropic have rolled out are on track to become a classic cautionary fable in how narrow and self-fulfilling notions of safety and control rarely work out.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>The smartest model in the world</h3><p>Before digging into the nuance of the safety facts, it is important to establish the quality of this model. The quality of the model paints the stakes of today &#8212; as these safety features are meaningfully changing the shape of access to frontier AI, something which has never happened with the modern LLMs we know. Second, the capabilities point to this story only accelerating. <a href="https://www.interconnects.ai/p/lossy-self-improvement">Recursive self-improvement isn&#8217;t quite the right mental model</a> of progress from here, but Claude Fable 5 should make it very clear that there are no immediate walls in training LLMs.</p><p>To start &#8212; Claude Fable 5 is definitely the smartest model available to the general public &#8212; a remarkable leap on pretty much every relevant benchmark of the day &#8212; at only 2X the price of current Opus models<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> (which is still less than GPT 5.5 Pro&#8217;s variant). This alone is a seminal moment for the field. To have a model iteration take such a substantial step in capabilities, a few years into the post-ChatGPT LLM race, is astounding. There&#8217;s no clear breakthrough associated with this model, such as inference-time scaling or RL, and public wisdom is that this is achieved by advances across the whole stack (of course, we can&#8217;t know for sure &#8212; it&#8217;s not documented). This is a major technical achievement and the employees who built the model should be very proud of their work.</p><p>This model was delayed 2+ months after it was done training before it was publicly available<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>. Given the competitive dynamics of the AI economy, the smarter version of this model is already well underway.</p><p>To continue, the benchmarks for the model are below.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zKZX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zKZX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp 424w, https://substackcdn.com/image/fetch/$s_!zKZX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp 848w, https://substackcdn.com/image/fetch/$s_!zKZX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp 1272w, https://substackcdn.com/image/fetch/$s_!zKZX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zKZX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp" width="1456" height="1607" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1607,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Benchmark table showing Claude Fable and Mythos compared to other leading models&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Benchmark table showing Claude Fable and Mythos compared to other leading models" title="Benchmark table showing Claude Fable and Mythos compared to other leading models" srcset="https://substackcdn.com/image/fetch/$s_!zKZX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp 424w, https://substackcdn.com/image/fetch/$s_!zKZX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp 848w, https://substackcdn.com/image/fetch/$s_!zKZX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp 1272w, https://substackcdn.com/image/fetch/$s_!zKZX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7caf7c30-6c3d-4735-b600-02d7c534525d_2600x2870.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>An asterisk on these scores is that these aren&#8217;t necessarily the scores that the public will get, as some of the prompts will be downgraded to Opus 4.8 with the current safety filters on the model.</p><p>This is the type of jump in benchmark scores where I don&#8217;t even need to substantially test the model to know it&#8217;s an incredible tool. Remember that Anthropic is also the AI lab with the track record of caring <em>the least</em> about benchmarks (in particular, when compared to OpenAI and Gemini). Recall a comment I made in <a href="https://www.interconnects.ai/p/summertime-outlook-o3s-novelty-coming">June of 2025</a>:</p><blockquote><p>This is a different path for the industry and will take a different form of messaging than we&#8217;re used to. More releases are going to look like <a href="https://www.interconnects.ai/p/claude-4-and-anthropics-bet-on-code">Anthropic&#8217;s Claude 4</a>, where the benchmark gains are minor and the real world gains are a big step. There are plenty of more implications for policy, evaluation, and transparency that come with this. It is going to take much more nuance to understand if the pace of progress is continuing, especially as critics of AI are going to seize the opportunity of evaluations flatlining to say that AI is no longer working.</p></blockquote><p>Clearly, a few pieces of the progress dynamics have changed, but that&#8217;s a post for another day. I&#8217;ve written <a href="https://www.interconnects.ai/p/opus-46-vs-codex-53">multiple</a> <a href="https://www.interconnects.ai/p/get-good-at-agents">posts</a> about new models this year specifically in how it&#8217;s hard to trust benchmarks (and partially because the benchmarks don&#8217;t move that much). Altogether, this is a major validation for AI-savvy workers who realized they&#8217;re likely never going to write meaningful code again and need to develop new workflows around agents. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>Smarter models spawn new safety games</h3><p>There are multiple pieces of safety tooling associated with this release, including but not limited to required data-retention policies and added prompt filters. Through this analysis it is particularly important to be precise and clear as to which pieces of these are causing harm, and why single elements being out of place in an otherwise comprehensive policy are so damning for the overall safety process.</p><p>For their focus areas of cybersecurity, targeted model distillation, and research biology, Anthropic details new safety classifiers in their <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">blog post</a>:</p><blockquote><p>Fable 5 comes with a new set of <em>classifiers</em>: separate AI systems that detect potential misuse, including jailbreak attempts, and prevent the main model (in this case Fable 5) from responding. We&#8217;ve been running classifiers on our models <a href="https://www.anthropic.com/research/next-generation-constitutional-classifiers">for some time</a>, and Fable 5&#8217;s classifiers are an extension of this previous work with extra coverage.</p><p>When Fable&#8217;s classifiers detect a request related to cybersecurity, biology and chemistry, or distillation, the response is automatically handled by Claude Opus 4.8 instead. Users will be informed whenever this occurs. Opus 4.8 is a highly capable model in its own right: a response that falls back to Opus is a far better experience than an outright refusal from Fable. Our early data shows that more than 95% of Fable sessions involve no fallback at all&#8212;for those sessions, Fable 5&#8217;s performance is effectively the same as that of Mythos 5.</p></blockquote><p>Examples of the primary cybersecurity and biology safety filters &#8212; which tell the users explicitly when they&#8217;re triggered &#8212; are <a href="https://x.com/DimitrisPapail/status/2064415276968333548">already</a> <a href="https://x.com/DeryaTR_/status/2064414826122866707">proliferating</a> <a href="https://x.com/acerfur/status/2064400810054680634?s=46">online</a> and appear quite sensitive. These can be a frustrating experience for users, but Anthropic is definitely within its power to do this and intellectually consistent for doing so. </p><p>The damaging part of the safety story falls under the fold in the <strong>Claude Fable 5 &amp; Claude Mythos 5 <a href="https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf">System Card</a>:</strong></p><blockquote><p><mark data-color="#e1edeb" style="background-color: rgb(225, 237, 235); color: rgb(0, 0, 0);">We have also added safeguards related to frontier LLM development.</mark> As discussed in Section 6.1 of our February 2026 Risk Report, we are concerned about the risks of accelerating the overall pace of AI development, though we remain uncertain about the severity of these risks. In particular, our concern is with&#8212;as we wrote then&#8212;&#8220;accelerating other AI developers in building powerful AI systems that pose similar risks to the ones ours pose - without necessarily having commensurate safeguards.&#8221; </p><p>In light of the ability of recent models to accelerate their own development, we&#8217;ve implemented new interventions that limit Claude&#8217;s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms. </p><p><mark data-color="#e1edeb" style="background-color: rgb(225, 237, 235); color: rgb(0, 0, 0);">Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user.</mark> <mark data-color="#e1edeb" style="background-color: rgb(225, 237, 235); color: rgb(0, 0, 0);">Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT).</mark> </p></blockquote><p>Anthropic documents on how this will impact a small percentage of users, which is true. I focus on the small amount of users supporting AI&#8217;s diffusion and understanding outside of the few frontier labs, as a crucial mechanism for the continued safety of the technology. </p><p>Anthropic is documenting how the proliferation of AI capabilities is a concern to them, but they are solving it by misleading their users. An AI model that gets less intelligent automatically without notifying me is categorically misaligned AI. The next step on this line &#8212; not that Anthropic did it, but they could &#8212; is to have a model silently manipulate a workplace when it thinks it is an unsafe use for AI. Second, the implementation here is more complicated than was documented for cybersecurity or biology &#8212; modifying the model itself or the data presented to it, all without notifying the user.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p>The duality of these policies is extremely confusing and paints a strong inconsistency that casts doubt over their safety policies. This &#8220;safety&#8221; measure is presented as being far more about maintaining their competitive position. Again, if all of the safety policies took one form, this would be far more cogent and easier to support intellectually.</p><p>Anthropic has been very vocal about their <em>concern</em> over distillation attacks from particularly Chinese actors. Their claims are not transparent enough with the facts &#8212; or context as to why they can&#8217;t prevent the behavior &#8212; <a href="https://www.interconnects.ai/p/how-much-does-distillation-really?utm_source=publication-search">to be fully believable</a>. Despite the limited information, in the broader AI and DC communities, there have been <a href="https://www.interconnects.ai/p/the-distillation-panic?utm_source=publication-search">serious discussions about taking action against the Chinese model builders</a> on the grounds of said distillation.</p><p>On the point of distillation, my hypothesis is that API builders don&#8217;t have an easy time preventing hacks or jailbreaking because it&#8217;s a deeply grounded property of reasoning models to want to output the reasoning traces, and it would make the model far less intelligent to fully patch the behavior. This is based on a few assumptions:</p><ol><li><p>Chinese labs are <em>not</em> just showing up as customers to Anthropic&#8217;s API and paying for tokens in the intended input-output form. If the Chinese labs are paying for intended use behaviors, despite being banned by the terms and conditions, I don&#8217;t have a lot of sympathy for the frontier labs manifesting policy actions against this.</p></li><li><p>Reasoning traces are disproportionately effective at seeding behavior in downstream models.</p></li><li><p>Leading labs work very hard to patch the pipeline of these jailbreaks.</p></li></ol><p>So, my logical conclusion is that the model companies would have to weaken their economic position to fully protect their IP. If this is the case, Anthropic would get a lot more sympathy from the AI research community by being transparent. It would also be far easier to have informed policy discussions, and not rely on me proposing Occam&#8217;s razor explanations for what the API jailbreaking looks like.</p><p>Building these safeguards is not something that Anthropic should do alone. Safety research should be built on common understanding and information sharing across both labs and public research efforts. </p><p>If the exact safety procedures were actually the top line item to the company &#8212; a true non-negotiable for the leadership &#8212; they wouldn&#8217;t permit the model to be released with an unclearly implemented safety filter in one of their areas of focus (frontier AI training). I am asking &#8212; why isn&#8217;t there a classifier to downgrade AI research requests? This is a mix of transparent and reasonable safety policies with quietly rolled-out market entrenchment tactics.</p><p>I personally cannot trust the best AI model in the world to work in my professional domains building models, which I&#8217;ve constructed entirely out of a passion for making sure the transition to very powerful AI systems goes well for society. This inevitably will feel like a declaration of superiority by the Anthropic leadership. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>The control problem and open-source as the only answer</h3><p>All of the actions Anthropic is taking, including calling out smaller Chinese companies for distillation, is well within their right. In fact, many people already expected the leading frontier models to be obviated from users so that labs can protect their IP. Today&#8217;s actions miss the big picture that AI will always be an ecosystem, and cultivating an us against them dynamic between the leading company and the other players is structurally unstable. </p><p>Remember, this is at a time when the AI ecosystem is seeing the first stirrings <a href="https://jasmi.news/p/warning-shots">of violence against AI leaders</a> &#8212; and I&#8217;ve heard from many people that they don&#8217;t expect it to abate. I wish I knew how to engage more to prevent this, and I see myself in the non-profit sector as someone who can hopefully independently represent AI to broader stakeholders.</p><p>I believe there was something misread, or at least misunderstood here, by the Anthropic leadership having a narrowly cultivated worldview around AI. An overwhelming sentiment I had today was one of obligation and confusion. I <a href="https://x.com/natolambert/status/2064412173527556298">shared</a> how I don&#8217;t really want to have to go to bat against Anthropic, but they&#8217;ve just been unnecessarily antagonistic to China, then not so subtly to open weight models, and now more broadly to open AI research. </p><p>I understand that Anthropic has a specific view of AI, but such a powerful technology will never have its final equilibrium be one of singular control by a private company. Anthropic showcased this earlier this year in the spat between the Department of Defense and themselves &#8212; which points to a long-term equilibrium where the government will either want AI to be controlled by them or to be open. This <a href="https://www.interconnects.ai/p/how-anthropic-vs-dow-impacts-open">made me believe</a> that an open ecosystem is a far safer outcome.</p><p>Many of these events make me feel that Anthropic&#8217;s leadership has a culture by which they can&#8217;t help but speedrun through these issues &#8212; going head to head with existing power structures. This adds substantial uncertainty into an AI ecosystem at a time when it is very much not needed.</p><p>Collectively, the last week could be seen as a major rallying point for a new open-source ecosystem in the U.S. Nvidia released their first flagship model last week &#8212; <a href="https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/">Nemotron 3 Ultra</a> &#8212; and these actions from Anthropic have galvanized a unanimous motivation and concern among my peers building open models. We need intelligence that we can trust, that we can modify, and that we can control. </p><p>The American open-source ecosystem has its feet underneath it and keeps being given more reasons to fight for its leadership, right from the hands of the companies it directly undercuts. That&#8217;s the moral of this fable.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Fable is at $10 per million input and $50 per million output tokens.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>based on the original Mythos roll-out, which is an imperfect metric.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Fable <a href="https://claude.ai/share/78d57944-e7d3-4b2c-ba9f-8febd163fbb3">confirmed</a> for me that these are different mechanisms.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Farewell Ai2]]></title><description><![CDATA[This was my last week at the Allen Institute for AI (Ai2), where I got the great privilege to work on the Olmo models, to grow, to learn, and to have broad lasting impacts.]]></description><link>https://www.interconnects.ai/p/farewell-ai2</link><guid isPermaLink="false">https://www.interconnects.ai/p/farewell-ai2</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Tue, 02 Jun 2026 14:15:59 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/04d1f8be-c112-47a8-9a5b-251653b5951e_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;m departing the Allen Institute for AI (Ai2), where I got the great privilege to work on the Olmo models, to grow, to learn, and to have broad lasting impacts. This post is an attempt to reflect on why what we did was influential, despite obviously being far from the frontier in performance (even when within size buckets), and how this reflects on various paths to impact in AI today.</p><p>To start, I shared the following note with the company yesterday:</p><blockquote><p>Dear Ai2.</p><p>As many of you know, today is my last day working at Ai2.</p><p>I joined Ai2 largely as an accident. I met Luca at ICML 2023 in Hawaii and realized I could level up my open post-training work dramatically if I got the chance to join. When I got an offer it was an absolute no-brainer, it was such a welcoming and exciting environment.</p><p>It has been a wonderful ride that has transformed my life, and I couldn&#8217;t be prouder of the work we did together. Ai2 has a wonderful scientific culture at its core and I&#8217;m excited to see this continue. I feel very lucky to have been here and that I personally have benefited massively from everyone who has worked so hard to cultivate that culture and environment. It is and has been a team effort. This includes all the people whose longest interactions with me were brief chats at the coffee machine. I drew so much energy and excitement from all the different ways people at Ai2 showed up for the mission.</p><p>I&#8217;ve already thanked much of the OE team directly, but I wanted to thank everyone else that went into this. Legal, IT, Comms, and the Office team all do a great job enabling and leveling up our research work. It&#8217;s often work that is forgotten, outside of the lime light, or remembered at the last minute, but it all has been crucial to achieving our goals. I&#8217;m excited to keep visiting the wonderful Northlake space in the coming years.</p><p>Even though I&#8217;m leaving, I&#8217;m more excited than ever about Ai2&#8217;s mission. Ai2 operates in such a rare niche between academia and industry, where we can explore and influence the most important technology of our lifetime. Doing this openly is the best way to ensure the technology diffuses safely to everyone who may benefit. Ai2 needs to stay as ambitious as possible, trying to influence the cutting edge of AI and the biggest issues of the field. Do not shy away from these challenges &#8211; AI needs independent voices as it only becomes more geopolitical, socially disruptive, and central to the economy.</p><p>I will still be working in this space, working to make the open ecosystem better coordinated and more useful.</p><p>So as I go off to try something new, don&#8217;t be strangers. I&#8217;ll always be reachable at <a href="mailto:nathan@natolambert.com">nathan@natolambert.com</a> and will still live in Seattle for most of the year.</p><p>Nathan</p></blockquote><p>I have loved and will still love Ai2. Ai2 has a deep culture of caring about the research process, the outputs that get shared, and most importantly the people who do the work. This is why the institution creates countless wonderful people that go and spread the gospel throughout the research community. This core culture will remain through the rebuild, and there are plenty of resources to do impactful research across the spectrum of AI.</p><p>In the last two years of my time at Ai2 I&#8217;ve done so much meaningful work. Of course Olmo is at the top and has been my priority, but making time for consistent practice here on Interconnects, weekend cram sessions for <a href="https://atomproject.ai/">ATOM</a>, and also the fun <a href="https://rlhfbook.com/">RLHF book</a> make for a list that makes me wonder how I did it all. I was obviously obsessed with work, but not in a way that made me lose sleep or lose my overall wellness. It was the right long-term approach.</p><p>This impressive list is one where I was ruthless in saying no to things that didn&#8217;t matter and got all my work out to see the light of day. I had no medium-sized projects that didn&#8217;t succeed in the last few years. It makes me wonder if I wasn&#8217;t taking enough risk. It shows you can truly do so much with your time, and it&#8217;s actually harder to find the right problems and environment to do it. Many people are in environments where their work never becomes public or they&#8217;re forced to change topics consistently.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/farewell-ai2?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/farewell-ai2?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2><strong>From zero to hero</strong></h2><p>To start, I&#8217;d like to do a short recap on my path to Ai2 to show what Ai2 was just as much a growth story for me as an execution story.</p><p>I studied electrical engineering in undergrad, focusing on linear systems math and microelectronics.</p><p>I was admitted to the UC Berkeley EECS Ph.D. program to study microelectromechanical systems (MEMS).</p><p>I showed up at Berkeley in August of 2017 and realized AI was obviously the thing I should be doing. I asked the likes of Sergey Levine or Pieter Abbeel if they could advise me &#8211; they said no.</p><p>I threw all my energy into learning what I could about AI. I got a break to get advised by one of Sergey&#8217;s post-docs in 2018 or 2019. I went all in on that, I fought for funding, I fought to have <em>an</em> AI paper.</p><p>This process worked out by the end of my Ph.D. in 2022: I had access to the Berkeley AI Research (BAIR) building and collaborations in the department. It was a bumpy road.</p><p>I wanted to go to industry research, to get a nice paying job with intellectual freedom, something like FAIR or Google Brain at the time. HuggingFace was the only job that fit that bill, it was easy to say yes to.</p><p>I joined HuggingFace in May of 2022 and wasted my time at the company until ChatGPT was released. I used my RL background to write a <a href="https://huggingface.co/blog/rlhf">blog post</a> on RLHF which went viral. HuggingFace decided it would be good for me to form a team around this success.</p><p>In 2023 I learned NLP and about language models. I had a lot of fun and built an initial community. I got burned out by working remote with a huge time difference. I met Luca Soldaini at ICML in Hawaii, where I was giving a tutorial on RLHF, and they told me Ai2 was hiring.</p><p>I got the job at Ai2 largely because of my excitement and how I was saying I wanted to do a lot of stuff that sounded cool to them but no one was likely to do (RL related things). My interviews were far from a sure thing &#8211; this is a great job to land!</p><p>I started at Ai2 in October of 2023. I worked remotely for a while. I was doing normal research, I made the first reward model evaluation, RewardBench. It was a solid success, but nothing like how the pretraining team was getting ready to release the first Olmo.</p><p>I helped coach Ai2 on how to release models well, helping the T&#252;lu 2 project land (the first model to do DPO well, publicly at the 70B scale).</p><p>The first Olmo was released in early 2024, I squeaked onto the papers just by trying to be helpful and doing some basic post-training. I was already good at paying attention to which projects are actually important.</p><p>That summer I started rounding everyone up to do a &#8220;big frontier post-training project.&#8221; This became T&#252;lu 3, one of my favorite projects ever released, in fall of 2024. The goal was to beat Llama 3&#8217;s post-training with their own base model. The team morale was incredibly high and the execution was so timely, allowing us to coin the term Reinforcement Learning with Verifiable Rewards (RLVR) in the paper.</p><p>The crazy lengths I went to get the T&#252;lu 3 and Olmo 2 post-training done had me sending 40% more slack messages than anyone at the company and got me the award &#8220;The Cat Herder.&#8221;</p><p>2025 was a much simpler year. We were too slow to react to reasoning models, given we had been doing similar stuff with T&#252;lu 3, but sometimes that happens.</p><p>Originally we wanted to release Olmo 3 by June or July of 2025. That obviously didn&#8217;t happen, but we got the slim chance to train a bigger model, and it really landed. We threaded the needle.</p><p>Since Olmo 3 was released, it was clear that some changes were coming and I personally never got a big post-training project off the ground after that. Many other people managed great work in the spring of 2026.</p><p>This all leaves me here today showing you that only about half of my story at Ai2 is what I was known widely for, and the rest was building momentum. It often takes a year of building relationships and direction before really big successes can happen in a career.</p><p>I was just about a nobody when I joined Ai2 and I got to join a team that was willing to learn from the skills I had brought from HuggingFace. With how media works, I often think I get more recognition than I deserve for Ai2&#8217;s success.</p><p>The likes of T&#252;lu 3, Olmo 2, and Olmo 3 felt like generational team efforts. The amount of personal successes and breakthroughs that happened for those projects is immense &#8211; and to sustain them over such a long time period is incredibly hard to replicate. The sum far exceeded the individual parts.</p><p>I&#8217;ve heard many times in the last few months how people wouldn&#8217;t know about Ai2 if it wasn&#8217;t for my writing. Statements like this are overblown, but they are partially true and reiterate how crucial building relationships and getting the word out is today. </p><p>When you write a plan that is feasible, the world bends towards that plan. When you convince people it&#8217;s going to happen it only becomes more likely. Vision and compelling explanations are one of the items in shortest supply in the tech industry. Often building the thing is easy and explaining it is hard. If no one knows about your work, the value is often close to 0. So much of building reputation is about building relationships with people who will receive your work.</p><p>Reflecting on all of this, I&#8217;ve had a shockingly linear path through my career to incremental success. I would expect the first 10 years of most careers to be in search of finding one opportunity as good as Ai2, and you will not always be able to seize it. There are some ways to create more opportunities. </p><p>I&#8217;ve <a href="https://www.interconnects.ai/p/my-path-into-ai">discussed before</a> how a large part of my rise is down to many more senior and more established scientists being drawn into the closed ecosystems at the same time as an immense swell in interest for AI. This created a power vacuum that I, and a few other prominent scientists that I think form my &#8220;generation&#8221;, got to grow rapidly into.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><strong>The role of public scientists</strong></h2><p>With my work at Ai2 and Interconnects, I summarize my role and mission as trying to accomplish three things:</p><ol><li><p><strong>Provide clarity in the evolution of frontier models.</strong> This is easiest when the science has caught up, but even applying a scientific lens to how the models are changing is very useful to building trust in the broader AI ecosystem.</p></li><li><p><strong>Create a vibrant and diverse open (model) ecosystem.</strong> This is crucial to mitigating some risks of AI, particularly with concentration of power and myopia in studying frontier safety, that has motivated me now for 3-4 years. The risks haven&#8217;t abated.</p></li><li><p><strong>To build institutions</strong> that create people and ideas that further the above missions, and generally mission-driven individuals that are willing to advocate and build a future they believe in. AI is a grand problem, and not one that I can do alone, so I need to build brands to rise through the noise and attract likeminded people. </p></li></ol><p>At my best, I have many avenues for impact. I help open researchers work on impactful problems &#8211; not wasting the precious compute and time they have during the AI boom. I help policymakers know what is true. I build models that people use. I tell stories that make people smile. I keep the list wide so that I can stay motivated.</p><p>I see all of this continuing, and have been thinking about the broader impacts of this repeatedly over the last few months. Hearing that Andrej Karpathy was joining Anthropic prompted me to finally share more of <a href="https://x.com/natolambert/status/2056772229120295416">my opinions</a>:</p><blockquote><p>For a long time, academic researchers being at the cutting edge of new technologies has been a great social equilibrium. Neutral, unbiased technologists have been the people to spread new ideas to the world.</p><p>As AI research takes off in velocity, it is also going behind closed doors. The tech industry has sowed distrust, and now they are the ones trying to tell the world about incredible changes coming. It&#8217;s a big loss to a form of social contract in America.</p><p>There&#8217;s been a history of scientists helping society understand new technologies. There is a public service in the culture of science that I want to see continue.</p><p>It&#8217;s being exacerbated by feelings of FOMO, especially financially driven, where I&#8217;m seeing many people who previously wanted to be professors -- and likely still do deep down -- feel a need to conform and chase money, in a pocket of industry. I get it, I grapple with this.</p><p>For those with a safety net, there will be great returns to some who choose to zag, and try to build something good, for people who need something different. For me, this is building interesting, fully-open models, to show what you can do with a variety of open weight sizes.</p><p>Yes, AI&#8217;s immediate future is dictated by the frontier, but it&#8217;s long-term trajectory still deeply includes academic institutions and open science. Knowledge will always diffuse, but to whom?</p><p>As of today, I think China is positioned to be the global home of AI research in a few years. The home of research is where ideas are accessible, spread rapidly, and are nurtured. The U.S. seems to be unwinding many institutions and relationships.</p><p>The largest returns go to people who build something differentiated, at least in reputation, and a lot of people are not being shown that this path exists.</p></blockquote><p>To elaborate on this, I don&#8217;t fault any of the individuals who are going to industry today. I&#8217;ve been very close to doing this myself in the past weeks of job searching, or rather job exploring. It&#8217;s a systematic problem where scientists cannot easily get the support to take bold stances, especially stances that are designed around the public good.</p><p>To go a step further and say that only the research within closed, frontier labs matters is very myopic. Yes, there&#8217;s a sort of research you can only do with vast compute resources, and they will directly impact the most revolutionary tools of the day. But, I see the relative opportunity to do good elsewhere as higher for plenty of people.</p><p>Open research will always be the standard that sets the language people use to understand AI. It&#8217;ll always be how the next generation is trained &#8211; even if it&#8217;s behind what industry has built. It&#8217;ll be the ecosystem where new long-shot ideas are built. Without investing in this open ecosystem, all of these cycles will be kneecapped.</p><p>At the end of the day, so much of my role now is just showing the path to impact in this domain. To show how clever, mid-sized open models can impact real problems in the world. To show how policy-makers and educators need open research to structure the rest of society around AI. This is a fun role too! It would be very sad for me to see this light diminish ever further, into the lightest embers of a fire that looks almost entirely out.</p><p>Even if the pace of research were to slow further, if the folks remaining like myself got financial offers they can&#8217;t refuse for their families&#8217; sake, the torch of open research will never fully go out. It&#8217;s core to how science is taught and done. There is a next generation coming, they just look for guidance and role-models.</p><h2><strong>What&#8217;s next</strong></h2><p>I see the best Ai2 work as research infrastructure. Building recipes in public gives countless researchers the ability to ask very specific questions of training processes. We need these researchers in the broader community, as Ai2 could never answer all the interesting questions themselves. One of my great joys in recent months has been visiting a top ML university and hearing so many graduate students say they&#8217;re building on Olmo. This is how the world should work!</p><p>Going forward, I still plan to operate in similar spaces, fighting for open-science, imagining what the future of the open model ecosystem can be, and doing my best to make the social transition to an AI-native era smooth. I&#8217;m most excited by how you can train medium sized open models on specific tasks that become useful tools in complement to the frontier models &#8211; massively winning on price. I want to invest in the ecological diversity of open models and coordination across builders.</p><p>For something that isn&#8217;t surprising given my past focus areas, I&#8217;m watching the pace of releases from all labs open &amp; closed, and how they&#8217;re hillclimbing on super ripe new post-training veins (on-policy distillation, agentic workflows, etc.), it&#8217;s clear that fully-open post training recipes are about as far behind as they ever have been &amp; falling further behind. I&#8217;d like to fix this. It&#8217;s not 100% clear yet if I will this year, but I&#8217;ll try.</p><p>To do this best and to execute, mostly personally, I needed a new start and fresh perspectives. I&#8217;ll be carefully building what I&#8217;m doing next over the next few months and am eager to share more about it when I can. One of my close teammates at Ai2 shared this quote with me in a farewell card, and I found it very apt in where I&#8217;m going next.</p><blockquote><p>The object of life is not to be on the side of the majority, but to escape finding oneself in the ranks of the insane. &#8212; Marcus Aurelius</p></blockquote><p>Thank you all for your continued support.</p>]]></content:encoded></item><item><title><![CDATA[Open and closed models are on different exponentials]]></title><description><![CDATA[Where marginally higher intelligence drives value, and where it doesn't.]]></description><link>https://www.interconnects.ai/p/open-and-closed-models-are-on-different</link><guid isPermaLink="false">https://www.interconnects.ai/p/open-and-closed-models-are-on-different</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 01 Jun 2026 13:03:48 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/6b994b97-89a2-4101-ad81-aa3ffe274b75_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The largest debate that&#8217;ll define the future balance of power between the open and closed AI model ecosystems is primarily economic &#8212; it&#8217;s if users of AI will continue to pay dramatically more, i.e. large margins, for the top closed models. Early 2026 is a seminal time for the AI industry, as the coding agents<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> have shown the first area where a huge AI market will continue to pay a substantial premium for better intelligence. </p><p>The other side of this dichotomy is the inevitable decay of API businesses at these same labs. These labs will realize they need to protect their best models, rolling them out later in APIs to both protect token supply, avoid distillation, and stick to use-cases with higher margins. All of these effects will be clearly visible in 5-10 year timelines, as in the near term markets, prices, margins, and demand will be dictated by a rapid buildout of compute (supply-limited in the near term) and mass subsidization of tokens (through continued investment in new AI companies).</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/open-and-closed-models-are-on-different?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/open-and-closed-models-are-on-different?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>The core of this argument rests in the obvious habit changes that are setting in with coding agents past the Opus 4.5 and Codex 5.2 thresholds. People are not making this switch because they are lazy, but because their net output is obviously higher when using an agent as an implementation aid for complex knowledge work. For people who rely on coding agents to work, they will always pay more for the best rather than settle for good enough. There are so many ways to make the product better, speed, intelligence, specialized models, etc. </p><p>I would pay $2000/month for the tools today, especially knowing they&#8217;ll get much better. At the same time, it is likely that many companies are forcing agents and usage onto people that actually will get very little out of them in their current form, which helps the AI buildout (or bubble) continue.</p><p>The best closed labs &#8212; right now this list is just Anthropic and OpenAI, but it&#8217;s reasonable to expect Google to catch up &#8212; will always make the most efficient models for intelligence at a given cost. Building models is a mass capital investment of talent, data, and compute. These systems, a combination of model weights, harnesses, tools, and serving infrastructure have massive returns on integration (where open models are designed to work across many, diverse serving situations). These integration benefits &#8212; the integration of hardware and new forms of software &#8212; can be expressed in any possible way of making models better. </p><p>The models in the near future may saturate on benchmark scores, but if that intelligence ceiling really is a cap on utility then the labs will optimize utility per second or per watt, serving users in another way. Improving the models is possible in every direction &#8212; there have been no walls in progress. We&#8217;re early in the mass buildout of intelligence, which involves harnessing the physical world to build numerous datacenters, organizing many AI researchers so that a large team can contribute to one model, and of course solving many small, low-level puzzles that unlock performance. Every indication is that there is still meaningful performance to be unlocked and the closed labs are the best set up to extract it.</p><p>The collective wisdom of the labs is that making the models smarter, in terms of the frontier of absolute intelligence, has the most value. This is the right call to me because it unlocks large new markets. Optimizing models at a fixed intelligence level locks in markets, expands accessibility over time, and increases return on investment for users (while potentially lowering margins for selling intelligence).</p><p>Many people are making this bet that models will keep getting better and are learning to work well in these harnesses, even though some workflows are still a bit clunky. This is the <a href="https://www.interconnects.ai/p/get-good-at-agents">right bet</a>. These people all will continue to use the absolutely best models available. It&#8217;s like buying an iPhone as a consumer. You could get an Android and suffer from a bunch of paper cuts to save money, but why would you? The returns to performance are even higher in the workplace, which drives pricing power.</p><p>In this mental model, the frontier labs as businesses, will look like new, reimagined forms of a mix of Apple and Microsoft. The Apple side is that they&#8217;re selling an integrated, extremely hard to replicate technology. The Microsoft side is selling high-leverage subscriptions across the economy. In 5-10 years I expect both OpenAI and Anthropic to be valued in the $2-10T range. The true frontier labs will be an oligopoly that looks like the cloud market today.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>On the other side of this equation is the <a href="https://www.interconnects.ai/p/the-next-phase-of-open-models?utm_source=publication-search">open model economy</a>. This isn&#8217;t to say that the frontier labs will dominate all aspects of AI use. Yes, I expect OpenAI and Anthropic to be the most representative companies of the AI boom (new companies, alongside Nvidia of course), but the collective value capture around open models will be far bigger overall, it&#8217;s just that the revenue and margins will be shared across a wide stack of companies.</p><p>Many businesses want to switch to open models but the models today are not good enough in out-of-distribution tasks. Eventually open model builders will stop chasing Claude and GPT on the Artificial Analysis index and fill this niche. This fork could be driven by economic factors, where they no longer have the revenue to support the growing R&amp;D costs for continuing to scale models. It can also be driven by pure demand, where certain AI solutions only can exist at low price points present in open models. Where closed labs are an oligopoly, open model builders and users will be far more diverse and numerous. The total market value will dramatically exceed the cumulative value of OpenAI and Anthropic.</p><p>Open models are by their nature <em>not</em> integrated, so they will rely on multiple companies coordinating to serve them. Each of these layers will have alternatives, driving prices down to commodity pricing. These low, predictable prices will be where many enterprises enter to build in-house agents and tools for niche tasks. The predominant mode of deployment here is that enterprises find a model that hits a sufficient performance threshold on a task of interest and does not replace the model later (setup costs are high). As customizing models becomes easier, again in the open model finetuning stack we are seeing emerge (Tinker, Fireworks, Prime Intellect, etc.), this market becomes even bigger.</p><p>What this will look like in the coming years is a steady rise in open model inference proportion across the entrenched hyper-scale clouds of Google, Amazon, Microsoft and new AI infrastructure companies of Together, Fireworks, OpenRouter, etc when compared to OpenAI and Anthropic.</p><p>The key is that the open and closed model economies are operating on different exponentials. I still believe that progress will continue at a fast pace <em>across the entire ecosystem</em>, but <a href="https://www.interconnects.ai/p/lossy-self-improvement?utm_source=publication-search">claims of recursive self improvement (RSI) giving the closed labs an unassailable advantage are overblown</a>. New forms of products like background agents can support both these open and closed models.</p><p>The closed models hit incredible product-market fit with the current agents, starting their integrated exponential by monetizing the top end of the knowledge work. The open model economy will take far longer, but it will also be far more satisfying to follow, as it tracks the broader diffusion of AI into the entire economy and world.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>The term coding agent is funny because we barely write code in them. They&#8217;re general agents that are so capable because they write a lot of code.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Some ideas for what comes next, May 2026]]></title><description><![CDATA[Gemini Flash 3.5, Mythos, open-closed balance, America's open-source surge, emerging power struggles and more.]]></description><link>https://www.interconnects.ai/p/some-ideas-for-what-comes-next-may</link><guid isPermaLink="false">https://www.interconnects.ai/p/some-ideas-for-what-comes-next-may</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Tue, 26 May 2026 15:39:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-711!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As the years of AI progress go by, it&#8217;s been accompanied by a slowly rising tide of consequence. Models are getting more capable, how we work is changing quickly, economics of AI are becoming real, just as real-world risks come to the forefront. 2026 is the first year where I don&#8217;t think there&#8217;ll be any breaks from this. The hard part to prepare for is that there&#8217;s a good chance things just continue to ratchet up from here &#8211; more disruption, more surprises, more stakes.</p><p>On my end, there&#8217;s been a growing list of topics that are very fateful to how I see the current state of AI, but I haven&#8217;t even gotten to write about them (at least not from all the angles I want to)! All of these are closely related to the implications of different models reaching new capability levels and how I use that to infer what may come next.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/some-ideas-for-what-comes-next-may?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/some-ideas-for-what-comes-next-may?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>1. Open models haven&#8217;t had their true agent moment like Opus 4.5</h3><p>The time gap between open and closed models is very often discussed, but the reality is that we have a nice time-gating that&#8217;s independent of debatable benchmarks &#8211; if open-weight models do or do not become super useful in agentic harnesses. The <a href="https://www.interconnects.ai/p/claude-code-hits-different">Opus 4.5 in Claude Code moment</a> of December 2025 was so loud and obvious, that if open models hit this performance level for price points as low as $5/month, there will be an explosion in usage.</p><p>Right now we are about 5-6 months in with no equivalent open model. I suspect the robustness of the best closed frontier models that I write about could make this moment take a good amount longer, say closer to 12+ months. In this time, Claude Code and Codex may seem like different categories of products. In the standard flurry of new, state-of-the-art open models from a variety of labs, benchmarks will definitely keep climbing, but the open-closed gap should become more interpretable as real-world use becomes the real litmus test.</p><h3>2. Gemini still doesn&#8217;t have a meaningful competitor for Claude Code and Codex</h3><p>The best exclamation point I can offer to reinforce my prediction that open models are further behind than the benchmarks claim is that even the mighty Google doesn&#8217;t have a clear competitor for Claude Code and Codex. I&#8217;m sure the Gemini team is pushing very hard on this.</p><p>I still need to do a lot more testing on Gemini 3.5 Flash, but reading reviews makes it clear that it&#8217;s not a substitute for how I&#8217;m working today. It&#8217;s maybe not the Gemini team explicitly specializing for Google&#8217;s existing products (search, YouTube, etc.), but the model seems to suit them. If Google doesn&#8217;t have a powerful tool here soon, I don&#8217;t expect the open model labs to either. The open models are going to be used more for automated, enterprise agents and low-cost domains, rather than being the driving tool of modern knowledge work. This will feed directly into the economic engine of funding future models, where the agents like Claude Code and Codex are the current best path to massive AI revenue growth.</p><p><em>I discussed how the current environment is quietly driving labs in China to specialize on <a href="https://aiproem.substack.com/p/nathan-lambert-reflects-on-chinas">AI Proem</a> with Grace Shao and this is central to my <a href="https://www.interconnects.ai/p/the-next-phase-of-open-models">expectations of open models specializing</a> over the next few years instead of competing with OpenAI, Anthropic, and Google.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>3. I don&#8217;t expect an open-weights Mythos this year</h3><p>While I don&#8217;t think Mythos is a general &#8220;god model&#8221; that will crush the competition in every domain, I do think it&#8217;s a remarkable technical achievement in software engineering and cybersecurity. Mythos is obviously a watershed moment for those fields. Having spoken to most of the Chinese labs &#8211; particularly those with the most prominent, large, open MoE models like Kimi, Z.ai, DeepSeek, and Qwen &#8211; I think they&#8217;re heavily resource limited and don&#8217;t have an immediate path to scaling up training processes like the big labs in the U.S. For the labs which are more corporate, which comes with more resources, such as Alibaba and Bytedance, they also have more conservative stances on safety and security.<br><br>Mythos is a bellwether of the massive acceleration in training and research compute available to the largest American companies.</p><p><em>Epoch AI recently had a nice <a href="https://epoch.ai/gradient-updates/frontier-labs-dont-use-most-ai-compute">piece</a> on the compute available to various labs (~Google 25%, Meta 11%, OpenAI 11%, Anthropic 6%). All of these numbers are vastly higher than any Chinese lab.</em></p><h3>4. American open models are slowly gaining steam</h3><p>Nvidia with Nemotron, Google with Gemma, Arcee AI and others are slowly stabilizing the open model ecosystem in the U.S. There&#8217;s a lot that&#8217;s hard to measure here, especially in the rise of local agents like OpenClaw and Hermes, but there are adoption numbers of American models that we haven&#8217;t seen since Llama 3.<br><br>Gemma 4&#8217;s models are all tying or outperforming the equivalently sized Qwen 3.5/3.6 models &#8212; where Qwen has for years now been the default open model at these sizes. These Qwen 3.5/3.6 models have been tricky to get working in a lot of post-training research, partially due to architecture/tooling and partially likely due to modeling (i.e. the model is not easy to finetune for some training decision). I&#8217;ve heard few complaints about Gemma, but it also could be because Gemma is not yet the <em>researcher</em> default.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-711!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-711!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png 424w, https://substackcdn.com/image/fetch/$s_!-711!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png 848w, https://substackcdn.com/image/fetch/$s_!-711!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png 1272w, https://substackcdn.com/image/fetch/$s_!-711!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-711!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png" width="1456" height="1008" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1008,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:207124,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.interconnects.ai/i/199119723?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-711!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png 424w, https://substackcdn.com/image/fetch/$s_!-711!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png 848w, https://substackcdn.com/image/fetch/$s_!-711!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png 1272w, https://substackcdn.com/image/fetch/$s_!-711!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F723ba9cb-c351-4b89-8860-2ac4eda7f335_2288x1584.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There's a simple reality that we've seen recently with models like GPT-OSS, Nemotron 3, and now Gemma 4, that if a model is in the right range of benchmarks and released by an American lab with a truly permissive license, it'll get a large amount of adoption (in this cycle, recall that Gemma 4 adopted the Apache 2.0 License, changing from one with use-case restrictions on earlier Gemmas). This early phase of American growth in open models is establishing key brands directly with developers. The consensus is that more neolabs like Reflection and Thinking Machines are likely to participate in this space, but being too patient will lose the time when new agentic workflows and enterprise relationships are built.</p><h3>5. Anthropic and OpenAI are just getting up to speed in model iterations</h3><p>I expect the rest of this year to be a ruthless competition between these two flagship companies. I&#8217;m at an interesting balance where I think GPT 5.5 is a bit smarter of a model and I love the Codex App, so I&#8217;m structuring much of my work to be possible there. At the same time, for a lot of writing-related and broader surface area tasks I really still love Claude. These models are rapidly changing how we work, I run Codex from my phone while doing other things, am setting up automated open model analysis jobs on the back of agents, and expect to be able to scale the research side of Interconnects widely.</p><p>AI is beginning to drive companies to the two extremes in the scaling era. The biggest companies will be way bigger than ever, using resources and mass talent to have sustained progress at the frontier of raw AI capabilities. On the other side, tiny businesses like Interconnects thrive by using agents to refine, present, and sell niche expertise. The mass social job displacement that&#8217;ll come is going to reduce employability for various knowledge workers that don&#8217;t fit into either of these extremes for the raw technical side (big or small companies), while sustaining and maybe even amplifying careers that interface directly with humans (e.g. doctors) or other power structures with means to sustain themselves (law/government).</p><h3>6. More existing power structures will assert themselves on AI</h3><p>Just in the last few days while writing this, we had the Pope release <a href="https://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.html">an over 40,000 word document</a> on where AI is going<em><strong> </strong></em>and <a href="https://www.bloomberg.com/news/articles/2026-05-26/china-expands-travel-curbs-to-top-ai-talent-at-private-firms">China expand personnel movement restrictions</a> on top AI researchers across industry. At the same time, the U.S. has <a href="https://www.axios.com/2026/04/19/nsa-anthropic-mythos-pentagon">designated Anthropic a supply chain risk and continues to use its models for national security</a>. The list of news like this is only going to grow. Existing power structures are realizing there&#8217;s a finite time window for them to exert themselves in the AI dynamic &#8212; an intuition that could be mapped to influence going down as AI models get more powerful. This intuition is potentially dangerous, as it sets up meaningful conflict in who controls the technology (as I <a href="https://www.interconnects.ai/p/how-anthropic-vs-dow-impacts-open">discussed</a> with Dean Ball after the Anthropic-DoW spat).</p><div><hr></div><h3>Next: Where technical becomes social</h3><p>These largely technical and <em>power</em> trends accelerating are going to put more pressure on the social and political anti-AI sentiments within the U.S. This is currently the most obvious barrier to continued AI development and beneficial diffusion. Reflecting on this, many people in the tech discourse get too focused on the details, where yes a lot of data-center-detractors are making genuinely wrong factual claims in defense of their position. </p><p>The real position that a large swath of Americans has is that they have a voice in saying no to the current trend &#8212; by not granting permission to build data centers. This is a voice that they haven&#8217;t been granted by the tech industry that changed the face of the global economy and power structures in the last few decades. </p><p>This is setting us up for a challenging year ahead for the industry. The labs are aggregating and concentrating talent to peak levels. There are few neutral messengers to communicate the reality of AI to the public. The frontier labs leadership is largely gearing up to IPO and stay ahead in the capabilities race. With the status quo, there are few actions to unwind this <a href="https://jasmi.news/p/warning-shots">path toward social conflict</a>. </p><p>It takes individuals in the AI ecosystem to zag and go against the groupthink of needing to make your wealth today, of needing to be at a lab to do impactful work, and so on. I&#8217;m personally continuing to bet on this, by trying to make a vibrant and diverse open model ecosystem supported by clear, unbiased information. If you agree with this and have been watching from the sidelines, it&#8217;s a good time to get involved, before the situation spirals into something uncontrollable.</p>]]></content:encoded></item><item><title><![CDATA[Latest open artifacts (#21): Open model bonanza! Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 & others. On CAISI's V4 assessment.]]></title><description><![CDATA[An eventful month with one flagship release after another]]></description><link>https://www.interconnects.ai/p/latest-open-artifacts-21-open-model</link><guid isPermaLink="false">https://www.interconnects.ai/p/latest-open-artifacts-21-open-model</guid><dc:creator><![CDATA[Florian Brand]]></dc:creator><pubDate>Sat, 16 May 2026 17:00:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!S79s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff60acac9-5993-474e-84f9-8805792bddef_1024x576.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This month was packed, with all open frontier labs, including DeepSeek, releasing new models. The latter prompted an evaluation by the <a href="https://www.nist.gov/news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro">Center for AI Standards and Innovation (CAISI)</a>, which has evaluated open models and their risks in the past. Their result is that open models lag behind the American frontier, with the gap becoming wider over time:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4DPW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4DPW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png 424w, https://substackcdn.com/image/fetch/$s_!4DPW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png 848w, https://substackcdn.com/image/fetch/$s_!4DPW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png 1272w, https://substackcdn.com/image/fetch/$s_!4DPW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4DPW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png" width="1456" height="965" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:965,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Comparison of aggregate capabilities over time of the most capable publicly released U.S. and PRC models according to a suite of benchmarks covering five domains.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Comparison of aggregate capabilities over time of the most capable publicly released U.S. and PRC models according to a suite of benchmarks covering five domains." title="Comparison of aggregate capabilities over time of the most capable publicly released U.S. and PRC models according to a suite of benchmarks covering five domains." srcset="https://substackcdn.com/image/fetch/$s_!4DPW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png 424w, https://substackcdn.com/image/fetch/$s_!4DPW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png 848w, https://substackcdn.com/image/fetch/$s_!4DPW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png 1272w, https://substackcdn.com/image/fetch/$s_!4DPW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>For the report, they calculate an Elo score based on <a href="https://en.wikipedia.org/wiki/Item_response_theory">Item Response Theory</a>, which is commonly used to compare different models, even when they were tested on a different set of benchmarks. For V4, CAISI used nine different benchmarks:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cVg4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cVg4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png 424w, https://substackcdn.com/image/fetch/$s_!cVg4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png 848w, https://substackcdn.com/image/fetch/$s_!cVg4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png 1272w, https://substackcdn.com/image/fetch/$s_!cVg4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cVg4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png" width="1456" height="1209" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1209,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:215444,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.interconnects.ai/i/197676648?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cVg4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png 424w, https://substackcdn.com/image/fetch/$s_!cVg4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png 848w, https://substackcdn.com/image/fetch/$s_!cVg4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png 1272w, https://substackcdn.com/image/fetch/$s_!cVg4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The huge Elo difference is explained by DeepSeek V4s bad score in CTF-Archive-Diamond (which was only run with a subset of the benchmark and extrapolated with IRT for V4), PortBench (a CAISI-private benchmark) and ARC-AGI-2 (with a different scoring method than the public leaderboards). The differences in these benchmark have a huge impact on the overall Elo, which can exacerbate the difference in capabilities. </p><p>When using <a href="https://epoch.ai/eci">Epoch AI&#8217;s ECI</a>, which also uses IRT over a set of different benchmarks, we see that the gap roughly stays between 3-7 months since R1:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qx4F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qx4F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png 424w, https://substackcdn.com/image/fetch/$s_!qx4F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png 848w, https://substackcdn.com/image/fetch/$s_!qx4F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png 1272w, https://substackcdn.com/image/fetch/$s_!qx4F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qx4F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png" width="1456" height="910" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:910,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2762591,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.interconnects.ai/i/197676648?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qx4F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png 424w, https://substackcdn.com/image/fetch/$s_!qx4F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png 848w, https://substackcdn.com/image/fetch/$s_!qx4F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png 1272w, https://substackcdn.com/image/fetch/$s_!qx4F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The open&lt;&gt;closed gap in ECI (from https://mcnair.center/china/)</figcaption></figure></div><p>However, both CAISI and ECI paint an incomplete picture, as both use standardized (and simple) setups to compare the capabilities of models. To be more concrete: Coding tasks are evaluated using access to bash and a for-loop with a fixed budget of tokens, not with a harness such as Claude Code or OpenCode, which models are trained in! These setups result in benchmarks claiming that porting applications to another language is <a href="https://programbench.com/">currently not possible</a>, while <a href="https://github.com/oven-sh/bun/pull/30412">Bun has been ported from Zig to Rust with 1 million LOC changes</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>.</p><p>Therefore, we would argue that a frontier comparison of open and closed models would also need to elicit the capabilities of all models better, which means the usage of the preferred harnesses, as well as model-specific prompting.</p><p>This section was written primarily by Florian. An interesting dynamic within Interconnects is that Florian believes more in the proximity of open frontier models to closed alternatives in true performance. Nathan thinks the benchmarks are imperfect as well, but thinks the closed models are ahead by more. We&#8217;re going to continue to unpack this in our future content.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/latest-open-artifacts-21-open-model?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/latest-open-artifacts-21-open-model?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3><strong>Our Picks</strong></h3><ul><li><p><strong><a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro">MiMo-V2.5-Pro</a></strong> by <a href="https://huggingface.co/XiaomiMiMo">XiaomiMiMo</a>: Avid Artifacts readers know that Xiaomi has been working on open models for a while; its debut was <a href="https://www.interconnects.ai/p/latest-open-artifacts-10-new-deepseek">exactly one year ago</a>. The progress of its releases is remarkable, with 2.5 Pro (released under Apache 2.0) being neck and neck with other flagship models such as Kimi K2.6 and GLM-5.1 in both benchmarks and <a href="https://x.com/Designarena/status/2054776484833952000?s=20">real-world usage</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!25fp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!25fp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg 424w, https://substackcdn.com/image/fetch/$s_!25fp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg 848w, https://substackcdn.com/image/fetch/$s_!25fp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!25fp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!25fp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg" width="1200" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!25fp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg 424w, https://substackcdn.com/image/fetch/$s_!25fp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg 848w, https://substackcdn.com/image/fetch/$s_!25fp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!25fp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></li><li><p><strong><a href="https://huggingface.co/google/gemma-4-26B-A4B-it">gemma-4-26B-A4B-it</a></strong> by <a href="https://huggingface.co/google">google</a> (full Interconnects post <a href="https://www.interconnects.ai/p/gemma-4-and-what-makes-an-open-model">here</a>): The long-awaited update to the Gemma series, featuring multiple sizes: <a href="https://huggingface.co/google/gemma-4-E2B-it">4B</a>, <a href="https://huggingface.co/google/gemma-4-E4B-it">9B</a>, and <a href="https://huggingface.co/google/gemma-4-31B-it">31B</a> dense models, as well as a 26B-A4B MoE. Even more importantly, with Gemma 4, Google has decided to use Apache 2.0 as its license, which removes the uncertainty and legal challenges around interpreting custom licenses.</p></li><li><p><strong><a href="https://huggingface.co/moonshotai/Kimi-K2.6">Kimi-K2.6</a></strong> by <a href="https://huggingface.co/moonshotai">moonshotai</a>: An update to the Kimi series, delivering stronger performance across the board and making it one of the best open models out there yet again. They also focus on long-horizon performance, showing that open models are capable of running over hours to complete tasks or optimize performance. Given the focus of everyone to build <a href="https://github.com/karpathy/autoresearch">autoresearch</a>-like systems, seeing open models catch up is important.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K_mJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K_mJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp 424w, https://substackcdn.com/image/fetch/$s_!K_mJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp 848w, https://substackcdn.com/image/fetch/$s_!K_mJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp 1272w, https://substackcdn.com/image/fetch/$s_!K_mJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K_mJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp" width="1456" height="932" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:932,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;K2.6 Qwen3.5-0.8B Mac inference optimization case&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="K2.6 Qwen3.5-0.8B Mac inference optimization case" title="K2.6 Qwen3.5-0.8B Mac inference optimization case" srcset="https://substackcdn.com/image/fetch/$s_!K_mJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp 424w, https://substackcdn.com/image/fetch/$s_!K_mJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp 848w, https://substackcdn.com/image/fetch/$s_!K_mJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp 1272w, https://substackcdn.com/image/fetch/$s_!K_mJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></li><li><p><strong><a href="https://huggingface.co/poolside/Laguna-XS.2">Laguna-XS.2</a></strong> by <a href="https://huggingface.co/poolside">poolside</a>: Poolside AI has released its first public coding-focused models, including the open-weight XS.2. Its size (33B-A3B) makes it attractive for local use, with performance on par with other models in that size range. The accompanying <a href="https://poolside.ai/blog/laguna-a-deeper-dive">blog post</a> is worth a read, as is <a href="https://poolside.ai/blog/through-the-looking-glass">the deep dive</a> into reward hacking during coding evaluations.</p></li><li><p><strong><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash">DeepSeek-V4-Flash</a></strong> by <a href="https://huggingface.co/deepseek-ai">deepseek-ai</a>: DeepSeek has finally released its successor to the V3 series, which it kept updating for months. It comes in two sizes: Pro, which is a 1.6T-A49B MoE, and Flash, a 284B-13B model. Based on others&#8217; experience, the latter model seems to be the real star of the show, as its performance is relatively strong, while Pro seems to underdeliver relative to its size. The <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf">tech report</a> goes into great detail, including the architectural changes used to achieve better and cheaper long-context performance.</p></li></ul><h3><strong>Models</strong></h3><h4>General Purpose</h4><ul><li><p><strong><a href="https://huggingface.co/Qwen/Qwen3.6-35B-A3B">Qwen3.6-35B-A3B</a></strong> by <a href="https://huggingface.co/Qwen">Qwen</a>: An update to the Qwen 3.5 series targeting one of the most widely used sizes.</p></li><li><p><strong><a href="https://huggingface.co/LiquidAI/LFM2.5-350M">LFM2.5-350M</a></strong> by <a href="https://huggingface.co/LiquidAI">LiquidAI</a>: With 28T tokens for 350M parameters, this model might be the most overtrained model out there.</p></li><li><p><strong><a href="https://huggingface.co/arcee-ai/Trinity-Large-Thinking">Trinity-Large-Thinking</a></strong> by <a href="https://huggingface.co/arcee-ai">arcee-ai</a>: The reasoning version of Trinity, one of the best Western open models. It has topped the OpenRouter charts for a while and can power agentic applications such as OpenClaw.</p></li><li><p><strong><a href="https://huggingface.co/zai-org/GLM-5.1">GLM-5.1</a></strong> by <a href="https://huggingface.co/zai-org">zai-org</a>: An update to GLM-5, improving scores across the board. The focus for this update is on long-horizon tasks.</p></li></ul>
      <p>
          <a href="https://www.interconnects.ai/p/latest-open-artifacts-21-open-model">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[How open model ecosystems compound]]></title><description><![CDATA[Further reflections on China's high-participation, open-first AI ecosystem.]]></description><link>https://www.interconnects.ai/p/how-open-model-ecosystems-compound</link><guid isPermaLink="false">https://www.interconnects.ai/p/how-open-model-ecosystems-compound</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Tue, 12 May 2026 15:54:47 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/39c7cc76-02ac-4c38-bbab-7d89dca53d0b_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h5>Note: Voice-overs for paywalled posts are available for paid subscribes in podcast apps if you click on settings on Interconnects, then manage your description. Thanks for listening!</h5><p>Most of the compute to build a leading frontier model comes from R&amp;D costs, rather than the compute to train the final, big model end-to-end. In an ecosystem like China, where all the leading players are open, this creates a potential meaningful advantage in cost structures that&#8217;ll let labs keep building longer than outside observers would expect.</p><p>There are two recent pieces of research, one from Ai2 <a href="https://arxiv.org/abs/2605.01158">documenting the development of Olmo 3</a> and one from Epoch AI <a href="https://epoch.ai/gradient-updates/r-and-d-vs-training-compute">studying public documentation of costs from various frontier labs</a>, that put the estimate of compute spent on R&amp;D rather than the final model at about 80% (with meaningful error bars).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VBXD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VBXD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png 424w, https://substackcdn.com/image/fetch/$s_!VBXD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png 848w, https://substackcdn.com/image/fetch/$s_!VBXD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png 1272w, https://substackcdn.com/image/fetch/$s_!VBXD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VBXD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png" width="1026" height="1283" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1283,&quot;width&quot;:1026,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!VBXD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png 424w, https://substackcdn.com/image/fetch/$s_!VBXD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png 848w, https://substackcdn.com/image/fetch/$s_!VBXD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png 1272w, https://substackcdn.com/image/fetch/$s_!VBXD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e980106-980c-4092-bbf1-f472323b8da0_1026x1283.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In a world where research and development is most of the compute, the Chinese system is designed around quickly learning from your peers and avoiding double-spending research compute &#8212; or infra effort.  It&#8217;s far from perfect, but it&#8217;s the closest analog to the OSS ecosystem that one can get for building LLMs. The public discussion of AI has always emphasized that the <em>models</em> are expensive in a way that naturally lets passive readers think this is compute just dedicated to the artifact &#8212; <a href="https://www.interconnects.ai/p/deepseek-v3-and-the-actual-cost-of">as we saw with DeepSeek V3</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/how-open-model-ecosystems-compound?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/how-open-model-ecosystems-compound?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>This had me revisiting the core issue of open-source AI, and how it doesn&#8217;t have the feedback loops akin to open-source software (OSS) users back to the creation itself, that creates immense value following <a href="https://en.wikipedia.org/wiki/Linus%27s_law">Linus&#8217;s law</a> of &#8220;given enough eyeballs, all bugs are shallow&#8221;. This self-reinforcement of OSS makes deployment at scale the cheapest possible outcome &#8212; all the users together share the costs of fixing bugs and adding features.</p><p>Within open-source AI, almost all the cost falls on the model developer. At the same time, there are huge benefits to releasing the model openly that do reduce costs, but they only help reduce <em>future</em> development and deployment costs for the creator themselves, but more importantly the ecosystem widely.</p><p>Open AI models, tools, infrastructure, and everything in between are a cost reduction in development, not plug and play cost reduction on apples to apples solutions or products. If someone is going to be just using AI off-the-shelf with minimal iteration or internal development, using open models will almost always be more expensive. Using closed, integrated, hosted solutions achieves low price points by economies of scale across general usage.</p><p>The open-source ecosystem can only try to mirror the OSS-style financial and performance gain in continued performance. The Chinese labs, through incredibly thorough technical reports and intentional knowledge sharing across labs effectively are de-risking ideas for their peer companies to not necessarily need to invest as many resources in.</p><p>For this to work, the current norm where AI companies <em>fork</em> open-source tools, to evolve them into internal-only versions, will likely need to fade out. It&#8217;s too common of a trope for open-source AI companies to have their selling point being better performance via enterprise agreements or internal tools, as the fully open tools that people start with are falling behind in accessibility. A prime example is at-scale RL training of MoE models &#8212; no truly open recipe exists. It&#8217;s unclear if the open-supporting, but partially closed tools like Thinking Machine&#8217;s <a href="https://thinkingmachines.ai/tinker/">Tinker</a> and Prime Intellect&#8217;s <a href="https://www.primeintellect.ai/blog/lab">Lab</a> can be open enough for the advantages of an open ecosystem to sustain themselves. The more open the stack is, and the more information is shared, the more costs are reduced in future iterations.</p><p>The same reasoning that causes companies to fork open-source tools to make internal versions applies to why there isn&#8217;t a shared, single foundation model that everyone builds on. Building the best model today becomes an art of integrating your hardware, data, and infrastructure, while evolving all of them at a relatively high rate that lets you keep up with the frontier of performance. Given that all signs point to LLMs continuing their steady march in performance improvements for years, it seems unlikely to expect this equilibrium to change in the near term. This is exactly why I wrote my post on the <a href="https://www.interconnects.ai/p/the-inevitable-need-for-an-open-model">inevitable need for an open model consortium</a> &#8211; this shared resource is far more efficient and may become the only financially viable way to compete at the future frontier scale with open models.</p><p>It&#8217;s worth noting that, of course, the closed labs also see the investigations of the open frontier model companies and can benefit from them, but with the assumption that the closed labs are <a href="https://www.interconnects.ai/p/reading-todays-open-closed-performance">some months ahead in the development tree</a>, they often naturally stand to benefit less from the shared insights. The stronger the open-source community is, the more cost incentive there is for the various companies to be relatively close together on the same Pareto curve of performance.</p><p>This realization of the difference between <em>development</em> costs, or a process-focused technology, rather than some shared foundation that all the labs build on directly was downstream of a question I got in feedback to <a href="https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs">my recent China trip summary</a>. The question was: &#8220;Was there any chance of the Chinese ecosystem converging on a single base model to save costs?&#8221; The follow-up to this question was on if any of the open-weight companies in China are using open-source in strategically meaningful ways. There are many more useful questions to ask here, especially when trying to understand the different operational patterns of the ecosystems.</p><h2>China&#8217;s foundation model development model</h2><p>I found the following interview conducted by Bill Gurley with Dan Wang, author of Breakneck, and Patrick McGee, author of Apple in China, (both books I strongly recommend &#8211; must reads) very thought provoking on the biggest differences between technology cultures in the U.S. and China.</p><div id="youtube2-XpyqKn_1ZP4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;XpyqKn_1ZP4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/XpyqKn_1ZP4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>I get a lot of exposure to these differences at this point in my open-source AI arc. There&#8217;s a deep yearning to influence Western audiences and thinking that has bubbled up out of the Chinese AI ecosystem in the last year. This was obviously a strong pretext for why the <a href="https://readsail.com/">SAIL</a> group got such access in our recent trip &#8211; it&#8217;s not a given that anyone in the AI ecosystem will talk to senior leadership at so many companies.</p>
      <p>
          <a href="https://www.interconnects.ai/p/how-open-model-ecosystems-compound">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Notes from inside China's AI labs]]></title><description><![CDATA[Lessons from my trip to talk to most of the leading AI labs in China.]]></description><link>https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs</link><guid isPermaLink="false">https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Thu, 07 May 2026 15:42:43 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2b353f46-1b83-4750-9dc0-72877a402f19_1024x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Staring out the window on a new, high-speed train from Hangzhou to Shanghai I&#8217;m gifted with views of dramatic ridgelines speckled with wind turbines that are silhouetted against the setting sun. The mountains cast a backdrop to a mix of spanning fields and clustered skyscrapers. I&#8217;m returning from China with great humility. It&#8217;s a very warming, human experience to go somewhere so foreign and be so welcomed. I had the honor of meeting so many people in the AI ecosystem who I knew from afar, and they greeted me with big smiles and cheer, reminding me how global my work and the AI ecosystem is.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The mentality of Chinese researchers</h2><p>The Chinese companies building language models are set up as the perfect fast-followers for the technology, building on long-standing cultural traditions in education and work, along with subtly different approaches to building technology companies. When you look at the outputs, the latest, biggest models enabling agentic workflows, and the ingredients, excellent scientists, large-scale data, and accelerated computing, the Chinese and American labs look largely similar. The lasting differences emerge in how these are organized and conditioned.</p><p>I&#8217;ve long thought that a reason that the Chinese labs are so good at catching up and keeping up with the frontier is that they&#8217;re culturally aligned for this task, but without talking to people directly I felt like it wasn&#8217;t my place to attribute substantial influence to this hunch. Speaking with many wonderful, humble, and open scientists at the leading Chinese labs has crystallized a lot of my beliefs.</p><p>So much of building the best LLMs today comes down to meticulous work across the entire stack, from data to architecture details and RL algorithm implementations. All points of the model can give some improvements, and fitting them in together is a complex process where the work of some brilliant individuals needs to get shelved in favor of the overall model maximizing a multi-objective optimization.</p><p>Where American researchers are obviously also brilliant at solving the individual components, there&#8217;s more of a culture of speaking up for yourself in the U.S. As a scientist, you&#8217;re more successful when you speak up for your work and modern culture is pushing the new path to fame of &#8220;leading AI scientists&#8221;. This results in direct conflict. The Llama organization is heavily rumored to have collapsed under the political weight of these interests embedding themselves in a hierarchical organization. I&#8217;ve heard of other labs saying that it can be needed to pay off a top researcher to get them to stop complaining about their idea not making it in the final model. Whether or not that&#8217;s exactly true, the idea is clear. Ego and desires for career advancement do get in the way of making the best models. A small, directional shift in this sort of culture between the U.S. and China can have a meaningful impact on the final outputs.</p><p>Some of this has to do with who is building the models in China. There&#8217;s an immediate reality at all of the labs that a large proportion of the core contributors are active students. The labs are quite young, and it reminds me of our setup at Ai2, where students are seen as peers and directly integrated in the LLM team. This is incredibly different from the top labs in the US, where the likes of OpenAI, Anthropic, Cursor, etc. simply don&#8217;t offer internships. Other companies like Google nominally have internships related to Gemini, but there&#8217;s a lot of concern about whether your internship will be siloed and away from anything real.</p><p>To summarize how the slight change in culture can improve the ability to build models:</p><ul><li><p>More willingness to do non-flashy work in order to improve the final model,</p></li><li><p>People new to building AI can be free of prior phases of AI hype cycles, allowing them to adapt to the new modern techniques faster (in fact, one of the Chinese scientists I talked to really actively attached to this strength),</p></li><li><p>Less ego enabling org charts to scale slightly, as there&#8217;s less gamifying the system, and</p></li><li><p>Abundant talent well-suited to solving problems with a proof of concept elsewhere, etc.</p></li></ul><p>This slight inclination towards skills that complement building today&#8217;s language models stands in contrast to a known stereotype that Chinese researchers tend to produce less creative, field-spawning, 0-to-1 academic style research. Among the more academic lab visits on our trip, many leaders talk about cultivating this more ambitious research culture. At the same time, some technical leaders we talked to were skeptical about whether such a rewiring in the approach to science is likely in the near term, because it&#8217;ll take a redesign of the education and incentive systems that is too big to happen within the current economic equilibrium. This culture seems to be training students and engineers that are excellent at the LLM building game. They also, of course, have an extremely abundant quantity.</p><p>These students told me about a similar brain drain happening in China as in the U.S., where many who previously considered academic paths now intend to stay in industry. The funniest quote was from a researcher who was interested in being a professor to be close to the education system, but remarked that education is solved with LLMs &#8211; &#8220;why would a student talk to me!&#8221;</p><p>The students have a benefit of coming at LLMs with fresh eyes. Over the last few years we&#8217;ve seen the key paradigm of LLMs shift from scaling MoE&#8217;s, to scaling RL, to enabling agents. Doing any of these well involves absorbing an insane amount of context quickly, both from the broader literature and the technical stack at your company. Students are used to doing this and excited to humbly drop all presumptions about what should work. They dive in head first and dedicate their life to getting the chance to improve the models.</p><p>These students are also so magically direct and free of some of the philosophical chatter that can distract scientists. When asking questions on how they feel about the economics or long-term social risks of models, far fewer Chinese researchers have sophisticated opinions and a drive to influence this. Their role is to build the best model.</p><p>This difference is subtle, and easy to deny, but it is best felt when having long conversations with an elegant, brilliant researcher who can clearly communicate well in English, basic questions on more philosophical aspects of AI hang in the air with a simple confusion. It&#8217;s a category error to them. One researcher even quoted the famous Dan Wang premise of China being run by engineers, relative to the lawyers of the U.S. when probing in these areas, to emphasize their desire to build. There&#8217;s no track in China that systematically enables the growth of star power for Chinese scientists, akin to mega mainstream podcasts like Dwarkesh or Lex.</p><p>Trying to get Chinese scientists to comment on the coming economic uncertainty fueled by AI, questions beyond the capabilities of simple AGI, or moral debates on how models should behave all served to capture the upbringing and education of these scientists (edited<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>). They are extremely dedicated to their work, but have grown up in a system where debates and opinions on how society should be structured and changed are not encouraged. </p><p>Zooming out &#8212; Beijing especially felt much like the Bay Area, where a competitive lab is a short walk or Uber away. I got off a flight and stopped by Alibaba&#8217;s Beijing campus on the way to the hotel. Then, in 36 hours we went to all of Z.ai, Moonshot AI, Tsinghua University, Meituan, Xiaomi, and 01.ai. Travel by Didi is easy, and if you select an XL in China you&#8217;re often paired with electric mini vans that have massage chairs. We asked the researchers about the talent wars, and they said it&#8217;s very similar to what we&#8217;re experiencing in the U.S. It&#8217;s normal for researchers to bounce around, and much of where people choose to go is based on the best current vibes.</p><p>In China, the LLM community feels far more like an ecosystem than battling tribes. Across many off the record conversations, it&#8217;s nothing but respect for peers. All of the Chinese labs fear Bytedance with their popular Doubao model, which is the only frontier closed lab in China. At the same time, all of the labs have massive respect for DeepSeek as the lab with the best research taste in execution. When you meet with lab members off the record in the States, sparks fly quickly.</p><p>The most striking part of the humility of Chinese researchers is how they also often shrug on the business side, saying it&#8217;s not their problem, where everyone in the U.S. seems to be obsessed with various ecosystem-level industrial trends, from data sellers to compute or fundraising.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2>Where China&#8217;s AI industry differs (and matches) the Western labs</h2><p>The thing that makes building an AI model today so interesting is that it&#8217;s not just about getting a group of great researchers in one building together to produce an engineering marvel. It used to be this, but to sustain AI businesses, the LLMs are becoming a mix of building, deploying, funding, and getting adoption for this creation. The leading AI companies exist in complex ecosystems that supply money, compute, data and more in order to keep pushing the frontier. </p><p>The integration of these various inputs to creating and sustaining LLMs is fairly well conceptualized and mapped for the Western ecosystem, as typified by Anthropic and OpenAI, so finding big differences in how the Chinese labs think about it points at where the different companies can be making meaningfully different bets on the future. Of course, these futures can be heavily dictated by the constraints on funding and/or compute.</p><p>I&#8217;ve documented the biggest &#8220;AI Industry&#8221; level take-aways from talking to these labs:</p><ol><li><p><strong>Early signs of domestic AI demand.</strong> There&#8217;s a much-touted hypothesis that the Chinese AI market will be smaller because Chinese companies don&#8217;t tend to pay for software &#8211; thus, never unlocking a giant inference market supporting labs. This is only true for software spend that maps to the SaaS ecosystem, which is historically tiny in China, where on the other hand there is obviously still a large cloud market in China. A crucial unanswered question &#8211; one which the Chinese labs themselves debate &#8211; on if spending for AI in the enterprise tracks the SaaS market (small) or the cloud market (fundamental). On net, it feels like AI is trending closer to the cloud, and no one was actively worried about a market growing around the new tools.</p></li><li><p><strong>Most developers are Claude-pilled.</strong> Most of the AI developers in China are obsessed with Claude and how it&#8217;s changed how they build software, despite Claude nominally being banned in China. Just because China has historically been hesitant to buy software does <em>not</em> give me the impression that there won&#8217;t be a massive surge in inference demand. Chinese technical staff are so practical, humble, and motivated &#8211; a fact that seems stronger than any commitment to previous habits in not spending.<br><br>Some Chinese researchers mention building with their own tools, such as the Kimi or GLM CLIs, but <em>all</em> of them mention building with Claude. There were also surprisingly few mentions of Codex, which is definitely surging in popularity in the Bay Area.</p></li><li><p><strong>Chinese companies have a technology ownership mentality.</strong> The Chinese culture is combining with a roaring economic engine to create unpredictable outcomes. I&#8217;m left with a lasting feeling that the numerous AI models reflect a practical, current equilibrium of the many technology businesses here. There&#8217;s no master plan. The industry is defined by a respect for ByteDance and Alibaba, the incumbents expected to win large portions of all markets with their substantial resources. DeepSeek is the respected technical leader, but far from a market leader. They set the direction, but aren&#8217;t set up to win economically.<br><br>This leaves companies like <a href="https://huggingface.co/meituan-longcat/LongCat-Next">Meituan</a> or <a href="https://huggingface.co/inclusionAI/Ling-2.6-1T">Ant Group</a>, where people in the West can be surprised they&#8217;re building these models. In reality, they see LLMs obviously as being central to future technology products, so they need a strong base. When they fine-tune the strong, general purpose model it hardens their stack from getting the open community to provide feedback on it, and they can keep internal, fine-tuned versions of the model for their products. The &#8220;open-first&#8221; mentality in the industry is largely defined by practicality &#8212; it helps make their models get strong feedback, it gives back to the open-source community, and empowers their mission.</p></li><li><p><strong>Government aid is real, but unclear how big.</strong> It&#8217;s often asserted that the Chinese government is actively helping with the open LLM race. This is a government that&#8217;s decentralized across many levels, each of which doesn&#8217;t have a clear playbook for what exactly they do. Neighborhoods in Beijing compete for tech companies to house their offices there. The &#8220;help&#8221; offered to these companies almost certainly involved removing bureaucratic red tape like permits, but how far does it go? Can levels of the government help attract talent? Can they help smuggle chips? Across the visit, there were many mentions of government interest or help, but far too little to report the details as assertive or have a confident worldview of how government can bend the trajectory of AI in China. <br><br>There were certainly no hints of the top levels of the Chinese government influencing any technical decisions in the models.</p></li><li><p><strong>The data industry is far less developed. </strong>Having heard so much about the likes of Anthropic or OpenAI spending $10M+ for single environments, with cumulative spend on the order of hundreds of millions per year to push the frontier of RL, we were eager to know if Chinese labs are either buying the same environments from companies in the U.S. or supported by a mirrored domestic ecosystem. The answer was not quite complete that there&#8217;s <em>no</em> data industry, but rather that their experience was that the data industry was relatively poor quality and it is often better to build the environments or data in-house. Researchers themselves spend meaningful time making the RL training environments, and some of the bigger companies like ByteDance and Alibaba can have in-house data labelling teams to support this. This all mirrors the build-not-buy mentality from the previous bullet.</p></li><li><p><strong>Desperation for more Nvidia chips. </strong>Nvidia compute is the gold-standard for training and everyone is limited in progress by not having more of it. If supply was there, it is obvious that they would buy it. Other accelerators, including but not limited to Huawei, were spoken positively of for inference. Countless labs have access to Huawei chips.</p></li></ol><p>These points paint a very different picture of an AI ecosystem, where quickly mapping how Western labs operate to their Chinese counterparts will often result in a category error. The crucial question is if these different ecosystems will produce meaningfully different types of models, or if the Chinese models will always be explained by being similar to the U.S. frontier models of 3-9 months ago.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/subscribe?"><span>Subscribe now</span></a></p><h2>Conclusion: The global equilibrium</h2><p>I knew so little about China going into the trip and came out with the feeling of just starting to learn. China isn&#8217;t a place that can be expressed by rules or recipes, but one with very different dynamics and chemistry. The culture is so old, so deep, and still completely intertwined with how domestic technology is built. I have much more learning ahead.</p><p>So much of the current power structures in the US use their current worldviews of China as crucial mental devices for decision making. Having talked, in person, either formally or informally to pretty much every leading AI lab in China, there are a lot of qualities and instincts in China that&#8217;ll be very hard to model with Western decision making. Even after asking directly about <em>why</em> these labs release their top models openly, the intersection between ownership mentality and genuine ecosystem support is hard for me to connect the dots on. </p><p>The labs here are practical and not necessarily absolutists around open-source, where every model they build would be released openly, but there&#8217;s a deep intentionality in supporting developers, the ecosystem, and using it as a way to learn more about their models.</p><p>Almost every major Chinese technology company is building their own general purpose LLMs, as we see with the likes of Meituan (delivery service) and Xiaomi (broad consumer technology company) releasing open weight models. The equivalent companies in the U.S. would just buy services. These companies aren&#8217;t building LLMs out of a race to be relevant with the hot new thing, but a deep fundamental yearning to control their own stack and develop the most important technologies of the day. When I look up from my laptop and always see bunches of cranes on the horizon, it obviously fits in the with the broader culture and energy around building in China.</p><p>The humanity, charm, and genuine warmth of Chinese researchers is extremely humanizing. At a personal level, the cut-throat geopolitical conversation we&#8217;re used to in the U.S. hasn&#8217;t permeated them at all. The world can use more of this simple positivity. As a citizen of the AI community, I currently worry more about the fissures appearing within members and groups around labels of nationality. </p><p>I&#8217;d be lying if I said I didn&#8217;t want US labs to be clear leaders in every part of the AI stack &#8212; especially with open models where I spend my time &#8212; I&#8217;m American, and that&#8217;s an honest preference. With this, I want the open ecosystem itself to thrive globally, as this can create safer, more accessible, and more useful AI for the world, and right now the question is whether American labs will take the steps to own that leadership position. </p><p>As of finishing this piece, more <a href="https://x.com/andrewcurran_/status/2052023542582292855?s=46">rumors</a> are swirling of executive orders influencing open models, which can further complicate this synergy between American leadership and the global ecosystem &#8212; it doesn&#8217;t fill me with confidence.</p><p>Thank you to all the wonderful people I got to talk to at Moonshot, Zhipu, Meituan, Xiaomi, Qwen, Ant Ling, 01.ai, and others. Everyone has been so welcoming and gracious with their time. I&#8217;ll keep sharing my thoughts on China as they crystallize, across culture generally and AI specifically. It is obvious that this knowledge will be directly relevant to the story unfolding at the frontier of AI development.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Edit 05/07: In this paragraph in the original I misattributed an unwillingness to speak on broader issues to humility, which can of course play a part, but this habit is also shaped by the system which they were trained and raised, a system they are successful in and adept at navigating.</p><p>What I removed: &#8230; capture the upbringing and education of these scientists extreme humility of these scientists. It&#8217;s more than just being dedicated to <em>their</em> work, but they don&#8217;t want to comment on issues they&#8217;re not informed on.&#8230;</p></div></div>]]></content:encoded></item><item><title><![CDATA[The distillation panic]]></title><description><![CDATA[&#8216;Distillation attacks&#8217; is a horrible term for what is happening right now.]]></description><link>https://www.interconnects.ai/p/the-distillation-panic</link><guid isPermaLink="false">https://www.interconnects.ai/p/the-distillation-panic</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 04 May 2026 15:56:44 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b94fcb01-2fef-44ce-9d1d-4b0d9b7a737e_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>&#8216;Distillation attacks&#8217; is a horrible term for what is happening right now. Yes, some Chinese labs are hacking or jailbreaking APIs to attempt to extract more signal from model APIs &#8212; stopping this is important to maintain the U.S.&#8217;s lead in AI capabilities. Referring to this as distillation attack is going to irrevocably associate all distillation with this behavior, and distillation generally is a core technique needed to diffuse AI capabilities broadly through academic and economic activities.</p><p>We went through this sort of language transition with the open source vs open weight debate. All the terms just reduced to open models &#8211; very few people in the large AI community know exactly how open-source differs from open-weights. And terminology matters, as the less informed people who still care about &#8212; and influence &#8212; the technology are bound by different terms they use. If we&#8217;re not careful with the discourse around distillation, many people could associate this broad technique used for research and development of new models as an act at the boundary of corporate manipulation and crime.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/the-distillation-panic?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/the-distillation-panic?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>I&#8217;ve recently written a more <a href="https://www.interconnects.ai/p/how-much-does-distillation-really">technical piece</a> on estimating how impactful state-of-the-art distillation methods are on leading Chinese models, and this piece follows to push for caution in any hasty actions to target the methods with policy. To set the stage, recall Anthropic&#8217;s recent blog post where they <a href="https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks">detailed &#8220;distillation attacks&#8221; made by 3 Chinese labs</a>.</p><blockquote><p>These labs used a technique called &#8220;distillation,&#8221; which involves training a less capable model on the outputs of a stronger one. Distillation is a widely used and legitimate training method. For example, frontier AI labs routinely distill their own models to create smaller, cheaper versions for their customers. But distillation can also be used for illicit purposes: competitors can use it to acquire powerful capabilities from other labs in a fraction of the time, and at a fraction of the cost, that it would take to develop them independently.</p></blockquote><p>This is a clever paragraph, where they normalize distillation generally and explain how a few people can use it illicitly, without detailing how illicit use often involves other more explicit behavior like jailbreaking, hacking, or identity spoofing of the API.</p><p>Distillation itself is an industry standard. It&#8217;s used extensively, primarily in post-training, by smaller players to create specialized or smaller models. In my <a href="https://rlhfbook.com/c/12-synthetic-data">book</a> coming this summer, I describe it as follows:</p><blockquote><p>The term distillation has been the most powerful form of discussion around the role of synthetic data in language models. Distillation as a term comes from a technical definition of teacher-student knowledge distillation from the deep learning literature.</p><p>Distillation colloquially refers to using the outputs from a stronger model to train a smaller model.</p><p>In post-training, this general notion of distillation takes two common forms:</p><ol><li><p>As a data engine to use across wide swaths of the post-training process: Completions for instructions, preference data (or Constitutional AI), or verification for RL.</p></li><li><p>To transfer specific skills from a stronger model to a weaker model, which is often done for specific skills such as mathematical reasoning or coding.</p></li></ol></blockquote><p>With this definition, it&#8217;s easy to see how distillation takes many forms. Of course, if you just take the outputs from GPT-5.5 and train a recent open-weight base model with them to host a competitive product, that&#8217;s one thing. But, a lot of the things that fall under the bucket of distillation are complex, multi-stage processes that muddle the exact impact of the model you distilled from.</p><p>Modern LLM processes could look like using a GPT API to build an initial batch of synthetic data to build a specialized small data-processing model. A good example is a model like olmOCR (or many other models in this category) that are trained to convert PDFs to clean text. This specialized model would be used to create large amounts of data. Finally, you train another model (often from scratch) with the new data you created. Is this final model distilled from GPT?</p><p>When done via a closed, API-based model, distillation sits in the grey area of the terms of service that you agree to when signing up to the Claude or GPT platform. They generally forbid the use of the API to create competing language model products, but this term has largely gone unenforced. The open-source community used to worry deeply at being cut off from these cutting-edge APIs for doing research or creating public datasets, but to date only <a href="https://www.theverge.com/2023/12/15/24003151/bytedance-china-openai-microsoft-competitor-llm">one prominent case of corporate accounts being restricted exists</a> (at least until the recent Chinese companies).</p><p>This is all to say that distillation is an industry standard technique, and the use of closed APIs to perform distillation has always been a grey area. Nvidia&#8217;s latest Nemotron models, as one of the only models with open post-training datasets, are technically in large part distilled from Chinese, open-weight models. The Olmo models we&#8217;ve built at Ai2 are distilled from a mix of open and closed models. This grey area was brought to the forefront again when it turned out that xAI has been distilling from OpenAI. Quoting from the recent trial <a href="https://x.com/MTSlive/status/2049886679876632724">proceedings</a> between Elon and OpenAI:</p><blockquote><p>OpenAI&#8217;s counsel asked Musk whether xAI has ever &#8220;distilled&#8221; technology from OpenAI.</p><p>Musk: &#8220;Generally AI companies distill other AI companies.&#8221;</p><p>&#8220;Is that a yes?&#8221; Savitt asked.</p><p>Musk: &#8220;Partly.&#8221;</p></blockquote><p>xAI is likely the largest, and most successful AI company willing to thread the grey area that is distillation from their competitors. On the other side, the majority of startups and research groups with fewer resources than them have very likely engaged in distillation of some capacity from Claude, GPT, or Gemini models.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In the above Anthropic blog post, the problem with the distillation attacks by a few Chinese labs is less the distillation and more the means of attack. It is documented that Chinese labs are actively working to get around the intended use of the API, e.g. to provide additional reasoning data that is very useful for training.</p><p>Of course no one should be able to access information from a model that a developer didn&#8217;t intend to reveal in their APIs (e.g., reasoning traces which would be helpful for training). Associating all of distillation with these attacks, which is to date an industry standard for post-training, from open and closed models alike will be a massive own goal.</p><p>What these few labs are doing should be referred to as jailbreaking or abuse, rather than distillation.</p><p>The discourse around these actions is creating a troubling discussion that&#8217;s marching towards a mix of regulatory capture or regulatory exuberance that&#8217;s most likely to harm the U.S.&#8217;s ecosystem more than China&#8217;s. Even if we ban, most likely through potential legal action and other penalties, this type of API abuse, the Chinese companies will likely still do it. We&#8217;ve seen this playbook with Chinese multimedia models taking a flexible view of copyrighted content that no U.S. player is willing to take the risk on.</p><p>This distillation discussion has quickly snowballed, with a <a href="https://www.congress.gov/bill/119th-congress/house-bill/8283/text">bill moving out of a committee in Congress</a>, an <a href="https://whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf">executive order</a> pushing for action, and <a href="https://www.semafor.com/article/04/29/2026/house-committee-probes-cursor-parent-airbnb-over-chinese-ai">congressional oversight</a> targeting U.S. companies building on Chinese models (which are downstream of distillation). This multi-pronged regulatory environment could yield truly horrible outcomes &#8211; such as figuring out a way to effectively ban open-weight models in the U.S. that are built in China by groups abusing closed LLM APIs.</p><p>It is obvious that no bill will literally ban open models, but they can create grey area that exposes entities to unwanted risk or require certain provisions that are bureaucratically very challenging to fulfill, squashing small open source contributors.</p><p>In that scenario, the groups who lose are Western academics and smaller companies building models for the long-tail of AI uses. The ecosystem here could be made permanently irrelevant with the removal of nearly all Chinese open-weight models. There is no immediate substitute and building new models with meaningful community adoption has a lead time measured in 6+ months. In the time it takes to build a new domestic open-source ecosystem, countless researchers would&#8217;ve moved onto closed training platforms or into new areas.</p><p>Altogether, I&#8217;m hoping this flurry of discussion around distillation becomes a nothing-burger and not a hasty, multi-pronged policy push. We need to avoid two things:</p><ol><li><p>A wholesale negative connotation of the word distillation, which is used extensively across the AI ecosystem.</p></li><li><p>A domestic ban of the open-weight models built by organizations engaged in some portion of distillation.</p></li></ol><p>In addition to this, I want the leading U.S. AI companies to be able to provide their APIs without having their IP leak. They should share more information on why it is hard for them to secure their APIs, but that&#8217;s an issue out of scope for my expertise.</p><p>I&#8217;ll conclude with a proposal from my friend Kevin Xu at <a href="https://www.interconnectedcapital.com/">Interconnected Capital</a> (and great <a href="https://interconnect.substack.com/">Substack</a>) on why this current distillation dynamic may actually be good for the leading labs.</p><p>If all the Chinese companies are addicted to distillation as a way of getting close to the frontier, then they&#8217;ll never actually learn the techniques needed to take an outright lead. If we cut off the Chinese&#8217;s obvious crutch in model building, we&#8217;ll gain a short-term lead in AI, but in the long-term that may be what they needed to get on a more competitive long-term trajectory. </p><p>This is the same debate we&#8217;re having with other technologies where the U.S. currently has a lead, e.g. with advanced semiconductor technologies. So I understand the trade-offs, but we not should crack down on all of distillation.</p>]]></content:encoded></item><item><title><![CDATA[Reading today's open-closed performance gap]]></title><description><![CDATA[The complex factors that determine the single evaluation number so many focus on. Plus, how this changes in the future.]]></description><link>https://www.interconnects.ai/p/reading-todays-open-closed-performance</link><guid isPermaLink="false">https://www.interconnects.ai/p/reading-todays-open-closed-performance</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Mon, 20 Apr 2026 18:25:02 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/707c69ea-fa59-4eba-8034-25b0af9b5443_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>It&#8217;s a clear, current equilibrium that open models will be in <a href="https://www.interconnects.ai/p/open-models-in-perpetual-catch-up">perpetual catch-up of closed models</a>, but this gap being viewed as a single number, a &#8220;distance&#8221;, covers up a nuanced and crucial dynamic at what capabilities the models are covering. The most popular benchmark to comment on this gap is the <a href="https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index">Artificial Analysis Intelligence Index</a> &#8212; a composite benchmark of ~10 sub-evals that they maintain over time to capture the &#8220;frontier&#8221; of current language model capabilities. </p><p>Particularly, I spend a lot of time understanding how dynamics that <em>feed into</em> that index are misunderstood by the natural tendency to reduce performance and trends to one number. Examples include:</p><ul><li><p>How benchmarks evolve over time, becoming more or less correlated with how people actually use models,</p></li><li><p>How different models&#8217; real-world performance relates to their benchmark rankings, and</p></li><li><p>How training regimes evolve over time to move said benchmarks.</p></li></ul><p>Agentic benchmarks are in a decent place, but benchmarks are no longer as trusted as a correlate to real-world performance. A key example to this gray area is Gemini 3&#8217;s incredible benchmarks and remarkable irrelevance in where AI tools currently are being tested and deployed (agents). These trends point to obvious and lasting flaws in our measurements.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/reading-todays-open-closed-performance?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/reading-todays-open-closed-performance?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>At the root of this dynamic &#8212; the dance of correlating model real-world performance and benchmark scores &#8212; is the constant shift of the industry. As all the models, i.e. both open and closed, evolve over time, the topics of focus for benchmarking shifts about every 12 to 18 months. All of the domains of interest have very different training domains associated with them, especially in post-training. The longer a single paradigm goes on, the better the industry gets at measuring performance. In a new era of rapid post-training improvements, I&#8217;m at a relative minimum in my personal confidence in benchmarks.</p><h3>Task evolution and LLM paradigms</h3><p>Right after ChatGPT the focus was a mix of chat, math, and simple code. Instruction tuning and RLHF dominated. Chat capabilities saturated and faded quickly, then mathematics became less focal. Through 2025 and to today, especially once reasoning models became the default, the focus shifted to more complex coding and other simpler agentic tasks. We&#8217;re at the tail end of this first era. Recent training recipes are all dominated by reinforcement learning with verifiable rewards (RLVR), but the domains it is applied in have shifted dramatically from basic question-answer checking to complex environments.</p><p>What we&#8217;re seeing is that the closed, frontier labs are investing astounding sums of money in mastering these current foci &#8212; i.e. code, terminal tasks, etc. &#8212; while starting to push into more diverse knowledge work tasks. These newer tasks encompass specialized domains, such as accounting, law, healthcare, etc. They are still agentic, but require more expertise and often integrations with existing software or domain-specific tools.</p><p>We have very limited evidence on the true balance of capabilities of these newer domains, but these are the areas I&#8217;m focusing on when I say open models will struggle to keep up. The problem is that evaluating <em>complex</em> language model workflows is also a challenging research problem in itself. </p><p>The tasks are getting harder and the data needed to hillclimb on them is getting more private (relative to code, which has swaths of code on GitHub). Leading open model labs are helped by dynamics happening in the data industry that are economically similar to building chip fabs. The few, leading labs in the U.S. pay astronomical sums to buy new environments and datasets, then the fast-following labs (often in China), buy these later at a steep discount. </p><p>This is a key missed point &#8212; that the levers non-frontier labs pull to keep up constantly shift over time. A focus on distillation as the key lever to Chinese models&#8217; progress reflects a blind-spot to the importance of RL environments to current training regimes. If an environment can be built either as a single evaluation in the Artificial Analysis Index, or to mirror it, currently the Chinese labs will be able to keep up. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. Consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>Economic pressure to reinvent &#8220;the frontier&#8221;</h3><p>The question worth dwelling on is: How crucial is the current set of tasks (again, coding and terminal tasks), where the likes of OpenAI and Anthropic have a massive business-adoption advantage over leading open weight models (and even Google alike), is crucial to maintaining revenue numbers? In order to maintain these record growth numbers and trajectories, there needs to keep being a meaningful edge in performance. Many companies would love to reduce their token expenditure cost if they can swap in a far cheaper, open model equivalent. </p><p>If agentic coding abilities saturate and the &#8220;frontier&#8221; of AI performance moves elsewhere, a large amount of the enterprise revenue could be reliant on well-formed customer relationships, inertia, and better product development, rather than the models being leaps and bounds better.</p><p>This precarious position is what I describe as the frontier labs needing to constantly reinvent themselves, and the field&#8217;s prospects, for monetizing the vast buildout of AI infrastructure. I still tend to fall on the side that the buildout will be worth it, and Anthropic and OpenAI will be astronomically profitable businesses, so I take this as a faith of a mix of them continuing to unlock compelling, new, valuable use-cases for the models, and that the benchmarks the open models are closing in on as <em>not being a complete signal</em>. </p><p>I operate with a sort of presumption where the leading open models from China are focused <em>slightly</em> more on benchmarks than the leading closed labs in the U.S. They&#8217;re incentivized to do so &#8212; they want to present the image as constantly being on the heels of the best closed models. Saying the Chinese labs are only in this narrative because they&#8217;re overfitting to benchmarks would be incredibly naive and incorrect. They&#8217;re genuinely strong models, and these dynamics of overselling and real innovation are a fine balance.</p><p>There are a few out-of-distribution benchmarks where open-weight models are very far behind, such as <a href="https://htihle.github.io/weirdml.html">WeirdML</a> or <a href="https://epoch.ai/benchmarks/arc-agi-2/">ARC AGI 2</a>, but there are countless random benchmarks that show these open models as being unexpectedly strong. When you use the models, you can pick up on this lack of robustness (e.g. in long-context capabilities, and needing to reset your agent context more often than Claude/Codex), but they&#8217;re not a category error in the sense that they&#8217;re fundamentally different classes of models. They&#8217;re far closer than many would&#8217;ve expected.</p><h3>How long can open models keep up?</h3>
      <p>
          <a href="https://www.interconnects.ai/p/reading-todays-open-closed-performance">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[My bets on open models, mid-2026]]></title><description><![CDATA[What I expect to come next and why, focused on the open-closed gap.]]></description><link>https://www.interconnects.ai/p/my-bets-on-open-models-mid-2026</link><guid isPermaLink="false">https://www.interconnects.ai/p/my-bets-on-open-models-mid-2026</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Wed, 15 Apr 2026 18:20:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8be08b29-d70a-43f3-8422-6b952816ddab_3182x1790.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We&#8217;re living through the period of time when we&#8217;ll learn if open models can keep up with closed labs. The obvious answer is that no, they won&#8217;t. This answer is a form of saying they won&#8217;t keep up in <em>every area</em>. This framing closes off a popular prediction where the open models completely <em>catch up</em>, as in all models saturate and open and closed models only become increasingly similar. In living through this, it&#8217;s evidently very unclear when the longer-term stable balance of capabilities will solidify. </p><p>This is a very complex dynamic, where the core point we monitor is a <a href="https://www.interconnects.ai/p/open-models-in-perpetual-catch-up">capability gap between models</a>. At the same time, this gap is intertwined with evolving dynamics in the funding of open models, who builds open models, how techniques like distillation that enable fast-following translate through new application domains, potential regulation hampering the open-source AI ecosystem, and of course who actually uses open models. </p><p>The capabilities gap is one signal in a complex sea of forces, pushing supply and demand into different shapes. In many cases the demand &#8212; where obviously tons of individuals, organizations, and sovereigns want, or need, open models &#8212; is largely separated from supply. Supply is fully dictated by economics. The question of &#8220;which business strategies support releasing open models&#8221; is still at stake.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Interconnects AI is a reader-supported publication. To receive new posts and support my work, consider becoming a subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>With this complexity, I wanted to distill my key beliefs down into a clear list. These are downstream of 10+ pieces I&#8217;ve written or recorded on open models this spring (which are linked throughout).</p><ol><li><p><strong>It&#8217;s surprising that the top closed models did </strong><em><strong>not</strong></em><strong> show a growing capability margin over open models</strong>, based on compute differences for training and research, especially in the second half of 2025 and through today.</p></li><li><p><strong>Open model labs are technically very strong</strong> at keeping pace on well-established benchmarks. This will continue and reflects a balance of abundant talent and sufficient computing power. </p></li><li><p><strong>Chinese open-weight labs focus </strong><em><strong>slightly</strong></em><strong> more on benchmark scores</strong> than comparable closed labs in the U.S. <a href="https://www.interconnects.ai/p/how-much-does-distillation-really">Distillation</a> helps the Chinese LLM companies do so, but it&#8217;s not a panacea. Changes in the distillation dynamic (e.g. regulation) will not be a determining factor on the balance of capabilities. This increase in focus is a natural evolution of their incentives in keeping the narrative on keeping up with the frontier alive, which is crucial to fundraising and adoption.</p></li><li><p><strong>To date, closed models tend to be more robust and generally useful than similarly scoring open models</strong>. Closed models have certain hard-to-measure qualities that are not well captured in current or past benchmarks. This will be key to enabling closed models to dominate in markets where an individual user constantly presents new challenges, i.e. supporting knowledge workers as a direct assistant.</p></li><li><p><strong>The open vs. closed model race, as monitored through benchmarks, will largely be a game of economic staying power</strong> and fast-following, until the market structure constricts. I expect Chinese open-weight labs to face funding difficulties first, as soon as later this year. Funding difficulties will be seen in different capability trajectories 3-9 months later.</p></li><li><p><strong>The RL dominated training era has increased the relevance of distribution to real-world use-cases as a key factor in continued capabilities improvements</strong>. These are tasks where users directly use tools like Claude Code or Codex to solve problems in their job with agents. This is the first clear technical area that closed labs can dominate open-weight models on capabilities, potentially <a href="https://cursor.com/blog/real-time-rl-for-composer">leveraging online RL directly</a> based on user feedback.</p></li><li><p><strong>Open models will be increasingly adopted in repetitive automation tasks</strong>, as measured in the relative share of the API market, for repetitive tasks across the ecosystem. This takes the form of many new AI-native applications, business backend automation, etc. The success of this will <a href="https://www.interconnects.ai/p/the-next-phase-of-open-models">drive more investment in domain-specific, efficient open models</a>.</p></li></ol><p>This is a complex picture, where the long-term trajectory is more of an economics question rather than an ability one. Many other outlets can paint a far more simplistic narrative that &#8220;<a href="https://www.nytimes.com/2026/04/13/opinion/china-ai-america-chipmakers.html">China will assuredly catch us in AI</a>&#8221; and get more distribution because it is a simple story. The reality is complex. Only real AI revenue begets more investment, eventually that&#8217;ll be linked to the ability to keep improving models at a rapid rate. Economic realities have not yet impacted scaling open models, as a general category.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/my-bets-on-open-models-mid-2026?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/my-bets-on-open-models-mid-2026?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>This economic-focused angle relates to my positions on the open model ecosystem more broadly.</p><ol start="8"><li><p><strong>Recurring calls to ban certain types of open models will continue to come but are in practice impossible to implement.</strong> Training strong AI models (i.e. near but not at the frontier) is a relatively small cost compared to large-scale deployments. E.g. if the U.S. bans open models over a certain compute threshold, another sovereign entity will eventually train them and release them publicly, with the models entering the U.S. market with less oversight.</p></li><li><p><strong>The second derivative of influence on open models has shifted, and the U.S. will slowly regain ground in <a href="https://www.interconnects.ai/p/8-plots-that-explain-the-state-of">adoption metrics</a></strong> of open models starting in early 2027 (it takes a long time for China&#8217;s velocity to slow, then flip). Examples include Google&#8217;s <a href="https://www.interconnects.ai/p/gemma-4-and-what-makes-an-open-model">Gemma 4</a> (a wild success), <a href="https://www.interconnects.ai/p/why-nvidia-builds-open-models-with">Nvidia&#8217;s Nemotron</a>, and <a href="https://www.interconnects.ai/p/arcee-ai-goes-all-in-on-open-models">Arcee AI</a>.</p></li><li><p>As ever-stronger closed models are built, previewed, and released, there will be more <strong>safety-shocks saying that open-weight versions of the strongest AI models never can be allowed to exist</strong>, similar to reactions to <a href="https://www.interconnects.ai/p/claude-mythos-and-misguided-open">Claude Mythos</a>. These can spur burdensome regulation on open models.</p></li><li><p>With the above, there will also be <strong>increased long-term interest in open models</strong>, as sovereign entities and existing power structures realize the coming, super powerful AI tools<a href="https://www.interconnects.ai/p/how-anthropic-vs-dow-impacts-open"> cannot land in the hands of only one or a few companies</a>. These entities will see open models as a different governance paradigm.</p></li><li><p><strong>New funding structures for open models will emerge</strong>, as many stakeholders realize <a href="https://www.interconnects.ai/p/the-inevitable-need-for-an-open-model">dependencies on single, for-profit companies for access to intelligence are unreliable</a>.</p></li><li><p><strong>Local agents, OpenClaw, and other personal agents represent a large, to date, mostly ignored market for open model usage</strong>. It is a sort of dark matter, with pervasive, massive potential for influence on the balance of open-to-closed models.</p></li></ol><p>A single word governs this post and is intentionally repeated &#8212; complex.</p><p>This complex reality has been driving me to think more deeply about how to clearly describe the open model gap, and why I can hold it in my head that I expect American closed labs to clearly draw ahead, despite the fairly unequivocal evidence in support of the capabilities of recent open-weight models. More on the nuance in the open-closed gap in another piece coming soon, so <a href="https://www.interconnects.ai/subscribe">please subscribe</a>!</p><p>Let me know any positions that I missed.</p>]]></content:encoded></item><item><title><![CDATA[What I’ve been building: ATOM Report, post-training course, finishing my book, and ongoing research]]></title><description><![CDATA[What I've been up to!]]></description><link>https://www.interconnects.ai/p/what-ive-been-building-atom-report</link><guid isPermaLink="false">https://www.interconnects.ai/p/what-ive-been-building-atom-report</guid><dc:creator><![CDATA[Nathan Lambert]]></dc:creator><pubDate>Tue, 14 Apr 2026 20:41:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Bv0Q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This post is a roundup of my recent efforts that did not warrant a standalone Interconnects post, why I&#8217;m spending time on them, and what they accomplished.</p><ol><li><p><a href="https://www.interconnects.ai/i/194224428/1-the-atom-report-measuring-the-open-language-model-ecosystem">The ATOM Report: Measuring the Open Language Model Ecosystem</a></p></li><li><p><a href="https://www.interconnects.ai/i/194224428/2-rlhf-book-is-done-and-ready-for-pre-order">RLHF Book is done &amp; ready for pre-order!</a></p></li><li><p><a href="https://www.interconnects.ai/i/194224428/3-a-post-training-course-im-making">A post-training course I&#8217;m making</a></p></li><li><p><a href="https://www.interconnects.ai/i/194224428/4-recent-technical-research">Recent technical research</a></p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/what-ive-been-building-atom-report?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/what-ive-been-building-atom-report?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2>1. The ATOM Report: Measuring the Open Language Model Ecosystem</h2><p><a href="https://arxiv.org/abs/2604.07190">https://arxiv.org/abs/2604.07190</a></p><p>To accompany The ATOM Project <a href="https://atomproject.ai/">memo</a>, arguably a manifesto, making the case for investment in open models in the U.S. &#8211; originally launched in August 2025 &#8211; we&#8217;ve released an updated technical report with our latest data, analysis, and storytelling within the open language model ecosystem. The ATOM Report is dense with the methods Florian and I use to keep track of the open ecosystem. It covers GPT-OSS&#8217;s rise, inference market share, the influence of China&#8217;s mid-tier players like Moonshot, Z.ai, &amp; MiniMax, signs of the U.S.&#8217;s progress on open models, and much more.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JZNn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JZNn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png 424w, https://substackcdn.com/image/fetch/$s_!JZNn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png 848w, https://substackcdn.com/image/fetch/$s_!JZNn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png 1272w, https://substackcdn.com/image/fetch/$s_!JZNn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JZNn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png" width="1456" height="1123" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1123,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:330823,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.interconnects.ai/i/194224428?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JZNn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png 424w, https://substackcdn.com/image/fetch/$s_!JZNn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png 848w, https://substackcdn.com/image/fetch/$s_!JZNn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png 1272w, https://substackcdn.com/image/fetch/$s_!JZNn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89b0ff17-1243-46dd-a81e-96c975f20a7b_2582x1992.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In particular, the paper details our updates to the <a href="https://atomproject.ai/relative-adoption-metric">Relative Adoption Metric (RAM)</a>, which we use to evaluate the adoption of recent models in a time-varying and size-normalized manner. Here&#8217;s a sampling of recent, primarily Chinese, models on the RAM score. The RAM score is designed so that a score &gt;1 indicates a model is, at that point in time, on track to be a top 10 most downloaded model of its size category, ever. It reduces a messy landscape to one, easily interpretable number!</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TeBR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TeBR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png 424w, https://substackcdn.com/image/fetch/$s_!TeBR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png 848w, https://substackcdn.com/image/fetch/$s_!TeBR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png 1272w, https://substackcdn.com/image/fetch/$s_!TeBR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TeBR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png" width="1456" height="1014" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1014,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:323626,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.interconnects.ai/i/194224428?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TeBR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png 424w, https://substackcdn.com/image/fetch/$s_!TeBR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png 848w, https://substackcdn.com/image/fetch/$s_!TeBR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png 1272w, https://substackcdn.com/image/fetch/$s_!TeBR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ef64b7a-04f2-4ed8-9cc4-966b775e9f59_1918x1336.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We used the data to also analyze the recent <a href="https://www.interconnects.ai/p/gemma-4-and-what-makes-an-open-model">Gemma 4</a> release, which is showing incredible early adoption numbers. We&#8217;ll stay tuned on it!</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!u86h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!u86h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg 424w, https://substackcdn.com/image/fetch/$s_!u86h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg 848w, https://substackcdn.com/image/fetch/$s_!u86h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!u86h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!u86h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!u86h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg 424w, https://substackcdn.com/image/fetch/$s_!u86h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg 848w, https://substackcdn.com/image/fetch/$s_!u86h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!u86h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9e2abfe-443e-48e8-bfda-bb5855dee388_1936x1056.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Subscribe to the (infrequent) <a href="https://atomproject.substack.com/">ATOM Project Substack</a> for more updates like this!</p><h2>2. RLHF Book is done &amp; ready for pre-order!</h2><p><a href="http://rlhfbook.com/">http://rlhfbook.com/</a></p><p>The goal of this book was to write the book I wished I had when I was getting started in post-training language models. This project has been on my mind for a long time. I bought the domain rlhfbook.com and started to take it more seriously on May 20th, 2024. Here we are!</p><p>Last week, it was sent to production with the Manning team. This means content edits are done, and it&#8217;ll be sent to print in ~2 months. In the meantime, I&#8217;m spending my time developing the accompanying code and course (more on that below).</p><p>You can preorder on <a href="https://amzn.to/4cwCDJQ">Amazon</a> or <a href="https://www.manning.com/books/the-rlhf-book">Manning</a> (currently cheaper).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Bv0Q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Bv0Q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Bv0Q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Bv0Q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Bv0Q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Bv0Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!Bv0Q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Bv0Q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Bv0Q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Bv0Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2d8ba64-922d-4000-9d57-12cb5524a238_1200x675.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>3. A post-training course I&#8217;m making</h2><p><a href="https://rlhfbook.com/course">https://rlhfbook.com/course</a></p><p>The goal of my book is for it to be the central resource for people looking to transition from beginner to expert in post-training. It&#8217;s not necessarily an entry-level book, but as AI models become stronger, it needs to be a <em>community</em>-building effort as well. The first step I&#8217;ve made to expand the scope from just a book to a complete learning experience is building a lecture series. The lectures will be freely available on YouTube and incorporate community questions &amp; answers (as standalone videos in between lectures).</p><p>You can watch the first batch of videos below, and subscribe on YouTube for future ones. I&#8217;m going to build on the book platform more this summer, as I develop the book <a href="https://rlhfbook.com/code">codebases</a> and host in-person events.</p><ul><li><p><a href="https://www.youtube.com/watch?v=jQPiH-KB4B0&amp;list=PLL1tdVxB1CpVpEtMHxwuR4uI4Lxjw00_y&amp;index=3">Welcome video &amp; YouTube playlist</a></p></li><li><p><a href="https://youtu.be/o6l6tJQgUg4">RLHF and Post-training Overview | RLHF Book Course, Lecture 1</a></p></li><li><p><a href="https://youtu.be/4gIwiSPmQkU">RLHF Foundations, IFT, Reward Modeling, Rejection Sampling | RLHF Course Lecture 2</a></p></li><li><p><a href="https://youtu.be/K_Sj_-1BUMM">Understanding Policy Gradient Algorithms for RL on LLMs | RLHF Course Lecture 3</a></p></li><li><p><a href="https://youtu.be/i-AIMpZHgeg">Implementing RL Algorithms for LLMs | RLHF Course Lecture 4</a></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VS0r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VS0r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png 424w, https://substackcdn.com/image/fetch/$s_!VS0r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png 848w, https://substackcdn.com/image/fetch/$s_!VS0r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png 1272w, https://substackcdn.com/image/fetch/$s_!VS0r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VS0r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1243661,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.interconnects.ai/i/194224428?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VS0r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png 424w, https://substackcdn.com/image/fetch/$s_!VS0r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png 848w, https://substackcdn.com/image/fetch/$s_!VS0r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png 1272w, https://substackcdn.com/image/fetch/$s_!VS0r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb238c68c-d7f4-4b2b-97fa-9a3cf773e72b_1280x720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>4. Recent technical research</h2><p>Long-time followers of Interconnects know that this blog has its roots in explaining fundamental research in the field. This has immense value in two ways. First, as AI moves incredibly fast, far more people need to be able to parse research to make the right bets on the technology. Research is the only early warning of some big changes coming. Second, it helps uplift the careers of my collaborators &#8211; the people I spend my life with! On that note, check out two papers I had the privilege of being part of below.</p><p><a href="https://arxiv.org/abs/2603.16759">https://arxiv.org/abs/2603.16759</a> -<em> TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities</em>,<em> </em>Graf et al. 2026</p><p>This work explores the strengths of various models in multi-turn dialogue settings, how to create training data to improve it, and other quirks in post-training. My interests here have fully shifted to agents, where I see multi-turn interactions as a very important user interface problem &#8212; what information do I show to the user to solve the task as soon as possible without cutting corners?</p><p><a href="https://arxiv.org/abs/2603.11327">https://arxiv.org/abs/2603.11327</a> - <em>Meta-Reinforcement Learning with Self-Reflection for Agentic Search</em>, Xiao et al. 2026</p><p>This paper frames solving hard problems with RLVR as a meta-learning problem, where context from previous attempts should be used to inform future rollouts. It&#8217;s a very obvious idea in some ways, where most of RL for LLMs is still very on-policy, but naive. The models learn from recent trials in parameters, but not in context. This research feeds into a ton of other recent work on ways that RL can be formulated to solve different forms of continual learning. Another great related paper is <em><a href="https://arxiv.org/abs/2601.16175">Learning to Discover at Test Time</a>.</em></p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.interconnects.ai/p/what-ive-been-building-atom-report/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.interconnects.ai/p/what-ive-been-building-atom-report/comments"><span>Leave a comment</span></a></p><p>I&#8217;m off to China (and then hopefully DC) in the next couple of months to learn even more about how the world sees progress in AI. I&#8217;m excited to talk to a broader range of people than I tend to in my focused technical job. Thanks for reading, as always!</p>]]></content:encoded></item></channel></rss>