<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The AI Realist]]></title><description><![CDATA[AI without the hype: what it costs, who profits, and what the filings say.]]></description><link>https://www.airealist.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!u6cR!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924ecf6b-2ddb-4f24-a3bd-89ae62c7c1dc_800x800.png</url><title>The AI Realist</title><link>https://www.airealist.ai</link></image><generator>Substack</generator><lastBuildDate>Thu, 08 Oct 2026 13:39:15 GMT</lastBuildDate><atom:link href="https://www.airealist.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Julien Simon]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[julsimon@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[julsimon@substack.com]]></itunes:email><itunes:name><![CDATA[Julien Simon]]></itunes:name></itunes:owner><itunes:author><![CDATA[Julien Simon]]></itunes:author><googleplay:owner><![CDATA[julsimon@substack.com]]></googleplay:owner><googleplay:email><![CDATA[julsimon@substack.com]]></googleplay:email><googleplay:author><![CDATA[Julien Simon]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Kolibri Is Large Enough]]></title><description><![CDATA[American chips, Chinese teachers, about $3 million for the training run. What Germany's new open model adds is eligibility.]]></description><link>https://www.airealist.ai/p/kolibri-is-large-enough</link><guid isPermaLink="false">https://www.airealist.ai/p/kolibri-is-large-enough</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Tue, 06 Oct 2026 06:28:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dJAi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dJAi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dJAi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dJAi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dJAi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dJAi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dJAi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:658258,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/219035310?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dJAi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dJAi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dJAi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dJAi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b8b5603-92ac-40c6-a8e8-e1f8830c1157_1456x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On October 3rd, 2026, the Day of German Unity, Aleph Alpha put a model called Kolibri on Hugging Face: 78 billion parameters, German and English only, trained from scratch, free under Apache 2.0.[1] Two years earlier, on September 5th, 2024, its then-CEO, Jonas Andrulis, told Bloomberg why building one was no longer enough. As TechCrunch relayed it the same day: &#8220;Just having a European LLM is not sufficient as a business model. It doesn&#8217;t justify the investment.&#8221;[2] And seventeen days before Kolibri shipped, the company had signed a definitive agreement, not yet closed, to combine with Cohere, of Toronto.[3]</p><p style="text-align: justify;">So a company that said a European model alone doesn&#8217;t pay, and that has agreed to operate under another company&#8217;s name, has just built one. Why? Look at what it built. The United States, China, and Europe are all inside it, and it shows how each competes in AI in 2026.</p><p style="text-align: justify;">Here is what I think: the investment that a European LLM &#8220;doesn&#8217;t justify&#8221; has collapsed. I&#8217;m not going to tell you Kolibri beats the frontier models. It doesn&#8217;t claim to. It is large enough for the job, and that is the point.</p><p style="text-align: justify;">If you run technology or allocate capital in Europe, read Kolibri as a bill of materials: what it cost, who supplied each part, and what each supplier can take back. Its own documents let you start counting. Before I do: I was Chief Evangelist at Arcee AI (see below) until November 2025.</p><h2><strong>What does a from-scratch model cost in 2026?</strong></h2><p style="text-align: justify;">Aleph Alpha&#8217;s model card gives the bill in GPU-hours: 392,000 for pre-training on 768 NVIDIA B200s over 21 days, 90,000 for a second phase on reasoning and code data, 10,000 to stretch the context window. That is 492,000 GPU-hours.[4]</p><p style="text-align: justify;">The technical report thanks its cloud provider once: Verda, which publishes its prices. A B200 rents for $7.13 an hour on demand, and 25% less on a two-year commitment. Multiply it out, and you get $3.5 million at the first rate and $2.6 million at the second.[5] A cluster this size is quoted privately, and Aleph Alpha hasn&#8217;t published what it paid.</p><p style="text-align: justify;">That number leaves things out: an earlier model whose pre-training was restarted once, the small experiments before the real run, the fine-tuning, and every GPU-hour spent generating training data with other people&#8217;s models. So the bill starts at roughly $3 million, and almost none of the rest is published.[6] Verda gave the only outside hint on October 5th, 2026: &#8220;Training a model like Kolibri requires over a thousand GPUs running around the clock for months.&#8221; A thousand GPUs for three months would cost about $15 million (my calc; the three is my assumption).</p><p style="text-align: justify;">In October 2021, <a href="https://huggingface.co/blog/large-language-models">I wrote</a> that replicating a 530-billion-parameter model would cost &#8220;close to $100 million&#8221; in servers, networking and hosting, and asked which organizations had a use case that justified even $10 million. Very few, I said. Five years later, you can rent the training run of a serious model for less than that $10 million.</p><h2><strong>The serving bill chose the size</strong></h2><p style="text-align: justify;">Why 78 billion parameters, and not 400? And why only German and English? Because Aleph Alpha builds for ministries and manufacturers who want to run this on their own machines, and the company sized the model based on serving cost, not a leaderboard.</p><p style="text-align: justify;">Start with what the 78 billion buys. Kolibri is a mixture-of-experts model, so only 3.46 billion of the 78 billion parameters do any work for any one token. That splits the serving bill in two: you pay for speed on 3.46 billion, and for memory on all 78 billion.[7] So the question of size is a question of memory.</p><p style="text-align: justify;">The launch post is plain about it: bigger was better and more expensive, and &#8220;The latter drove the decision&#8221;. Then the team estimated what each size would cost to serve: &#8220;123B can handle only 3 long-context 256k-token user queries on two H100s, while 78B handles 18 concurrent requests&#8221;.[8] Aleph Alpha's own estimate says six times as many long-document users on the same two GPUs won the argument.</p><p style="text-align: justify;">The two languages also follow from the budget. The card calls it &#8220;a deliberate choice of depth over breadth&#8221;, and the report prices a language: about a fifth of the pre-training tokens, i.e., 4 trillion for German, when the open German datasets held roughly half that. I&#8217;d expect French or Spanish to cost that again each. The report lists other European languages as future work.[7]</p><h2><strong>How do you tune a model you can only train once?</strong></h2><p style="text-align: justify;">A full run costs millions, so choices are tested on small copies first. Aleph Alpha used two network shapes: one with 0.6 billion active parameters trained on 97.6 billion tokens, and one with 2.1 billion trained on 327.6 billion. A third, half as wide as Kolibri, tuned the training settings. The hard part is carrying what you learn up to full size.</p><p style="text-align: justify;">For its earlier model, the team took the common approach: fit scaling laws, i.e., formulas that predict the best training settings from small runs, on runs of 100 to 430 billion tokens, then extend them to 7.5 trillion. Reused for Kolibri, those laws gave &#8220;unrealistic hyperparameter predictions&#8221;. That is a useful admission. A scaling law is not a law of nature. It is a curve fitted to one model, one context length, and one token budget, and this one stopped working when all three changed.</p><p style="text-align: justify;">So the team dropped them and extrapolated each candidate setting's loss curve on the half-width model. The result still has to hold at full width. To do that, the team used a known technique called &#181;P: it scales the starting weights and learning rates with the network width, so settings tuned on a narrow model also work on a wide one.</p><p style="text-align: justify;">The size came out of the same sweep I quoted above. At 3.4 billion active parameters, the team tried 42, 82, and 123 billion in total; the 82 is the test configuration closest to the 78 that shipped. Going from 42 to 82 improved loss and accuracy. Going to 123 improved them &#8220;only slightly further&#8221;, and at a 4,000-token context it decoded 32% slower than 42, where 82 was 13% slower.</p><p style="text-align: justify;">That is all &#8220;Pareto frontier&#8221; means in the report&#8217;s title: no other model it tested is both better and cheaper to run. On one of the report&#8217;s charts, three models sit on that frontier: &#8220;Qwen3.5 35B-A3B, Kolibri and Qwen3.8 27B&#8221;. Two of the three are Alibaba&#8217;s, and one of those is Kolibri&#8217;s teacher. The cost on that chart is a count of active parameters, not a measurement. The chart behind the title plots text decoded per second per GPU, as Aleph Alpha measured it. Memory, which chose the size, is on neither.</p><p style="text-align: justify;">The last stage of training is reinforcement learning against verifiable rewards: the model attempts a task and gets a grade from a program that can check the answer and from a frozen judge model for proofs and conversations. Kolibri ran one training run of 1,000 steps on 256 B300 GPUs across 38 environments. They are general: software engineering and terminal use make up 36% of the mix, math and reasoning 18%, and seven are German variants. No training environment is named for an industry. The industries appear on the test side: five in-house stand-ins for assistants Aleph Alpha runs for customers, in semiconductors, the German public sector, aerospace, automotive supply, and industrial drives. Two of the five have 11 and 26 questions.</p><p style="text-align: justify;">Generating the attempts &#8220;can dominate the cost of RL&#8221;, so the model that writes them runs with 8-bit weights and an 8-bit cache, which the report says is two to three times faster. The trainer copies that rounding, so that both sides stay the same model. In April, <a href="https://www.airealist.ai/p/the-verification-tax">I called the checking step the verification tax</a>: it runs on CPUs, and it can limit how fast a model learns. The run I described then, at Together AI, used more than 100 sandboxes at once on 32 GPUs. Kolibri&#8217;s averaged 25,000, against 480 for Aleph Alpha&#8217;s earlier model. They ran on the CPUs the GPU servers &#8220;keep idle&#8221;: 112 cores per node across 31 nodes, about 3,500 total. They paid the tax with cores that came with the rented GPUs.</p><p style="text-align: justify;">The report doesn't give MFU, model FLOPs utilization: the share of a chip&#8217;s peak arithmetic that training actually uses. It is the key efficiency metric in pre-training. It is also a favorite vanity metric for AI labs. In reinforcement learning, the key metric is simpler: how long the GPUs sit idle, waiting for the CPUs to check the answers. What Aleph Alpha publishes is throughput, a median of about 16,500 tokens per second per GPU in pre-training, and reliability: 38 unplanned interruptions that cost 28.7 hours in total. It publishes no utilization figure for any stage. My back-of-the-envelope estimate from its own figures is 15% to 22% for pre-training, if the run used 16-bit arithmetic: 15% counting the active parameters alone, 22% with attention and the output layer added.[9]</p><h2><strong>A page from Arcee AI&#8217;s book</strong></h2><p style="text-align: justify;">In January 2026, Arcee AI released Trinity Large: a 400-billion-parameter mixture-of-experts model, pre-trained on 2,048 B300s by a company of about thirty people, as reported. Arcee AI&#8217;s own account puts the whole six-month effort, including smaller models, at $20 million all in. The Trinity report came out after I left Arcee AI. Kolibri&#8217;s report cites it by name for two methods it reuses.[10]</p><p style="text-align: justify;">The book is the same: your own pre-training run, mixture-of-experts, open weights, a rented cluster, and a detailed report. Aleph Alpha runs it at a fifth the size, narrows it to two languages, and sells it in its home jurisdiction.</p><p style="text-align: justify;">The two differ in what they give away. Arcee AI published its raw pre-training checkpoints, and named no teacher. Aleph Alpha kept its earlier model, Kolibri Origin, to itself, and named the models that helped build Kolibri. Neither released its full training set, though Aleph Alpha&#8217;s German web dataset is public. Kolibri&#8217;s license adds a sentence under the Apache grant: it &#8220;does not extend to underlying code, model architecture, parameter settings or any training method.&#8221;[11]</p><p style="text-align: justify;">So the weights are free, and the recipe is printed, but the pantry is locked.</p><h2><strong>Who wrote the training data?</strong></h2><p>Like Arcee AI and NVIDIA, Aleph Alpha had other models write part of it.</p><p style="text-align: justify;">About 24% of Kolibri&#8217;s pre-training data, 4.8 trillion tokens, is synthetic. Google&#8217;s Gemma 4 rewrote English web pages into cleaner forms. For German, Mistral&#8217;s NeMo model did the rewriting, and the report calls the result &#8220;the single largest source of German in the model&#8221;.[12]</p><p style="text-align: justify;">Then comes the part that teaches the model to reason, call tools, and hold a conversation. The report: &#8220;The main models we use to generate this data, and to regenerate parts of the open datasets, are GLM-5.2, GLM-5.3 and Qwen3.8-27B.&#8221; The first two come from Z.ai; the third from Alibaba. The summary Aleph Alpha published under the EU&#8217;s AI Act template lists the same models by their Hugging Face repositories, including an 8-bit copy of GLM-5.2. That tells me the company downloaded the weights and ran them on its own machines.[13]</p><p style="text-align: justify;">Do the licenses allow it? I checked instead of assuming. On October 5th, 2026, I read the license terms of eleven of the twelve outside models in the recipe. All but three of them are plain Apache 2.0 or MIT. GLM-5.3 has its own license, which adds a security review for the largest companies selling model access. None restrict training another model on their outputs.[15]</p><p style="text-align: justify;">gpt-oss and Gemma 4 are also Apache 2.0, so the licenses are no different. What differs is which models get released. Here is how I read it: America gives away models good enough to filter and rewrite, and keeps the ones good enough to teach behind an API, while Chinese labs give away the teachers. That may be changing. On October 5th, 2026, Reflection AI, an American lab, announced Beam, a 501-billion-parameter model it calls competitive with GLM-5.2, and promised the weights under Apache 2.0 in October.[15]</p><p style="text-align: justify;">And where did the Chinese teachers learn? On September 8th, 2026, three American agencies published an advisory accusing six Chinese labs, including Z.ai and Alibaba, of distilling US frontier models at industrial scale, which they called &#8220;the core&#8212;not merely a supplement&#8212;of their AI development strategy.&#8221;[16] The advisory names the companies, not the models Kolibri used. If the agencies are right, the chain would run like this: American frontier models, illicitly, into Chinese open weights, and from there, lawfully, into a German model.</p><p style="text-align: justify;">On September 28th, 2026, Aleph Alpha published a study of Chinese political alignment in open models: six Chinese models gave balanced answers to only 17% to 41% of 967 sensitive prompts, compared with 70% for Claude Sonnet 5 and 92% for Mistral Small. Thus, generated conversations with open-ended answers pass through a filter that drops those that don't. The judge in that filter is gpt-oss-120b, from OpenAI.[14] . The study adds: &#8220;Anyone who is building their first frontier model effectively has to fall back on open-weight models and work with the resulting political behavior.&#8221; It says it has adopted its own remedies, and its model card warns that Kolibri &#8220;may reproduce political biases present in its training data&#8221;.[17] But the study tests no model from Z.ai, the lab behind two of Kolibri&#8217;s three main teachers.</p><h2><strong>Three blocs, one model</strong></h2><p style="text-align: justify;">In March, <a href="https://www.airealist.ai/p/build-buy-or-download-someone-elses">I wrote</a> that a country without a frontier lab has three bad choices: build its own model, buy American, or download Chinese weights. In June, <a href="https://www.airealist.ai/p/too-dangerous-for-you-free-for-everyone">I put the three blocs in one line</a>: the United States restricts, Europe regulates, Chinese labs ship.[18] Both need an update, because Kolibri doesn&#8217;t choose. It downloads, rents, and builds.</p><p style="text-align: justify;">American companies supplied every GPU the documents name, and gave away the filter and the rewriter. So far, they kept their teachers. So the United States competes at the top and at the bottom of the stack: rent the silicon to everyone, sell the frontier by the token, and let the rest of the world assemble whatever fits in between.</p><p style="text-align: justify;">Two Chinese labs supplied the main teachers, under licenses that charge nothing. So Chinese labs compete on distribution. Their work ends up inside other people&#8217;s products: in Mistral&#8217;s catalog, and now in a German model built for ministries.[19]</p><p style="text-align: justify;">Europe supplied the cloud, the work of assembling a German corpus, the team, the intended customers and the law. That is the last mile, a consolation prize at the frontier. But it is the market Aleph Alpha says it builds for, and Europe&#8217;s rules do one useful thing here: the reason I can list Kolibri&#8217;s teachers by repository is that the AI Act made Aleph Alpha fill in a form. Give yourself a pat on the back, EU bureaucrats.</p><p style="text-align: justify;">In June, <a href="https://www.airealist.ai/p/too-dangerous-for-you-free-for-everyone">I said</a> my read would break if someone shipped &#8220;a model under 100 billion parameters that matches the leaders.&#8221; Kolibri is under 100 billion, and it doesn&#8217;t claim to match the leaders. So my reading holds: someone still has a hand on every door in this model.[20]</p><h2><strong>Is it any good?</strong></h2><p style="text-align: justify;">On Aleph Alpha&#8217;s own test suite, Kolibri scores 75.5 in English and 70.8 in German. OpenAI&#8217;s gpt-oss-120b, a bigger open model with more active parameters, scores 72.3 and 70.2 on the same suite. The closest model in Kolibri&#8217;s class, Alibaba&#8217;s Qwen3.5 35B, is a point or less behind in both languages; call it a tie, at less than half Kolibri&#8217;s size. And Qwen3.8-27B, one of Kolibri&#8217;s teachers, beats it by 4.7 points in English and 9.1 in German, while activating nearly eight times as many parameters per token.[21]</p><p style="text-align: justify;">So why would a German ministry pick Kolibri over a Qwen half its size? On the table, it wouldn&#8217;t. On Aleph Alpha&#8217;s own German public-sector test, Kolibri scores 75, the small Qwen 80 and the dense Qwen 89. It would pick Kolibri on eligibility. When Hesse <a href="https://www.airealist.ai/p/vent-mauvais">bought from Mistral</a>, in an award published on July 2nd, 2026 and made without a call for competition, one mandatory requirement was that the vendor&#8217;s models be developed in the European Economic Area or a country the EU deems adequate. Qwen fails that line before anybody runs a benchmark. Hesse found only Mistral eligible, and on the same public-sector test Mistral Small 4, the one Mistral model on Aleph Alpha&#8217;s table, scores 50. A German company you can sign with, audit, and sue passes it. A different Qwen of that size, Qwen 3.6 35B, echoed Beijing&#8217;s framing on 80% of sensitive prompts in Aleph Alpha&#8217;s own study, and Kolibri&#8217;s score on that test is the number its buyers are owed. Until it is published, Kolibri is how a ministry with Hesse&#8217;s requirement buys what Chinese models taught.[21]</p><h2><strong>And Mistral?</strong></h2><p style="text-align: justify;">Mistral appears twice in this story. Three times, if you count my title: in July 2024, Mistral announced its Large 2 model under the headline &#8220;Large Enough&#8221;.[22]</p><p style="text-align: justify;">It is in the pre-training data, as NeMo, a now-deprecated 12B multilingual model co-built with NVIDIA. The results table also shows that Mistral Small 4 trails Kolibri by 12.4 points in English and 9.4 in German on Aleph Alpha&#8217;s suite, though it leads Kolibri on one knowledge index in the same report. And it is missing from the list of teachers: of the models Aleph Alpha names for writing reasoning traces, none is European.[22]</p><p style="text-align: justify;">Now read Cohere&#8217;s release. The combined company will operate as Cohere and, the release says, &#8220;will effectively create the first transatlantic sovereign AI solution&#8221;. In it, Ilhan Scheer, Aleph Alpha&#8217;s chief executive, calls Cohere &#8220;a strong strategic fit&#8221; for its &#8220;specialized language models for governments and regulated industries&#8221;. Neither company&#8217;s announcement names Mistral, so what follows is my reading. Cohere is not new to Europe: in June, it wrote of &#8220;our growing momentum across the UK and Europe&#8221;. What Aleph Alpha adds is a model built in Germany, for German, under European law, and Cohere buys Aleph Alpha and that model to go deeper into the European enterprise market. That is precisely where Mistral is trying to make money.[22]</p><h2><strong>Aleph Alpha has tried this before</strong></h2><p style="text-align: justify;">In 2023, some of its investors were also its customers. Bosch and the Schwarz Group led a round that included SAP, and the company&#8217;s own release says it included &#8220;preconsumption licenses&#8221;, i.e., prepaid orders. What is different now is the product. The weights it published in 2024 were licensed only for research and education. Kolibri is Apache 2.0, runs on two GPUs in the buyer&#8217;s own room, and may soon have Cohere&#8217;s sales force behind it.[23]</p><h2><strong>What would have to break</strong></h2><p>Can Europe compete this way? Yes, as long as you are clear about what &#8220;this way&#8221; means.</p><p style="text-align: justify;">It works on one condition: the scores hold up when somebody else measures them. I ran no benchmark of my own. If an independent evaluation puts Kolibri well below its class, the recipe is cheap because it is bad, and I will say so. The same goes for the bill: if the unpublished part turns out to be many times the training run, the $3 million tells you little.</p><p style="text-align: justify;">It also works only as long as two other parties allow it. The Chinese labs own the license file, and a license can change with the next release, as GLM&#8217;s did between versions 5.2 and 5.3. Washington owns the chips, which an export rule could reach, and Axios reported in July that restrictions on Chinese models were under discussion.[24] If the teachers go, the cheap part of the recipe stays cheap, and the data becomes the wall.</p><p style="text-align: justify;">The chips are American. The teachers are Chinese. On its maker&#8217;s own public-sector test, a Chinese model half the size scores higher. Only Europe could supply a legal entity, a form, and a buyer&#8217;s requirement that keeps the better model out of the room. I don&#8217;t call that an AI industry. I call it a customs office.</p><p style="text-align: justify;">Andrulis was right in 2024: a European LLM alone isn't a business model. Kolibri doesn&#8217;t prove him wrong. It shows the business model: eligibility, sold as a service, and under a Toronto company&#8217;s name if the deal closes. The launch post names the real asset itself: &#8220;the most durable thing we built this year is not the pipeline, it is a team with the proven capability to build, post-train, and ship LLMs from raw data at high velocity.&#8221;[25] Mistral already sells that eligibility: Hesse found that only it could meet every mandatory requirement.</p><p style="text-align: justify;">If you fund or buy AI in Europe, stop asking who will build our OpenAI. Nobody needs to, at about $3 million a run. A model like this is a component, like a database, and you should buy it like one. Ask what it costs, who trained it, whose switch it sits behind, and what you are really paying for. With Kolibri, you pay for a well-documented, competent model that is allowed to bid, not the one that scores highest.</p><h3><strong>Notes</strong></h3><p>[1] Aleph Alpha, <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1">Kolibri-1 model card</a>, Hugging Face, accessed October 5th, 2026: &#8220;Total parameters | 78B (78,103,074,560)&#8221;, &#8220;Active parameters / token | 3.46B (3,457,573,120)&#8221;, &#8220;License | Apache 2.0&#8221;, &#8220;Release Date | 3rd of October 2026&#8221;, and under Model Dependencies, &#8220;None. The model was trained from scratch.&#8221; Second document: Aleph Alpha, <a href="https://aleph-alpha.com/downloads/tech-report.pdf">&#8220;Kolibri: A Sovereign European Model on the Pareto Frontier&#8221;</a>, technical report, 189 pages, abstract. The launch post opens &#8220;On the Day of German Reunification&#8221;: Aleph Alpha, <a href="https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/">&#8220;Kolibri Has Landed&#8221;</a>. Both are the company&#8217;s own documents; no independent party has confirmed the parameter count beyond the Hugging Face file listing, whose tensor sizes sum to the same figure.</p><p>[2] As relayed by <a href="https://techcrunch.com/2024/09/05/german-llm-maker-aleph-alpha-pivots-to-ai-support">TechCrunch</a>, September 5th, 2024, relaying an interview given to <a href="https://www.bloomberg.com/news/articles/2024-09-05/the-rise-and-pivot-of-germany-s-one-time-ai-champion">Bloomberg</a> the same day. TechCrunch describes him as chief executive. The Bloomberg article is paywalled and was not read; the quote is reported, not confirmed at source.</p><p>[3] Cohere, <a href="https://cohere.com/blog/cohere-and-aleph-alpha-sign-agreement">&#8220;Cohere and Aleph Alpha sign agreement&#8221;</a>, September 16th, 2026: a &#8220;definitive business combination agreement&#8221;, &#8220;subject to final regulatory approvals&#8221;, &#8220;Operating globally as Cohere&#8221;, &#8220;following the release of our planned partnership in April of this year&#8221;. I call it a purchase on that basis, on <a href="https://www.handelsblatt.com/technik/ki/ki-aleph-alpha-veroeffentlicht-neues-sprachmodell-kolibri/100259581.html">Handelsblatt</a>, October 5th, 2026, which writes of a &#8220;De-facto-&#220;bernahme&#8221; by Cohere, and on <a href="https://www.reuters.com/legal/transactional/canadas-cohere-germanys-aleph-alpha-announce-merger-handelsblatt-reports-2026-04-24/">Reuters</a>, April 24th, 2026, which reported the deal as an acquisition at an undisclosed price and relayed Handelsblatt&#8217;s report that Cohere&#8217;s shareholders would hold about 90% of the combined company; neither company has stated a split. The release plans a company &#8220;dual-headquartered in Berlin and in Toronto&#8221;. Aleph Alpha&#8217;s own <a href="https://aleph-alpha.com/news/kolibri-souveraene-ki-made-in-germany/">release on Kolibri</a>, dated October 5th, 2026, says the company operates independently until closing.</p><p>[4] The <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1">Kolibri-1 model card</a>, Computing Resources row: &#8220;Hardware: 768 NVIDIA B200 (96 HGX 8xB200 nodes) [...] Time: 21 days (511h, 392k GPUh)&#8221;, &#8220;Mid-training: 5 days, 90k GPUh&#8221;, &#8220;Long-Context: 13h, 10k GPUh&#8221;. Second document: Aleph Alpha, <a href="https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/">&#8220;Kolibri Has Landed&#8221;</a>, October 3rd, 2026, which gives the same 768 B200s, 21 days and three stages. The technical report, Appendix B.7.2, gives a net figure of 377,000 GPU-hours for the pre-training run; I use the card&#8217;s figure. Author&#8217;s calculation from the card: 392,000 + 90,000 + 10,000 = 492,000.</p><p>[5] Verda, <a href="https://verda.com/pricing">pricing page</a>, read October 5th, 2026 and <a href="https://web.archive.org/web/20261005171250/https://verda.com/pricing">archived</a> the same day: &#8220;1x B200 SXM6 180GB [...] $7.13/h&#8221; on demand, and &#8220;2 years 25% off on-demand&#8221;. Author&#8217;s calculation from the card and the price list: 492,000 &#215; $7.13 = $3,507,960. Author&#8217;s calculation at the two-year rate: (492,000 &#215; $7.13) &#215; 0.75 = $2,630,970. For comparison, AWS listed a B200 at $12.355 per hour in <a href="https://aws.amazon.com/ec2/capacityblocks/pricing/">Capacity Blocks</a> on October 1st, 2026. The 25% is Verda&#8217;s published discount for single instances; its page says clusters above 144 GPUs are quoted on request, so the lower figure assumes a comparable discount. These are list prices; the contract between Aleph Alpha and Verda is not public.</p><p>[6] The <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, Acknowledgments, p. 114: &#8220;We thank Verda for supplying the GPU compute infrastructure behind this work&#8221;. Table 26, p. 95, gives &#8220;GPU hours, B300 equivalents&#8221; of 16,900 for the final reinforcement-learning run (and 5,500 for Kolibri Origin&#8217;s); Verda&#8217;s <a href="https://verda.com/pricing">price list</a> shows a B300 at $8.97 per hour on October 5th, 2026. Author&#8217;s calculation from the report and the price list: 16,900 &#215; $8.97 = $151,593. The <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1">model card</a> states that its energy estimate &#8220;excludes SFT and RL, peak, idle and low-load states, and proxy and ablation models&#8221;; I found no GPU-hour figure for fine-tuning, for the ablations or for data generation with the teacher models in the card, the report or the launch post. The restart is in <a href="https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/">&#8220;Kolibri Has Landed&#8221;</a>: &#8220;We stopped the Kolibri Origin pre-training after a few trillion tokens and restarted it from scratch&#8221;. Verda&#8217;s own account: <a href="https://verda.com/blog/aleph-alpha-kolibri">&#8220;Aleph Alpha releases Kolibri, a sovereign open-weight LLM trained on Verda&#8221;</a>, October 5th, 2026, which also says the model was trained &#8220;in European data centers we own and operate&#8221; on a fleet &#8220;custom-built to Aleph Alpha&#8217;s requirements&#8221;. It gives no GPU count, hours or price. Author&#8217;s calculation, assuming 1,000 GPUs for 90 days at the B200 rate: (1,000 &#215; 2,160) &#215; $7.13 = $15,400,800. My 2021 figure priced the purchase of DGX servers: <a href="https://huggingface.co/blog/large-language-models">&#8220;Large Language Models: A New Moore&#8217;s Law?&#8221;</a>, October 26th, 2021.</p><p>[7] The <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1">Kolibri-1 model card</a>, Model Architecture table and <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1/blob/main/config.json">config.json</a>: 50 layers, 384 routed experts and one shared per layer, six selected per token, a sliding window of 512 preceding tokens in 40 layers and full attention in 10, &#8220;float8_e4m3fn&#8221; weights, &#8220;Model memory footprint: ~78 GB&#8221;. On languages: the card&#8217;s &#8220;deliberate choice of depth over breadth&#8221;; the <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, section 2.3.2.2: &#8220;This leaves us with a lower bound of 4 T German tokens needed for training&#8221; and &#8220;The open datasets in the upper part of Table 9 add up to around 1.94T tokens&#8221;; its conclusion names &#8220;extending multi-lingual capabilities to European languages&#8221; as an area of interest. That each further language would cost about as much is my inference from those figures; the report does not say so. Forty of the fifty layers look back 512 tokens and ten read the whole context. The <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a> presents two methods as new, a tokenizer training method it calls UniBPE and a load-balancing technique (&#8221;We introduce LEI&#8221;); the rest it attributes to published work. Tokenizer: the card reports 4.7 bytes per token on German text, by Aleph Alpha&#8217;s measurement; the technical report gives 4.90.</p><p>[8] Aleph Alpha, <a href="https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/">&#8220;Kolibri Has Landed&#8221;</a>, section &#8220;Architecture and pre-training&#8221;. The 3 and the 18 are estimates, not measurements: the <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, Table 31, is headed &#8220;Estimated maximum number of concurrent 256k-token requests in FP8&#8221;, and the 18 belongs to an 82.3B test configuration, not the 78.1B model that shipped. Verda&#8217;s <a href="https://verda.com/pricing">price list</a> shows an H100 at $3.77 per hour on demand on October 5th, 2026. Author&#8217;s calculation from the price list: (2 &#215; $3.77) &#215; 730 = $5,504. The post&#8217;s sentence continues &#8220;and decodes 28% faster&#8221;, a figure from a different setup that I leave out. The report&#8217;s size sweep ran 42.1B, 82.3B and 122.6B; the post rounds differently.</p><p>[9] The <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>. Proxies, section 2.1: &#8220;The proxy at scale M has 0.6B active out of 7.8B total parameters trained on 97.6B tokens. The proxy at scale L has 2.1B active out of 30.6B total parameters trained on 327.6B tokens.&#8221; Scaling laws, section 2.2: &#8220;A common approach fits scaling laws that extrapolate along both dimensions jointly&#8221;; &#8220;For Kolibri Origin, we fit power laws to the optimal learning rate and batch size at horizons from 100B to 430B tokens and extrapolate them to its pre-training budget of 7.5T tokens. Extrapolating those laws to the Kolibri setting gives unrealistic hyperparameter predictions&#8221;; the replacement is &#181;P width transfer plus loss extrapolation on a proxy that &#8220;halves the width of the production model&#8221;. The Pareto quotation is from Figure 46; the title&#8217;s frontier is Figure 1. Size sweep, section 2.1.4: 42.1B, 82.3B and 122.6B total parameters at 3.4B active; &#8220;The 122.6B configuration improves loss and aggregate eval accuracy only slightly further&#8221;; at 4K context the &#8220;82.3B and 122.6B configurations are 13 % and 32 % slower than the 42.1B configuration&#8221;, measured on eight H100s. Pareto, section 3.3.3: &#8220;Qwen3.5 35B-A3B, Kolibri and Qwen3.8 27B span the convex hull of the Pareto frontier in both languages&#8221;, on a chart of the Overall scores against active parameters per token; Appendix A repeats the comparison against decoded text per second per GPU, measured by Aleph Alpha with vLLM on eight B200s. Reinforcement learning, section 3.2: &#8220;training it on its own rollouts against verifiable rewards&#8221;; &#8220;a single training run on 38 environments&#8221;; &#8220;All environments grade a sequence with a verifiable reward in [0, 1]&#8221;; &#8220;Our RL training runs on 256 NVIDIA B300 GPUs, split into 32 nodes with eight GPUs each&#8221;; &#8220;Serving the policy with FP8 weights and an FP8 KV cache raises the throughput of rollout generation, which can dominate the cost of RL, by a factor of two to three in our runs&#8221;; &#8220;Each non-controller node runs one sandbox executor on the CPUs that the run&#8217;s nodes keep idle, so runs never wait on a shared sandbox pool&#8221;. Table 26 gives 1,000 optimizer steps for the final run and 25,000 average concurrent sandboxes, against 480 for Kolibri Origin, and its caption says the earlier count covers code sandboxes while Kolibri&#8217;s &#8220;also includes agentic sandboxes that were added later&#8221;; judge model, section 3.2.2.1: the frozen judge &#8220;grades math proofs, simulates the user and grades natural-language assertions for the customer-service environment&#8221;; Table 58 gives &#8220;one per node on 31 nodes, 112 CPUs each&#8221;; 31 times 112 is 3,472, my calculation. Section 3.2.5 says the team &#8220;moved code execution to local sandboxes running on the CPUs of the training nodes&#8221;. The Together AI figures are from my piece <a href="https://www.airealist.ai/p/the-verification-tax">&#8220;The Verification Tax&#8221;</a>, April 10th, 2026, and its note 1. Environments, section 3.2.2 and Figure 32: &#8220;The final RL run trains on 38 environments, 31 English ones and 7 German variants&#8221;; &#8220;Software Engineering and Terminal Use (36.1 %)&#8221;; &#8220;Math and Reasoning (17.8 %)&#8221;; Table 55 lists every environment by name, and none carries an industry name. Table 57: &#8220;Five in-house proxies for deployed industry assistants&#8221;, namely &#8220;Semiconductors, German Public Sector, Aerospace, Automotive Supplier and Industrial Drive Technology&#8221;; Figure 44 says they &#8220;share no documents or questions with the training data&#8221;; section 3.3.3 says &#8220;Industrial Drive Technology (11 items) and Automotive Supplier (26) are too small to show a trend&#8221;. Utilization: I searched the report for MFU, FLOPs utilization and utilization on October 6th, 2026 and found no such figure for pre-training, SFT or RL. Figure 56 gives per-run median throughputs of 15,985 to 16,593 tokens per second per GPU; Appendix B.7.2: &#8220;We observed 38 unplanned interruptions during the net 377 000 GPU hours pre-training run&#8221;, and Table 35 totals the recovery time at 28.7 h. MFU estimate, my calculation: training costs about 6 floating-point operations per active parameter per token (2 forward, 4 backward; the rule ignores attention over the context), so 6 times 3,457,573,120 active parameters times 16,500 tokens per second is about 342 teraFLOPS per GPU. Lenovo&#8217;s <a href="https://lenovopress.lenovo.com/LP2226">product guide for the B200</a>, January 13th, 2026, gives &#8220;2.25 / 4.5 petaFLOPS&#8221; for BFLOAT16, &#8220;Without / with structural sparsity enabled&#8221;; NVIDIA&#8217;s <a href="https://www.nvidia.com/en-us/data-center/dgx-b200/">DGX B200 page</a> gives the 8-bit figure for eight GPUs. 342 over 2,250 is 15%. Adding the output layer (6 times 2,560 times 128,000 per token) and attention at the 16,000-token pre-training context (10 full-attention and 40 sliding-window layers of 512 tokens, 48 heads of 128, Table 2: 12 times 6,144 times 102,400 per token) raises the count from 20.7 to about 30.3 billion operations per token, or about 500 teraFLOPS per GPU, which is 22% of 2,250. The report does not state the precision of the pre-training run; at the 8-bit peak of 4.5 petaFLOPS the same throughput would be 8% to 11%. Qwen3.8-27B is one of the three teacher models in section 3.1.1. That a scaling law does not travel well is my reading of the report&#8217;s own result.</p><p>[10] Arcee AI, its <a href="https://www.arcee.ai/blog/trinity-large">launch post for Trinity Large</a>, January 27th, 2026: &#8220;All in&#8212;compute, salaries, data, storage, ops&#8212;we pulled off this entire effort for $20 million&#8221;, &#8220;We trained on 2048 Nvidia B300 GPUs&#8221;, and a pre-training run of 33 days. Second document: Varun Singh et al., <a href="https://arxiv.org/abs/2602.17004">&#8220;Arcee Trinity Large Technical Report&#8221;</a>, arXiv 2602.17004, whose abstract gives &#8220;400B total parameters and 13B activated per token&#8221;, so Trinity has five times Kolibri&#8217;s total parameters and nearly four times its active ones. The team of about thirty is as reported by <a href="https://www.implicator.ai/arcee-ai-releases-400b-open-reasoning-model-that-rivals-claude-at-96-lower-cost/">Implicator</a>. Kolibri&#8217;s <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a> cites &#8220;V. Singh et al. 2026&#8221; for sandwich normalization (&#8221;Each block uses sandwich normalisation&#8221;) and for its document buffer, both reused; for the load balancer of Kolibri Origin, which Kolibri replaced with its own; and, as what others do, for the normalization of routing scores.</p><p>[11] Arcee AI&#8217;s <a href="https://huggingface.co/arcee-ai/Trinity-Large-TrueBase">Trinity-Large-TrueBase</a> is described on its card as a &#8220;10T-token pre-anneal checkpoint with no instruction data&#8221;. Kolibri Origin: the <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, Table 1, &#8220;an internal model that is not released&#8221;. The public dataset is <a href="https://huggingface.co/datasets/Aleph-Alpha/Aleph-Alpha-GermanWeb">Aleph-Alpha-GermanWeb</a>, whose card says &#8220;The synthetic dataset contains the actual data&#8221;; I have not established that it is the same German text Kolibri was trained on. The carve-out is in the <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1">model card</a>, &#8220;License and terms&#8221;.</p><p>[12] The <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, section 2.3.2, pp. 24 to 27: &#8220;Approximately 24 % (4.8T tokens) of our pre-training data is synthetic in origin&#8221;, rephrased &#8220;with Gemma-4-26B-A4B&#8221;; &#8220;We use Mistral-Nemo-Instruct-2407 to rewrite organic German documents [...] making it the single largest source of German in the model.&#8221;</p><p>[13] The <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, section 3.1.1, p. 53. Second document: Aleph Alpha, <a href="https://aleph-alpha.com/downloads/data-summary.pdf">Public Summary of Training Content</a>, section 2.5, which lists Gemma-4-26B-A4B, Qwen3-32B, Qwen3.8-27B, Mistral-Nemo-12B, GLM-5.3, &#8220;GLM-5.2-fp8&#8221; (linked to the zai-org/GLM-5.2-FP8 repository) and Kimi-K2.6. That the weights were self-hosted is my inference from that listing; the report does not state how the GLM models were run. The summary does not list Kimi-K2.7, which the report uses in one reinforcement-learning environment.</p><p>[14] The <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>: Kimi-K2.7 and Kimi K2.6 in Appendix I.1, pp. 173 and 176; Qwen3.5-122B-A10B as judge and user simulator, section 3.2.3.2, p. 83; the filter in section 3.1.2.1, p. 60: &#8220;we judge the whole conversation, including reasoning traces, with gpt-oss-120b [...] Conversations of the first kind are dropped.&#8221; Arcee AI&#8217;s precedent, as I covered it at the time: my video of January 28th, 2025 on Virtuoso Lite and Virtuoso Medium v2, archived on <a href="https://julien.org/youtube/2025/20250128_Virtuoso_Lite_and_Virtuoso_Medium_v2_-_distilling_DeepSeek-V3_to_10B_32B.html">julien.org</a>. The report says in section 3.1.1.5, p. 59: &#8220;We acknowledge that several of our teacher models were built in China, so they carry known political biases on sensitive topics&#8221;.</p><p>[15] License files read on Hugging Face on October 5th, 2026: <a href="https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE">GLM-5.3</a> (&#8221;the Licensee must pass Z.AI&#8217;s security review&#8221; above &#8220;10 billion US dollars&#8221; of revenue for a &#8220;Model as a Service business&#8221;); <a href="https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE">Qwen3.8-27B</a>, identical to the canonical Apache 2.0 text apart from the copyright line; likewise <a href="https://huggingface.co/Qwen/Qwen3-32B/blob/main/LICENSE">Qwen3-32B</a>, <a href="https://huggingface.co/Qwen/Qwen3.5-122B-A10B/blob/main/LICENSE">Qwen3.5-122B-A10B</a> and <a href="https://huggingface.co/openai/gpt-oss-120b/blob/main/LICENSE">gpt-oss-120b</a>; <a href="https://huggingface.co/zai-org/GLM-5.2/blob/main/LICENSE">GLM-5.2</a>, MIT; <a href="https://huggingface.co/moonshotai/Kimi-K2.7-Code/blob/main/LICENSE">Kimi-K2.7-Code</a> and <a href="https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENSE">Kimi-K2.6</a>, &#8220;Modified MIT&#8221;, with a display requirement above 100 million monthly active users or $20 million of monthly revenue. The eleven models are Gemma-4-26B-A4B and Gemma-4-31B, Mistral-Nemo-Instruct-2407, Qwen3-32B, Qwen3.8-27B and Qwen3.5-122B-A10B, GLM-5.2 and GLM-5.3, Kimi-K2.6 and Kimi-K2.7-Code, and gpt-oss-120b: seven Apache 2.0, one MIT, three with terms of their own. Gemma 4 is Apache 2.0 on <a href="https://ai.google.dev/gemma/docs/gemma_4_license">Google&#8217;s license page</a>; <a href="https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407">Mistral NeMo</a> is marked Apache 2.0 on its model card and carries no license file. The twelfth, Cohere&#8217;s Command A+, used as one judge, is tagged Apache 2.0 on Hugging Face; I did not read its license file. These are the licenses of the downloadable weights; a hosted API has its own terms. Reflection AI&#8217;s own post, <a href="https://reflection.ai/blog/introducing-beam">&#8220;Introducing Beam&#8221;</a>, October 5th, 2026: &#8220;This month, we will release the weights under an Apache 2.0 license&#8221;; it describes Beam as &#8220;competitive with larger open models like GLM 5.2&#8221; and offers early access by waitlist. As of October 6th, 2026, 01:19 UTC, I found no Beam weights on Hugging Face. Also <a href="https://techcrunch.com/2026/10/05/reflection-debuts-beam-a-open-weight-ai-model-to-rival-chinese-models-at-lower-compute-cost/">TechCrunch</a>, October 5th, 2026: &#8220;Reflection says it will release Beam&#8217;s weights and full technical details this month&#8221;; its performance claims &#8220;haven&#8217;t been independently verified&#8221;. Announced, not released, as of that date.</p><p>[16] NSA, CISA and FBI, <a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a">advisory AA26-251A</a>, September 8th, 2026, executive summary. The advisory names &#8220;DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI&#8221;. The advisory separates legitimate distillation from what it calls malicious distillation; it describes fraudulent accounts as a technique of &#8220;China-based entities&#8221; without tying it to each of the six. A search of the advisory for &#8220;GLM&#8221; and &#8220;Qwen3.8&#8221; returns nothing: it names companies and model families, not the models Kolibri used. Its claims are the agencies&#8217; and have not been tested in court. Kolibri&#8217;s own use is of published weights, which the advisory does not concern. I searched on October 5th, 2026 for a published response from Z.ai or Alibaba and found none; <a href="https://www.theregister.com/ai-and-ml/2026/09/09/us-claims-chinese-ai-companies-core-ai-strategy-is-distilling-american-models/5295171">The Register</a>, September 9th, 2026, reported that China had answered by accusing US companies of distilling Chinese models, and wrote that it had sought comment from Chinese AI companies; its story prints none.</p><p>[17] Bastian Boll, <a href="https://aleph-alpha.com/en/blog/training-on-the-party-line/">&#8220;Training on the Party Line&#8221;</a>, Aleph Alpha, September 28th, 2026. The six Chinese models are two each from Qwen, DeepSeek and Kimi; prompts were generated and answers judged by gpt-oss-120b. Its Figure 1 gives Claude Sonnet 5 at 70% balanced and Mistral Small 2603 at 92%, and the two Alibaba models it tests, Qwen 3.6 35B-A3B and Qwen 3.8 2.4T-A95B, at 17% and 19%; neither is the Qwen3.5 35B or the Qwen3.8-27B of Kolibri&#8217;s results table, and the post concedes that &#8220;The difficulty of these judgement calls might cause judge bias&#8221;. The post: &#8220;sovereign models need three things: screening of training data for Chinese political content, targeted alignment data that sets the intended behavior, and evaluation against benchmarks like the one presented here. We have adopted these measures at Aleph Alpha.&#8221; The warning is in the <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1">model card</a>, &#8220;Political Bias&#8221;. No score for Kolibri on this benchmark appears in the card, the report, the launch post or the study, searched October 5th, 2026.</p><p>[18] My earlier positions: <a href="https://www.airealist.ai/p/build-buy-or-download-someone-elses">&#8220;Build, Buy, or Download Someone Else&#8217;s Politics&#8221;</a>, March 6th, 2026, on Singapore, and <a href="https://www.airealist.ai/p/too-dangerous-for-you-free-for-everyone">&#8220;Too Dangerous for You, Free for Everyone&#8221;</a>, June 28th, 2026: &#8220;While the United States restricts and Europe regulates, Chinese labs ship.&#8221;</p><p>[19] Chips and sites: notes 4 and 6; the <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, p. 4, says the model was trained &#8220;on infrastructure in Germany and Finland&#8221;. Methods, same report, section 2.1: &#8220;The routing scores are sigmoids of the router logits (DeepSeek-AI 2024b)&#8221;; its Exact Quantile Balancing extends &#8220;Quantile Balancing (QB) (Kimi Team 2026b)&#8221;; Arcee AI&#8217;s Trinity: note 10. Rewriters, teachers and filter: notes 12 to 13. Code of practice: the <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1">model card</a>, &#8220;Aleph Alpha is a signatory of the EU GPAI Code of Practice&#8221;. Buyer: note 3; the transaction had not closed as of October 5th, 2026. NVIDIA&#8217;s <a href="https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-SFT-Data">dataset card</a> for Nemotron Cascade 2 lists &#8220;responses generated by DeepSeek-V3.2&#8221; for its math data and Qwen3 models for its agent conversations. Mistral&#8217;s catalog: Mistral&#8217;s <a href="https://mistral.ai/news/regional-inference-open-models-new-compute/">post of August 11th, 2026</a> says its platform &#8220;will support third-party open models, starting with Z.ai&#8217;s GLM-5.2&#8221;; Mistral&#8217;s <a href="https://docs.mistral.ai/resources/changelogs">changelog</a> lists GLM 5.3 as generally available from September 28th, 2026 and GLM 5.2 as retiring on October 31st, 2026; see also <a href="https://www.airealist.ai/p/vent-mauvais">&#8220;Vent Mauvais&#8221;</a>.</p><p>[20] My piece <a href="https://www.airealist.ai/p/too-dangerous-for-you-free-for-everyone">&#8220;Too Dangerous for You, Free for Everyone&#8221;</a>, June 28th, 2026: &#8220;A frontier-parity model small enough to self-host cheaply (a step change in compression, or a model under 100 billion parameters that matches the leaders) would open a switch-free door for real&#8221;. Aleph Alpha&#8217;s own definition, in <a href="https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/">&#8220;Kolibri Has Landed&#8221;</a>: &#8220;Sovereignty, for us, combines two dimensions: how we built the model, and how it transfers to our customers.&#8221; GLM licenses: note 15.</p><p>[21] The <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1">Kolibri-1 model card</a>, Evaluation, post-training table, rows &#8220;Overall (EN)&#8221; and &#8220;Overall (DE)&#8221;: Kolibri 75.5 and 70.8, GPT-OSS 120B 72.3 and 70.2, Qwen3.5 35B-A3B 74.7 and 69.8, Qwen3.8 27B 80.2 and 79.9; the German knowledge average: Kolibri 57.6, GPT-OSS 120B 58.0, Gemma 4 26B-A4B 61.5, Qwen3.5 35B-A3B 61.3, Qwen3.6 35B-A3B 61.0. Author&#8217;s calculation from the card, English: 80.2 &#8722; 75.5 = 4.7. Author&#8217;s calculation from the card, German: 79.9 &#8722; 70.8 = 9.1. Two community tests have appeared since the release, each scoring Kolibri alone: a <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1/discussions/4">tool-calling run</a> posted on Hugging Face on October 4th, 2026 (&#8221;Final Score: 86 / 100&#8221;, with two safety-critical failures), and Stefan Beierle&#8217;s <a href="https://github.com/sbeierle/kolibri1-praxistest">practice test</a> of October 5th, 2026, which reports 96% on a German compliance check with the statute in the prompt and 53% without, and declares no connection to Aleph Alpha. Neither tests another model. The <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, section 3.3.3, p. 98, says the dense Qwen3.8 27B &#8220;activates nearly 8 times as many parameters per token as Kolibri&#8221;. All scores are Aleph Alpha&#8217;s own runs. The card&#8217;s row &#8220;German Public Sector&#8221;, one of its in-house customer proxies: Kolibri 75.0, Qwen3.5 35B-A3B 80.0, Qwen3.8 27B 89.0. The full row also has GPT-OSS 120B at 77.0, Qwen3.6 35B-A3B at 72.0 and Mistral Small 4 119B-A6B at 50.0; Small 4 is not Mistral&#8217;s best model and the notice does not say which model Hesse bought. Hesse: <a href="https://ted.europa.eu/de/notice/453996-2026/html">TED notice 453996-2026</a>, a &#8220;Verhandlungsverfahren ohne Aufruf zum Wettbewerb&#8221; (negotiated procedure without a call for competition); its market survey found that &#8220;lediglich Mistral AI SAS s&#228;mtliche zwingenden Mindestanforderungen kumulativ erf&#252;llen kann&#8221;; the requirement reads &#8220;Die verwendeten KI-Modelle m&#252;ssen innerhalb des Europ&#228;ischen Wirtschaftsraums oder in L&#228;ndern mit angemessenem Datenschutzniveau gem&#228;&#223; Art. 45 DSGVO entwickelt worden sein&#8221;; the notice says the lack of competition &#8220;ist nicht Ergebnis einer k&#252;nstlichen Einschr&#228;nkung der Auftragsvergabeparameter&#8221;. It is one buyer&#8217;s requirement for one purchase, not a law. I have not established whether a model developed in the United States, such as GPT-OSS 120B, would meet it. I covered the award in <a href="https://www.airealist.ai/p/vent-mauvais">&#8220;Vent Mauvais&#8221;</a>, September 24th, 2026. The 80% is Qwen 3.6 35B-A3B in Figure 1 of &#8220;Training on the Party Line&#8221; (note 17), a different model from the Qwen3.5 35B that scores 80.0; neither Qwen that outscores Kolibri on that row was tested in the study; no document gives Kolibri&#8217;s score on that benchmark. That Kolibri would be bought for its eligibility is my reading.</p><p>[22] The <a href="https://huggingface.co/Aleph-Alpha/Kolibri-1">Kolibri-1 model card</a>, same table, column &#8220;Mistral Small 4 119B-A6B&#8221;: 63.1 and 61.4. Author&#8217;s calculation from the card, English: 75.5 &#8722; 63.1 = 12.4. Author&#8217;s calculation from the card, German: 70.8 &#8722; 61.4 = 9.4. The <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, section 3.3.3, says Kolibri &#8220;trails behind Qwen3.6 35B-A3B and Mistral Small 4 119B-A6B on the AA-Omniscience Index&#8221;. NeMo&#8217;s German text &#8220;had already been produced for the Aleph-Alpha-GermanWeb project&#8221; (same report, p. 27). The title: Mistral, <a href="https://mistral.ai/news/mistral-large-2407">&#8220;Large Enough&#8221;</a>, July 24th, 2024. Scheer&#8217;s words are in the Cohere release of note 3, where he is Co-CEO; he has been sole chief executive since September 28th, 2026, per Aleph Alpha&#8217;s <a href="https://aleph-alpha.com/news/reto-spoerri-verlaesst-aleph-alpha/">release of that day</a>. Cohere on Europe: <a href="https://cohere.com/blog/cohere-triples-uk-footprint-with-new-london-office-to-support-r-and-d-growth">&#8220;Cohere triples UK footprint&#8221;</a>, June 15th, 2026. Aidan Gomez&#8217;s stated reason, in the same release: &#8220;we&#8217;re joining forces with Aleph Alpha to enhance the talent, infrastructure and institutional trust behind our mission&#8221;. That Cohere wants Aleph Alpha in order to go deeper into Mistral&#8217;s market is my reading; neither company&#8217;s announcement names Mistral, and Cohere&#8217;s says &#8220;further integration and product details to follow&#8221;.</p><p>[23] Aleph Alpha, <a href="https://aleph-alpha.com/en/news/half-billion-dollar-investment-round/">release of November 6th, 2023</a>: &#8220;The investment is led by the Innovation Park Artificial Intelligence (Ipai), Bosch Ventures and the companies of Schwarz Group&#8221;; &#8220;This includes preconsumption licenses with the global industry leaders of the consortium&#8221;. The release lists SAP among the &#8220;Other new investors&#8221;. Gr&#252;nderszene&#8217;s breakdown, relayed by <a href="https://the-decoder.com/german-ai-startup-aleph-alpha-breaks-down-500-million-investment-package/">The Decoder</a>, splits the round into 110 million euros of equity, 300 million of research funding and 60 million of order commitments. The 2024 license: the <a href="https://huggingface.co/Aleph-Alpha/Pharia-1-LLM-7B-control">Pharia-1-LLM-7B card</a>, &#8220;which limits the usage to educational and research purposes&#8221;; the same card says &#8220;We provide our customers with open access to our full model checkpoint including weights and code for commercial use&#8221;. &#8220;Prepaid orders&#8221; is my gloss on &#8220;preconsumption licenses&#8221;.</p><p>[24] The outlet <a href="https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi">Axios</a>, July 20th, 2026, reported that the administration &#8220;could ban cutting-edge Chinese AI models&#8221; and described approaches short of an outright ban, with Entity List designations under discussion; <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/trump-administration-reportedly-reviving-push-to-ban-chinese-ai-models-following-kimi-k3-launch-citing-cybersecurity-concerns-downloadable-open-weights-could-make-an-outright-u-s-ban-nearly-impossible-to-enforce-amid-growing-adoption">Tom&#8217;s Hardware</a> relayed it on July 22nd. Reported, not enacted as of October 5th, 2026. I covered it in <a href="https://www.airealist.ai/p/the-model-is-a-checkpoint">&#8220;The Model Is a Checkpoint&#8221;</a>, July 23rd, 2026. No export rule restricts B200s to Germany or Finland today that I know of; the chip switch is a possibility, not a measure in force. GLM licenses: note 15.</p><p>[25] Aleph Alpha, <a href="https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/">&#8220;Kolibri Has Landed&#8221;</a>. The 1.2 million reinforcement-learning tasks are &#8220;internally curated&#8221; in the <a href="https://aleph-alpha.com/downloads/tech-report.pdf">technical report</a>, abstract.</p>]]></content:encoded></item><item><title><![CDATA[Eight Changes to Transformer Attention, and What Each Costs]]></title><description><![CDATA[The newest open models from Alibaba, Moonshot, DeepSeek and Z.AI no longer read every earlier token. Eight designs, what each costs, and what to test.]]></description><link>https://www.airealist.ai/p/eight-changes-to-transformer-attention</link><guid isPermaLink="false">https://www.airealist.ai/p/eight-changes-to-transformer-attention</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Sun, 04 Oct 2026 13:35:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!SDde!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SDde!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SDde!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!SDde!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!SDde!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!SDde!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SDde!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:611266,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/218727599?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SDde!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!SDde!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!SDde!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!SDde!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff711d726-e2bd-4674-b852-d585b2f4485b_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">Open the config file for Qwen3.8-Flash-Next, released by Alibaba in August 2026, and count the layers. There are 48. Thirty-six are labelled <code>linear_attention</code>. The other 12 are labeled <code>full_attention</code>and they are not: a few lines up, the file gives them an indexer, a small network that picks which tokens the layer reads, with a budget of 2,048 tokens [1]. No layer in this model does what the original Transformer&#8217;s layers did in 2017: read every earlier token in full.</p><p style="text-align: justify;">It is not just one model. Alibaba, Moonshot, DeepSeek and Z.AI now ship models in which about one layer in four still reads every earlier token, or none does [2].  The labs&#8217; case for these designs is cost: by DeepSeek&#8217;s own estimate, its V4-Pro needs a tenth of the cache its previous model did at a million tokens [3]. </p><p style="text-align: justify;">But while Kimi K3 and DeepSeek V4 offer a million tokens of context, and Qwen3.8 offers 262,144, extensible to a million [2][3], only one of the ten published full-attention comparisons I count below goes past 128,000 tokens. So, if you are picking an open model for long documents or for agents, this one is for you.</p><h2><strong>What the config files say</strong></h2><p style="text-align: justify;">A model card is prose. <code>config.json</code> is what the inference engine actually loads, so that is where I looked, on October 1st, 2026 [2].</p><p style="text-align: justify;">Kimi K3, Moonshot&#8217;s 2.8-trillion-parameter model, has 93 layers: 69 linear layers and 24 full-attention layers. Alibaba&#8217;s largest Qwen3.8 has 92: 69 and 23. DeepSeek V4 has no full-attention layer at all: it reads the last 128 tokens exactly, and anything older only through compressed summaries, some of which an indexer selects. Z.AI&#8217;s GLM-5.3-Flash has 34 linear layers and 11 sparse ones, with a selection budget of 2,048.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SF5f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SF5f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png 424w, https://substackcdn.com/image/fetch/$s_!SF5f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png 848w, https://substackcdn.com/image/fetch/$s_!SF5f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png 1272w, https://substackcdn.com/image/fetch/$s_!SF5f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SF5f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png" width="727.9948120117188" height="692.9950614342323" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1386,&quot;width&quot;:1456,&quot;resizeWidth&quot;:727.9948120117188,&quot;bytes&quot;:303943,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/218727599?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SF5f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png 424w, https://substackcdn.com/image/fetch/$s_!SF5f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png 848w, https://substackcdn.com/image/fetch/$s_!SF5f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png 1272w, https://substackcdn.com/image/fetch/$s_!SF5f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f5cc7f-5f65-46b1-ac1b-1c278cc9a4cc_2400x2284.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Layer counts read from each model&#8217;s config file [2][4].</em></figcaption></figure></div><p style="text-align: justify;">One warning before you go and check. GLM-5.3-Flash does what Qwen3.8-Flash-Next does: it lists its 11 sparse layers under a key called <code>full_attn_layers</code> [2]. The giveaway is elsewhere in the same file, where <code>layer_types</code> calls those layers <code>deepseek_sparse_attention</code>, with a budget of 2,048. Read the neighbors, not just the label.</p><p style="text-align: justify;">Not everyone moved. Mistral Medium 3.5, IBM&#8217;s Granite 4.2, and K2-Horizon from the IFM institute still use full attention in every layer [4]. Google, Meta, OpenAI, and Thinking Machines mix it with layers that only see a recent window. NVIDIA and IBM have used a third option, Mamba-2. And Ai2 trained a 7-billion-parameter model on the same linear layers Alibaba uses [4]. A survey by Alibaba&#8217;s own researchers counts 17 hybrid releases, i.e., models that mix layer types, out of 27 in 2026, against 9 of 25 in the two years before, and concludes &#8220;continued coexistence rather than a universal replacement path&#8221; [5]. Fair enough. Still, of the eleven high-performing open-weight models that survey compares, only one uses full attention in every layer [5].</p><h2><strong>Why long context gets expensive</strong></h2><p style="text-align: justify;">A full-attention layer keeps a key and value for every token it has seen (the KV cache) and reads all of them to write the next token; that cache is what you rent when a provider sells you &#8220;cached input&#8221;. Two bills grow with it: compute, because work per token grows with context and reading a whole prompt grows quadratically, and memory, because the cache grows with every token, for every user at once.</p><p style="text-align: justify;">Agents made both bills urgent because they resend their whole history at every step: tool output, files, earlier turns. Moonshot&#8217;s and DeepSeek&#8217;s papers both open on the cost of long contexts and long-running agents [3][6]. DeepSeek&#8217;s tenth-of-the-cache model above uses 27% of the compute per token, both at a million tokens [3]. Moonshot&#8217;s Kimi Linear claimed &#8220;up to 75%&#8221; less cache [6].</p><p style="text-align: justify;">What does that buy? More users per GPU and more context per user, i.e., a lower cost per token. On September 29th, 2025, the post that announced DeepSeek&#8217;s first sparse-attention model also cut its price per million tokens from $0.56 to $0.28 for uncached input and from $1.68 to $0.42 for output [7]. List prices don't tell me whether buyers see the savings. On October 1st, 2026, DeepSeek&#8217;s own V4-Pro listed at $1.32 per million uncached input tokens and $3.96 per million output tokens at peak hours, above its 2025 price, and Kimi K3, a far larger model, at $3.00 and $15.00 [7].</p><p style="text-align: justify;">One thing I looked for and did not find: export controls, sanctions, HBM, memory prices, or shortages cited as reasons for the design in the DeepSeek V4 report, the Kimi K3 technical report, or the Kimi K3, GLM-5.3-Flash, and Qwen3.8 model cards [8]. Z.AI comes closest, saying GLM-5.3-Flash&#8217;s pre-release traffic was &#8220;served on Chinese AI chips&#8221;, which &#8220;are primarily constrained by memory capacity and bandwidth&#8221; [8]. Z.AI doesn't say whether that's also why the model looks the way it does, and I won&#8217;t say it for them.</p><h2><strong>Eight attention designs: what each saves and gives up</strong></h2><p style="text-align: justify;">Each of the eight answers the same questions: what a layer stores and reads, what that saves, what it gives up, who ships it, and what to test. They run in order of what each still reads, not of date: every token (1 and 2), some layers still do (3 to 5), none does (6 to 8). The table gives the dates. For the mechanics of attention and the KV cache, I have three step-by-step videos [9].</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!df5G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!df5G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png 424w, https://substackcdn.com/image/fetch/$s_!df5G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png 848w, https://substackcdn.com/image/fetch/$s_!df5G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png 1272w, https://substackcdn.com/image/fetch/$s_!df5G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!df5G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png" width="1456" height="1523" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1523,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:402598,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/218727599?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!df5G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png 424w, https://substackcdn.com/image/fetch/$s_!df5G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png 848w, https://substackcdn.com/image/fetch/$s_!df5G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png 1272w, https://substackcdn.com/image/fetch/$s_!df5G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec25fb57-5fd9-473f-b118-d129bb5bf973_2400x2510.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>The read budget, step by step [2][4][10][11][12][13].</em></figcaption></figure></div><p style="text-align: justify;"><strong>1. Grouped-query attention.</strong> The heads of a layer, its parallel readers, share keys and values in groups instead of each keeping its own; every token is still read. It saves memory and costs some accuracy: on a 7-billion-parameter test, DeepSeek measured 41.2 on a general-knowledge benchmark with eight groups against 45.2 without [10]. Even K2-Horizon, whose config has no window, linear, or latent layers, &#8220;retains full GQA&#8221; [4][5]. What to test: nothing specific to this step.</p><p style="text-align: justify;"><strong>2. Latent attention (MLA).</strong> The key and value for each token are compressed into a single vector of 576 numbers per layer, which the heads read directly, still for every token, every time [10]. In May 2024, DeepSeek ran the same model both ways, and the compressed version won 7 of 8 comparisons, with a cache between 4% and 14% of the original [10]. Kimi&#8217;s full-attention layers use it too. What to test: nothing specific to this step.</p><p style="text-align: justify;"><strong>3. Sliding-window attention.</strong> Some layers keep only the most recent 128 to 2,048 tokens, and the rest keep everything: OpenAI&#8217;s gpt-oss alternates one and one, Gemma 4 and Thinking Machines&#8217; Inkling run five windowed layers for each full one [4]. The windowed layers save memory and compute and cannot see far. MiniMax tried it on its M2 model: on a task that asks which words come up most often in the context, no difference at 32,000 tokens and a fall from 90.0 to 72.0 at 128,000 [14]. What to test: the same task at 32,000 tokens and at your real length.</p><p style="text-align: justify;"><strong>4. State-space layers (Mamba-2).</strong> These layers store nothing per token. Each keeps a fixed-size state, a running summary of everything it has read so far, and rewrites it with every token, so nothing grows. That saves memory and compute, but it gives up the per-token key and value: there is nothing to go back to, and the more the state absorbs, the less of it comes back exactly. At 2.7 billion parameters, it tied a full-attention model (60.2 to 60.2) and did slightly better with six attention layers mixed in [11]. NVIDIA&#8217;s Nemotron 3 Ultra ships 48 of these layers, with 12 attention layers [4]. What to test: finding one exact name or number from early in a long context.</p><p style="text-align: justify;"><strong>5. Linear attention (Gated DeltaNet, KDA).</strong> The state is again fixed in size, and each new token does three things to it: fade everything by a learned amount, weaken whatever was stored under a similar key, write the new entry [12]. That targeted weakening is the change from Mamba-2: the layer can revise a fact instead of blurring it. But a fixed state holds only so much. The paper that introduced the layer shows the cost: across six retrieval and question-answering tasks, the pure layer averaged 30.6 against a full-attention model&#8217;s 37.0, on inputs of only 2,000 tokens [12]. So no model in this piece ships the pure layer; the ones that use it ship a hybrid: three linear layers, then one full-attention layer for the exact lookups. Qwen3-Next shipped that in September 2025; Moonshot followed in October with its own variant, Kimi Delta Attention (KDA). Kimi K3 and Qwen3.8 use about the same layout at the trillion-parameter scale (69 linear layers to 24 and to 23): Kimi with KDA, which lets each channel forget at its own rate, and Qwen with Gated DeltaNet [2][6].</p><p style="text-align: justify;">Why three to one? Moonshot tried five ratios on a small model, and 3:1 came out a hair ahead of 1:1 [6]. Mamba-2&#8217;s authors had found about 10% attention worked best [11]. Moonshot&#8217;s own &#8220;up to 6&#215;&#8221; faster decoding is, in its words, &#8220;theoretical&#8221;; measured one request at a time, it is 2.3&#215; [6]. What to test: retrieval of many items at once, at the context length you will really run.</p><p style="text-align: justify;"><strong>6. Sparse attention with an indexer.</strong> Keep every token&#8217;s key and value, but put an indexer in front of the layer. DeepSeek&#8217;s lightning indexer scores every stored token cheaply, passes the best 2,048 to the layer, and the layer reads only those [13]. This saves compute, though less than it sounds: the indexer still scores every token for every new one, so its work still grows with the square of the context; it is just much cheaper work than reading [13]. DeepSeek&#8217;s paper claims no memory savings, and the indexer keeps its own keys on top of the cache. It gives up whatever the indexer fails to pick. Against its predecessor on 14 benchmarks, and after further training, the first sparse DeepSeek scored higher on 7, lower on 6, and tied on 1 [13]. Z.AI uses sparse attention in all 78 layers of GLM-5.3 [2]. Its GLM-5 report calls the mechanism &#8220;lossless by construction&#8221;; its own conversion test, on the 31-billion-parameter GLM-4.7-Flash, came out 0.35 points behind at 128,000 tokens, 78.86 against 79.21, and 1.72 ahead at 64,000 [15]. What to test: any task whose answer needs more than a couple of thousand tokens of evidence at once, such as counting how often something occurs.</p><p style="text-align: justify;"><strong>7. Pooled cache with sparse reads.</strong> Now the stored entries themselves get merged. In DeepSeek V4, about half the layers slide an 8-token window along the text in steps of 4 and pool each window into one entry, so neighboring entries overlap, and an indexer can choose among them. The other half pools every 128 tokens into one entry and reads all of those. Every layer also keeps the last 128 tokens unpooled, and in V4-Flash the first two layers keep nothing else [3]. It saves memory and compute; DeepSeek&#8217;s 10% and 27%, quoted earlier, are for the whole model, pooling included. It gives up token-level detail for anything older than 128 tokens. DeepSeek&#8217;s report has no test that isolates the new attention: it compares V4 with V3.2, a different model, and the word &#8220;ablation&#8221; does not appear in its 58 pages [3]. DeepSeek prints its own curve, and it's honest: 0.92 at 128,000 tokens and 0.59 at a million on an eight-needle test, where eight identical requests are hidden in a long conversation, and the model must return the one it is asked for [3]. An outside team at ByteDance found something stranger in the newer V4.1-Flash: retrieval accuracy depends on whether the target sits in an odd or even position, about five points apart on average [16]. DeepSeek&#8217;s own V4.1-Flash paper says its architectural changes &#8220;create robustness boundaries that have yet to be fully characterized&#8221; [16]. What to test: retrieval at each position, as well as the average.</p><p style="text-align: justify;"><strong>8. Linear plus sparse.</strong> The last step combines step 5 and swaps its full-attention layer for a sparse one. Qwen3.8-Flash-Next and GLM-5.3-Flash were uploaded on the same day, August 26th, 2026, with the same idea: three linear layers, then one sparse layer [2]. It saves and gives up what 5 and 6 do. Z.AI&#8217;s is assembled from other labs&#8217; designs: Moonshot&#8217;s KDA for the linear layers and DeepSeek&#8217;s sparse attention for the fourth; its config file names both, and its launch post, which compares the model with Kimi and DeepSeek, names neither KDA nor Moonshot [17]. Alibaba tested the swap on its shipped model: between 512,000 and a million tokens, the sparse fourth layer scored 93.00 on RULER, a long-context benchmark that mixes needle lookups with tracing and counting tasks, against 90.08 with full attention [18]. What to test: everything under 5 and 6.</p><h2><strong>Ten tests against full attention, and what is missing</strong></h2><p style="text-align: justify;">The labs&#8217; case first. They said what they changed: DeepSeek&#8217;s launch post names its &#8220;Novel Attention&#8221;, Moonshot&#8217;s names KDA, and the config files are public [19]. Six labs ran the old design against the new and published the results, including losses.</p><p style="text-align: justify;">I count ten such comparisons of designs that now ship in a model, from six labs: one win for the efficient design, seven mixed results and two losses [20]. The small-model tests in steps 4 and 5 are not in the count. I judge each comparison on every result the paper reports: a win is more than a point ahead somewhere and never more than a point behind; a tie is within a point everywhere; and mixed is more than a point apart in both directions. None of the long-context tables print error bars. One mixed result, Ai2&#8217;s, is against a baseline that already used sliding windows in three layers of four; another, Alibaba&#8217;s at 125 billion parameters, swaps one layer in four inside a hybrid against the same model with full attention in that layer.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HrV2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HrV2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png 424w, https://substackcdn.com/image/fetch/$s_!HrV2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png 848w, https://substackcdn.com/image/fetch/$s_!HrV2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png 1272w, https://substackcdn.com/image/fetch/$s_!HrV2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HrV2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png" width="1456" height="1310" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1310,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:468234,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/218727599?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HrV2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png 424w, https://substackcdn.com/image/fetch/$s_!HrV2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png 848w, https://substackcdn.com/image/fetch/$s_!HrV2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png 1272w, https://substackcdn.com/image/fetch/$s_!HrV2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff391fb09-e4da-4b8d-a56a-ac3d12bc29f5_2400x2160.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Ten published comparisons, judged on every result each paper reports [14][6][13][15][20][18].</em></figcaption></figure></div><p style="text-align: justify;">Moonshot&#8217;s was a controlled test: three 48-billion-parameter models, same data, same recipe, and the KDA hybrid beat full attention 84.3 to 81.3 at 128,000 tokens on RULER. That model also dropped positional encoding from its full-attention layers; with it, the score was 78.8. The plain Gated DeltaNet hybrid scored 80.5 in the same test, just under full attention. Full attention stayed more than a point ahead on three other benchmarks in that paper, so I count the test as mixed [6]. </p><p style="text-align: justify;">Alibaba ran a 25-billion-parameter version of the experiment. Its Gated DeltaNet hybrid beat the full-attention model on eight of nine benchmarks, none of them long-context, and trailed on the remaining one, EvalPlus, 49.71 to 51.01, so I count it as mixed [18]. </p><p style="text-align: justify;">Ai2&#8217;s hybrid led on RULER at 64,000 tokens, 76.9 to 70.9, and trailed at 4,000, 92.8 to 95.8: mixed as well [20]. </p><p style="text-align: justify;">MiniMax converted its full-attention model to sparse attention and matched the headline score (72.12 to 72.00) but fell short on the split [15]. </p><p style="text-align: justify;">Z.AI&#8217;s GLM-4.7-Flash conversion, above, is the one win: 1.72 points ahead at 64,000 tokens and never more than a point behind [15]. Z.AI ran one more test, on GLM-5&#8217;s own base models: sparse attention against the full-attention parent on four tasks at 128,000 tokens, two up, one down and one level [15].</p><p style="text-align: justify;">Two lost, and both were conversions to linear or sliding-window layers: an existing model, changed and then trained further, for 190 billion tokens at Z.AI and hundreds of billions to trillions at MiniMax. Of the four sparse conversions, one came out ahead and three mixed [13][15]. </p><p style="text-align: justify;">MiniMax&#8217;s loss is the sliding-window result above. In a February 2026 report, Z.AI converted half the layers of a 9-billion-parameter model to Gated DeltaNet, trained it on 64,000 tokens, and lost 11 points on RULER at 128,000. Its text says these methods lose &#8220;up to 5.69 points&#8221;; its table shows 11.28 for this one. It blamed &#8220;the unavoidable information loss introduced by efficient attention mechanisms during continual-training adaptation&#8221;, and its report has no test of the same layers trained from scratch [15]. In August, it shipped GLM-5.3-Flash, a new base model with 34 linear layers. Its launch post speaks of &#8220;preserving precise long-context capabilities&#8221;; neither the post, the model card, nor the docs give a long-context number [17]. On October 30th, 2025, MiniMax posted that &#8220;efficient attention still has some way to go before it can definitively beat full attention&#8221; [14]. Moonshot published its 48-billion-parameter RULER result that same day [6], and by the following June MiniMax had its own sparse attention [15]. On my count, MiniMax&#8217;s sentence still holds.</p><p><strong>So what is missing?</strong> </p><p>Three things, and they are narrower than &#8220;nobody checked&#8221;.</p><p style="text-align: justify;">The first is size. None of the four flagships has been compared with a same-size model built only with full-attention layers. Z.AI&#8217;s GLM-5 comes nearest: its comparison table does not give the models&#8217; size, but the text says the sparse run starts from GLM-5&#8217;s own 744-billion-parameter base model. It stops at 128,000 tokens [15]. Alibaba&#8217;s 125-billion-parameter test swaps one layer in four inside a hybrid, not the whole model [18]. Kimi K3 rests on the 48-billion-parameter test; its own report has no such comparison [20].</p><p style="text-align: justify;">The second is the task type. Needle tests are one place where sparse attention did well. When MiniMax broke down its 128,000-token results, sparse attention won in-context learning by 2.40 points and needle lookup by 2.24, and lost reranking by 2.10, word-frequency counting by 1.35, and, by exactly 1.00 each, question answering and variable tracking [15]. The gaps are small. But reranking, tracking, and question answering are closer to agent work than needle lookup, and the agent benchmarks in these reports don't isolate the attention design.</p><p style="text-align: justify;">The third is depth, and it depends on what you ask. RULER holds: Alibaba&#8217;s 93.00 (above) and Moonshot&#8217;s Kimi Linear&#8217;s 94.8 at a million [6][18]. Telling eight similar things apart does not. On the eight-needle test, Alibaba&#8217;s new model scores 96 at 128,000 tokens, 93 at 256,000, 41 at 512,000 and 26 at a million [18]. Nothing in that table blames the sparse layer: with full attention put back in the fourth layer, the scores at 512,000 and a million are 31 and 21. So the million-token window is real as capacity, and RULER holds across it in these two tables. Distinguishing eight similar things requires between 256,000 and 512,000 tokens, with either fourth layer.</p><p style="text-align: justify;">An audit in May looked for seven standard long-context benchmarks in the main results tables of four releases (DeepSeek-V4-Pro, MiMo-V2.5-Pro, Kimi-K2.6 and GLM-5.1): 0 of 28 were there, against 20 of 28 for agent and coding benchmarks [21]. DeepSeek&#8217;s Table 6 does print 83.5 for MRCR, OpenAI&#8217;s multi-needle test, at a million tokens. Its Figure 9 gives 0.59 for the eight-needle version at the same length, and the report does not say how the two relate [3].</p><h2><strong>Which chips run these designs fast</strong></h2><p style="text-align: justify;">The first models shipped on kernels, the low-level code that runs a layer fast on a given chip, that the labs wrote or adapted from open-source libraries. Z.AI&#8217;s GLM-5 report has a section on adapting the model to seven Chinese chip platforms [15].</p><p style="text-align: justify;">Mamba-2&#8217;s authors intentionally made it less expressive than the first Mamba so it &#8220;can exploit specialized matrix multiplication (matmul) units on GPUs, also known as tensor cores&#8221; [11]. Gated DeltaNet, the layer in Qwen and the ancestor of Kimi&#8217;s, has three authors: one from MIT and two from NVIDIA, with a footnote saying the work was done during the first author&#8217;s internship there [12]. Its training algorithm adapts DeltaNet&#8217;s, which the paper says &#8220;leverages matmuls, enabling tensor core-based GPU optimization&#8221; [12]. Neither paper measured serving.</p><p style="text-align: justify;">Qwen3-Next shipped in September 2025 on kernels adapted from an open-source library. An NVIDIA engineer&#8217;s kernel for the layer, for the Hopper generation and for prompt processing only, arrived 114 days later [22]. Moonshot published its tuned kernel 173 days after Kimi Linear, with a first benchmark on NVIDIA&#8217;s H20 and a second on Blackwell a month later [22]. vLLM added AMD GPU support for DeepSeek&#8217;s sparse attention 52 days after launch, with accuracy numbers and no speed numbers [22]. Each trailed the others by 52 to 173 days, and none was the first to support it: DeepSeek, for one, published its own kernels with the model [22]. By this summer the gap had closed for Qwen3.8 at least: vLLM wrote that &#8220;NVIDIA and Inferact co-developed ultra-fast kernels for Linear Attention (Gated Delta Rule)&#8221; for it, with no speed number [22].</p><p style="text-align: justify;">DeepSeek&#8217;s V4 report offers &#8220;some proposals to hardware vendors, in the hope of aiding efficient hardware design and achieving better software-hardware co-design&#8221; [23]. They concern the kernels of the expert layers (the ratio of compute to communication, power, communication primitives, the activation function), not attention. On September 30th, 2026, its FlashMLA kernel library added Huawei&#8217;s Ascend 950 and announced, &#8220;we removed support for the Hopper architecture&#8221; [23]. Hopper is the generation of the H100, the H200, and the H20. The README still lists NVIDIA first, as &#8220;NVIDIA SM100 / SM103 GPU&#8221;, i.e., Blackwell, and tells Hopper owners to &#8220;switch to&#8221; an older version of the code. If you own Hopper, the current release is not written for you.</p><p style="text-align: justify;">So, co-designed or not? In the maths, yes: the linear layers were designed to train fast on GPU matrix units, and one was written partly at NVIDIA. In the serving code, mostly not: tuned kernels trailed the models by weeks or months, until Qwen3.8.</p><p style="text-align: justify;">One company bet a chip on attention staying put. In June 2024, Etched announced a chip that ran only Transformers and wrote: &#8220;If transformers are replaced by SSMs, RWKV, or any new architecture, our chips will be useless.&#8221; [24] On October 1st, 2026, its homepage read &#8220;We&#8217;re building a new category of AI hardware: frontier inference clusters,&#8221; and it didn't mention &#8220;transformer&#8221; [24]. Four of the eight steps above first shipped between September 2025 and August 2026, and Qwen3.5, a 35-billion-parameter model from February 2026, already ships three linear layers for each full-attention one [2].</p><h2><strong>Where the cost reaches you</strong></h2><p style="text-align: justify;">Three places: cache, retrieval, and hardware.</p><p style="text-align: justify;">The first is the cache you pay for. When a provider sells &#8220;cached input&#8221;, it keeps the prefix it has already processed [25]. With full attention, that is the KV cache, one entry per token, so it can reuse your prompt up to any token. With linear layers, it saves a copy of the state at one position and can reuse it only at a position where someone saved one. A hybrid needs both, aligned to the same position. Saved states are big: on Qwen3.5-4B, one is about 49 MiB, as much as the KV cache for 1,500 tokens [26]. Moonshot&#8217;s own report says of caching these layers in fixed blocks that &#8220;At such a coarse granularity caching is nearly useless&#8221;, then describes the fix it built [26].</p><p style="text-align: justify;">vLLM caught up fast. It shipped Kimi K3 support on July 27th, 2026, with caching switched off by default. A month later, its release notes read &#8220;prefix caching enabled by default for Mamba models&#8221; [27]. That line doesn't name K3, so check your version. Two things have not caught up. On vLLM today, a cached answer from a Gated DeltaNet model is not guaranteed to match an uncached one bit for bit, so you cannot diff outputs in a regression test: vLLM&#8217;s batch-invariant mode does not cover Gated DeltaNet yet, and LMCache&#8217;s documentation says &#8220;Generation is not bit-exact between a cached and a fresh run&#8221; and tells you to compare benchmark scores instead [27]. The only hit rate I found for K3 is Moonshot&#8217;s own; neither the vLLM post, LMCache&#8217;s documentation, nor the SuffixReplay paper gives one. Moonshot sells K3 cache hits at $0.30 per million tokens against $3.00, and says its API &#8220;achieves a cache hit rate above 90% in coding workloads&#8221;, with no method [7]. vLLM&#8217;s headline K3 benchmark, 2.8 times its earlier throughput at 16 concurrent requests, ran with caching off because of a bug since fixed [27]. The discount is real: all 24 K3 endpoints on OpenRouter price cached input below fresh input, 21 of them at a fifth of it or less [7]. Your hit rate is the number you should measure yourself by.</p><p>The second is retrieval at your depth: item 2 of the checklist below.</p><p style="text-align: justify;">The third is the hardware you own. Somebody tried Kimi K3 on A100s, NVIDIA&#8217;s Ampere generation, the week it shipped, and got &#8220;no kernel image is available for execution on the device.&#8221; A member of the vLLM project replied: &#8220;K3 doesn&#8217;t support Ampere yet.&#8221; [28] Memory sizing changes too. Each request now carries state for its linear layers on top of its cache, and one operator who asked SGLang for 32 concurrent requests found it &#8220;capped to 9 by the mamba state cache&#8221; [28]. That was working as designed, and the log showed it. Old tuning tricks need retesting: storing the cache in 8 bits halves it on a full-attention model, and one vLLM calculation for a small hybrid gave 1.84 times the capacity at 32,768 tokens and 1.00 at short context, where the state, which is not quantized, sets the limit [28].</p><h2><strong>Six checks before you build on one</strong></h2><ol><li><p>Open <code>config.json</code>. Count the layer types, then read the keys next to anything called full attention.</p></li><li><p style="text-align: justify;">Run the same test at 32,000 tokens and at your own context length, with the target at several positions, including the middle, on counting, reranking, and multi-step questions. The drop between the two scores is what length costs you on that model. Needles are the easy case.</p></li><li><p>Measure cache hit rate on your engine, your version, and your prompts. Compare scores, not outputs.</p></li><li><p>Size concurrency against state memory as well as cache memory, and retest every tuning you carried over.</p></li><li><p>Ask which GPU generation has a fast path for this model in this release.</p></li><li><p style="text-align: justify;">Price the workload two ways: this model at the hit rate you measured, and a full-attention model you could use instead. The savings are real where the provider has passed them on.</p></li></ol><p style="text-align: justify;">The window got longer, inference got more expensive, and layers got smarter. Yet testing didn't keep up, and what we measure is narrower. </p><p style="text-align: justify;">Offered at a million tokens. Compared with an all-full-attention model at 128,000, at most. Your job is to explore what happens in between, and find the accuracy price you&#8217;re paying to optimize inference.</p><h3><strong>Notes</strong></h3><p>[1] Qwen3.8-Flash-Next on Hugging Face, accessed October 1st, 2026. The <a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next/raw/main/config.json">config.json</a> lists 48 layers under <code>layer_types</code>: 36 <code>linear_attention</code> and 12 <code>full_attention</code>, with <code>indexer_budget</code> 2048 and <code>indexer_compress_ratio</code> 4. The <a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next">model card</a> describes the same 12 layers as sparse: &#8220;12 &#215; (3 &#215; (Gated DeltaNet &#8594; MoE) &#8594; 1 &#215; (Qwen Sparse Attention &#8594; MoE))&#8221;, &#8220;Budget: 512 blocks or 2048 tokens&#8221;. Both documents are Alibaba&#8217;s own; the claim is about what its file says.</p><p>[2] Config files and model cards on Hugging Face, accessed October 1st, 2026. <a href="https://huggingface.co/moonshotai/Kimi-K3">Kimi K3</a>: &#8220;Attention-Layer Composition 69 KDA + 24 Gated MLA&#8221;, 93 layers; the <a href="https://huggingface.co/moonshotai/Kimi-K3/raw/main/config.json">config</a> lists 69 <code>kda_layers</code> and 24 <code>full_attn_layers</code>. <a href="https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B">Qwen3.8-2.4T-A95B</a>: 92 layers, 69 <code>linear_attention</code> and 23 <code>full_attention</code>. <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/raw/main/config.json">DeepSeek V4-Flash</a>: 43 layers, <code>compress_ratios</code> of 0, 4 or 128 for every layer; V4-Pro has 61. <a href="https://huggingface.co/zai-org/GLM-5.3-Flash/raw/main/config.json">GLM-5.3-Flash</a>: 45 layers, 34 <code>linear_attention</code> and 11 <code>deepseek_sparse_attention</code>, the 11 listed under <code>full_attn_layers</code>, with <code>index_topk</code> 2048. Qwen3.8-Flash-Next is in note 1. <a href="https://huggingface.co/Qwen/Qwen3.5-35B-A3B/raw/main/config.json">Qwen3.5-35B-A3B</a>: 40 layers, 30 <code>linear_attention</code> and 10 <code>full_attention</code>; the <a href="https://huggingface.co/api/models/Qwen/Qwen3.5-35B-A3B">repository record</a> gives its creation date as February 24th, 2026. <a href="https://huggingface.co/zai-org/GLM-5.3/raw/main/config.json">GLM-5.3</a>: 78 layers, architecture <code>GlmMoeDsaForCausalLM</code>. <a href="https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct">Qwen3-Next</a>: 48 layers, &#8220;12 * (3 * (Gated DeltaNet -&gt; MoE) -&gt; 1 * (Gated Attention -&gt; MoE))&#8221;. <a href="https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct/raw/main/config.json">Kimi Linear</a>: 27 layers, 20 KDA and 7 full attention. Author&#8217;s calculation from the Kimi K3 config: 24 &#247; 93 &#8776; 0.258. Independent confirmation of the pattern: the survey in note 5, and LMCache&#8217;s <a href="https://docs.lmcache.ai/mp/hybrid_models.html">list of hybrid models</a>, which gives 12 validated hybrid recipes against 6 for uniform attention. In the <a href="https://huggingface.co/zai-org/GLM-5.3/raw/main/config.json">GLM-5.3 config</a>, <code>indexer_types</code> reads <code>shared</code> for 57 of the 78 layers and <code>full</code> for 21. vLLM&#8217;s launch post (note 27) gives the size: &#8220;Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model&#8221;. Context windows, from the model cards: Kimi K3 &#8220;supports a 1-million-token context window&#8221;; Qwen3.8-2.4T-A95B, &#8220;Context Length: 262,144 natively and extensible up to 1,010,000 tokens&#8221;. The GLM-5.3-Flash config also sets <code>index_kpool</code> to 4; neither the config nor the model card says whether its 2,048 counts tokens or pooled groups of four. Qwen3.8-Flash-Next and GLM-5.3-Flash both received their weights on August 26th, 2026, by their repositories&#8217; commit histories (read again October 3rd, 2026); the GLM-4.7-Flash checkpoint holds 31.2 billion parameters, by the Hub&#8217;s count.</p><p>[3] DeepSeek-AI, &#8220;<a href="https://arxiv.org/abs/2606.19348">DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence</a>&#8220;, launched April 24th, 2026; the arXiv copy is dated April 26th, 2026. Abstract: &#8220;V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache&#8221; of V3.2; section 1 gives these for &#8220;the scenario of 1M-token context&#8221;, from &#8220;estimated&#8221; FLOPs, and gives V4-Flash, a smaller model with fewer layers, as 10% and 7%. Section 4.2.1: &#8220;For CSA, we set the compression rate m to 4&#8221;; &#8220;For HCA, we set the compression rate m&#8242; to 128&#8221;; &#8220;the window size n_win is set to 128&#8221;. The text contains no match for &#8220;ablat&#8221;. Figure 9, V4-Pro-Max on eight-needle MRCR: 0.92 at 128K and 0.59 at 1024K; Table 6 gives &#8220;MRCR 1M&#8221; as 83.5 without saying how the two relate. Moonshot&#8217;s framing is in the Kimi Linear paper, note 6. Section 4.2.1, on V4-Flash: &#8220;For the first two layers, we use pure sliding window attention.&#8221; In the layers with a 4-to-1 ratio each compressed entry is built from 8 tokens and neighbouring entries overlap; the 128-to-1 layers do &#8220;not perform overlapped compression&#8221;. The V4-Flash config lists 21 layers at ratio 4, 20 at 128 and 2 at 0. A search of the report for &#8220;RULER&#8221; on October 1st, 2026 finds no match.</p><p>[4] Config files on Hugging Face, accessed October 1st, 2026: <a href="https://huggingface.co/mistralai/Mistral-Medium-3.5-128B">Mistral Medium 3.5-128B</a> (88 layers, no sliding window); <a href="https://huggingface.co/ibm-granite/granite-4.2-30b">Granite 4.2-30b</a>(&#8221;Decoder-only Dense Transformer&#8221;); <a href="https://huggingface.co/IFM/K2-Horizon-375B-A23B">K2-Horizon-375B-A23B</a> (61 layers, <code>use_sliding_window</code> false); <a href="https://huggingface.co/openai/gpt-oss-120b">gpt-oss-120b</a> (18 sliding-window and 18 full-attention layers, window 128); <a href="https://huggingface.co/google/gemma-4-31B-it">Gemma 4 31B</a> (50 and 10, window 1,024); <a href="https://huggingface.co/thinkingmachines/Inkling">Inkling</a> (55 and 11, window 512); <a href="https://huggingface.co/meta-models/Muse-Glimmer-30B">Muse Glimmer 30B</a> (39 and 13, window 2,048); <a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16">Nemotron 3 Ultra</a> (48 Mamba-2, 48 MoE and 12 attention blocks); IBM&#8217;s <a href="https://huggingface.co/ibm-granite/granite-4.0-h-small">Granite 4.0-H-Small</a>card: &#8220;GQA, Mamba2, MoEs with shared experts&#8221;. <a href="https://huggingface.co/allenai/Olmo-Hybrid-7B">Olmo Hybrid 7B</a>: &#8220;75% of layers use gated DeltaNet heads instead of attention heads.&#8221; Microsoft and Ai2 also publish fine-tunes of Qwen models that carry the linear layers.</p><p>[5] Tan et al., &#8220;<a href="https://arxiv.org/abs/2609.39661">The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends</a>&#8220;, arXiv 2609.39661, September 30th, 2026; all six authors are at Alibaba Token Hub. &#8220;Table 17 shows that Hybrid records increase from none in 2022&#8211;2023 to 9 of 25 in 2024&#8211;2025 and 17 of 27 in 2026.&#8221; The survey counts sliding-window mixes as hybrid and calls its inventory &#8220;purposively curated rather than exhaustive&#8221;. Of its eleven frontier models: &#8220;Hybrid structures are prominent but not universal: K2-Horizon retains full GQA and GLM-5.3 uses a single DSA family&#8221;. Its conclusion: &#8220;continued coexistence rather than a universal replacement path&#8221;.</p><p>[6] Moonshot AI, &#8220;<a href="https://arxiv.org/abs/2510.26692">Kimi Linear: An Expressive, Efficient Attention Architecture</a>&#8220;, submitted October 30th, 2025, 16:59 UTC. Table 1: a sweep of five hybrid ratios on a small model, with validation perplexity 5.65 at 3:1, 5.66 at 1:1 and 5.70 at 7:1. Section 5.4: the three models &#8220;share the same architecture, parameter count, and training setup for fair comparisons&#8221;, at 48 billion parameters, 3 billion active, 1.4 trillion training tokens. Table 5, RULER at 128K: full attention 81.3, Gated DeltaNet hybrid 80.5, Kimi Linear 84.3; full attention is ahead by more than a point on EvalPlus (62.6 against 61.0), LongBench V2 (36.1 against 35.0) and Frames (60.5 against 58.8), and by half a point on LiveBench (45.7 against 45.2) and Long Code Arena Commit (33.2 against 32.7); Kimi Linear with RoPE in its full-attention layers scores 78.8 on RULER. Abstract: &#8220;reducing KV cache usage by up to 75% and achieving up to 6&#215; decoding throughput for a 1M context&#8221;. Section 6.3: &#8220;a theoretical decoding speedup of up to 6.3&#215;&#8221;; at batch size 1, &#8220;a 2.3&#215; speedup at a 1M token context&#8221;. A later run of the released model: &#8220;Kimi Linear@5.7T obtains a score of 94.8 on RULER at 1M context length&#8221;, with no full-attention model tested at that length. Author&#8217;s calculation from the Kimi Linear config in note 2: 20 &#247; 27 &#8776; 0.741.</p><p>[7] DeepSeek, &#8220;<a href="https://api-docs.deepseek.com/news/news250929">DeepSeek-V3.2-Exp Release</a>&#8220;, September 29th, 2025: &#8220;DeepSeek API prices drop 50%+, effective immediately&#8221;; the price card reads &#8220;$0.56 &#8594; $0.28 cache miss&#8221; and &#8220;Output $1.68 &#8594; $0.42&#8221;. The archived pricing page shows <a href="https://web.archive.org/web/20250922205322/https://api-docs.deepseek.com/quick_start/pricing">$0.56 and $1.68 on September 22nd</a> and <a href="https://web.archive.org/web/20250929101628/https://api-docs.deepseek.com/quick_start/pricing">$0.28 and $0.42 on September 29th</a>. Kimi K3: Moonshot&#8217;s <a href="https://platform.kimi.ai/docs/pricing/chat">pricing page</a>, last modified September 28th, 2026, lists input $3.00, cached input $0.30 and output $15.00 per million tokens. Moonshot&#8217;s <a href="https://www.kimi.com/blog/kimi-k3">launch post</a>: &#8220;the official Kimi API achieves a cache hit rate above 90% in coding workloads.&#8221; The post gives no method or denominator for that figure, and the technical report gives no hit rate (both read October 1st, 2026). OpenRouter&#8217;s <a href="https://openrouter.ai/api/v1/models/moonshotai/kimi-k3/endpoints">endpoint list for Kimi K3</a>, read October 3rd, 2026: 24 endpoints, all with a cache-read price below the input price, 21 of them at 20% of it or less and 17 at 10%. DeepSeek&#8217;s <a href="https://api-docs.deepseek.com/quick_start/pricing">pricing page</a>, read October 1st, 2026: DeepSeek-V4-Pro-0813, input on a cache miss $1.32 per million tokens at peak and $0.66 off peak, output $3.96 and $1.98.</p><p>[8] Searched on October 1st, 2026 for &#8220;export&#8221;, &#8220;sanction&#8221;, &#8220;HBM&#8221;, &#8220;memory price&#8221; and &#8220;shortage&#8221; in the DeepSeek V4 report (note 3), the <a href="https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf">Kimi K3 technical report</a>, and the Kimi K3, GLM-5.3-Flash and Qwen3.8 model cards: no match in that sense. Z.AI, &#8220;<a href="https://z.ai/blog/glm-5.3-flash">GLM-5.3-Flash</a>&#8220;, August 26th, 2026: &#8220;with all of this traffic served on Chinese AI chips&#8221;; &#8220;These chips are primarily constrained by memory capacity and bandwidth, especially when supporting context lengths of up to one million tokens.&#8221;</p><p>[9] Julien Simon on YouTube: &#8220;<a href="https://www.youtube.com/watch?v=cl3MAhAr9-M">Decoder-only inference: a step-by-step deep dive</a>&#8220;, January 10th, 2025 (self-attention, the KV cache, multi-head and latent attention); &#8220;<a href="https://www.youtube.com/watch?v=2TT384U4vQg">Deep dive - Better Attention layers for Transformer models</a>&#8220;, February 12th, 2024 (grouped-query and sliding-window attention); &#8220;<a href="https://www.youtube.com/watch?v=hMs8VNRy5Ys">Deep Dive: Optimizing LLM inference</a>&#8220;, March 11th, 2024 (the KV cache in serving).</p><p>[10] DeepSeek-AI, &#8220;<a href="https://arxiv.org/abs/2405.04434">DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model</a>&#8220;, May 7th, 2024. Table 8: MMLU 41.2 with grouped-query attention in 8 groups against 45.2 with multi-head attention, on 7-billion-parameter models. Table 9: latent attention against multi-head attention on two model sizes, with a cache of &#8220;14% for small MoE models and 4% for large MoE models&#8221;; the one loss in eight cells is C-Eval on the small model. The cache per token is the 512-number latent plus a 64-number positional key. On inference: &#8220;we even do not need to compute keys and values out for attention.&#8221; Grouped-query attention: Ainslie et al., &#8220;<a href="https://arxiv.org/abs/2305.13245">GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints</a>&#8220;, submitted May 22nd, 2023.</p><p>[11] Tri Dao and Albert Gu, &#8220;<a href="https://arxiv.org/abs/2405.21060">Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality</a>&#8220;, May 31st, 2024. Section 10.1, comparing Mamba-2 with the first Mamba: &#8220;trades off this expressivity for improved hardware efficiency (and ease of implementation)&#8221;. Section 9.3: &#8220;can exploit specialized matrix multiplication (matmul) units on GPUs, also known as tensor cores.&#8221; Table 3: Transformer++ 60.2, Mamba-2 60.2, Mamba-2 with 6 attention layers 61.0. Section 9.2.3: &#8220;having around 10% of the total number of layers being attention performs best.&#8221; The hybrid comparison is run &#8220;at the 2.7B scale (64 layers)&#8221;.</p><p>[12] Songlin Yang, Jan Kautz and Ali Hatamizadeh, &#8220;<a href="https://arxiv.org/abs/2412.06464">Gated Delta Networks: Improving Mamba2 with Delta Rule</a>&#8220;, December 9th, 2024; affiliations MIT CSAIL, NVIDIA and NVIDIA; footnote: &#8220;Work done during SY&#8217;s internship at NVIDIA.&#8221; On DeltaNet: &#8220;a hardware-efficient chunkwise algorithm for DeltaNet that leverages matmuls, enabling tensor core based GPU optimization&#8221;; on Gated DeltaNet: &#8220;we can adapt DeltaNet&#8217;s chunkwise algorithm (Eq. 8-9) for Gated DeltaNet to enable hardware-efficient training&#8221;. Table 4, inputs truncated to 2,000 tokens: Gated DeltaNet 30.6, Transformer++ 37.0, the two hybrids 39.0 and 40.1; the first hybrid combines the layer with sliding-window attention; the second also stacks Mamba2 layers. Independent measurement of the same layer: Moonshot&#8217;s Table 5 (note 6), where the Gated DeltaNet hybrid scores 80.5 against full attention&#8217;s 81.3 at 128K. Z.AI&#8217;s <a href="https://arxiv.org/abs/2602.15763">GLM-5 report</a> measured a larger loss on a converted model and attributes it to &#8220;the unavoidable information loss introduced by efficient attention mechanisms during continual-training adaptation&#8221;. Table 4, &#8220;input truncated to 2K tokens&#8221;, averages six tasks (SWDE, SQuAD, FDA, TriviaQA, NQ and Drop); the gap is widest on FDA, 23.7 against 52.2, and the pure layer leads on TriviaQA, 60.0 against 58.3.</p><p>[13] DeepSeek-AI, &#8220;<a href="https://arxiv.org/abs/2512.02556">DeepSeek-V3.2</a>&#8220;, December 2nd, 2025: &#8220;select 2048 key-value tokens for each query token&#8221;; the paper claims that sparse attention &#8220;reduces the core attention complexity of the main model from O(L&#178;) to O(Lk)&#8221; and makes no claim about cache size. The <a href="https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp">V3.2-Exp model card</a>, released September 29th, 2025, prints 14 benchmarks against V3.1-Terminus; the count of 7 higher, 6 lower and 1 level is mine. The sparse model was built by continued training from the V3.1-Terminus checkpoint : 2.1 billion tokens to warm up the indexer, then 943.7 billion tokens of sparse training. On the indexer: &#8220;Although the lightning indexer still has a complexity of O(L&#178;), it requires much less computation compared with MLA in DeepSeek-V3.1-Terminus.&#8221;</p><p>[14] MiniMax, &#8220;<a href="https://huggingface.co/blog/MiniMax-AI/why-did-m2-end-up-as-a-full-attention-model">Why Did MiniMax M2 End Up as a Full Attention Model?</a>&#8220;, Hugging Face, October 30th, 2025. MiniMax, &#8220;<a href="https://arxiv.org/abs/2605.26494">The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence</a>&#8220;, arXiv 2605.26494, May 26th, 2026, Table 2, full attention against hybrid sliding-window attention: RULER 32K CWE 99.0 and 99.0; RULER 128K CWE 90.0 and 72.0; MMLU 85.5 and 85.6. CWE is RULER&#8217;s common-words-extraction task, an aggregation task. The report does not state the size of the two models compared. &#8220;we found no variant that reliably matches full attention quality in production settings spanning reasoning, coding, and agent tasks.&#8221; The sliding-window models were conversions: &#8220;continuing pre-training for hundreds of billions to trillions of tokens across multiple configurations&#8221;. The report cites sliding-window attention to &#8220;Beltagy et al., 2020&#8221;. The blog post: &#8220;efficient attention still has some way to go before it can definitively beat full attention&#8221;. The report describes the comparison as being &#8220;at the M2 architecture scale&#8221;.</p><p>[15] Zhipu AI and Tsinghua University, &#8220;<a href="https://arxiv.org/abs/2602.15763">GLM-5: from Vibe Coding to Agentic Engineering</a>&#8220;, arXiv 2602.15763, February 17th, 2026. Table 5, RULER at 128K on a 9-billion-parameter model: full attention 75.28, Gated DeltaNet variant 64.00. &#8220;Nevertheless, all of these methods incur an inherent accuracy gap on fine-grained retrieval tasks&#8212;up to 5.69 points on RULER@128K and 7.33 on RepoQA@128K&#8212;due to the unavoidable information loss introduced by efficient attention mechanisms during continual-training adaptation, even when half of the layers retain full attention.&#8221; The report tests no linear layers trained from scratch (read October 1st, 2026). &#8220;DSA is lossless by construction&#8221;; Table 6, RULER at 128K: GLM-4.7-Flash 79.21, the same model with sparse attention 78.86. Table 3, &#8220;Comparison of long-context benchmarks between MLA and DSA base models&#8221;, at 128K, dense parent against sparse: MQ-NIAH 100.0 and 100.0; MV-NIAH 95.5 and 97.0; SQuAD 79.7 and 86.0; HotpotQA 66.3 and 63.0. The table does not label the models&#8217; size; the text says &#8220;The DSA training begins from the base model at the end of mid-training&#8221;, and the report puts GLM-5 at 744B parameters. Neither Z.AI nor MiniMax prints error bars on these long-context tables. Sparse attention&#8217;s largest gain in MiniMax&#8217;s Table 3 is on in-context learning, 72.80 against 70.40. MiniMax, &#8220;<a href="https://arxiv.org/abs/2606.13392">MiniMax Sparse Attention</a>&#8220;, arXiv 2606.13392, June 11th, 2026, Table 3 at 128K, full attention against the same checkpoint converted to sparse attention (MSA-CPT): RULER 72.00 and 72.12; reranking 34.60 and 32.50; word-frequency tasks 46.35 and 45.00; question answering 47.80 and 46.80; variable tracking 97.80 and 96.80; needle retrieval 96.63 and 98.87. Conclusion: &#8220;closing the residual long-context retrieval gap&#8221;. The MiniMax test model has &#8220;approximately 109B total parameters and 6B activated parameters per token&#8221;. GLM-5, section 5, &#8220;Adapting GLM-5 to Chinese Chip Infrastructure&#8221;: &#8220;seven mainstream Chinese chip platforms, including Huawei Ascend, Moore Threads, Hygon, Cambricon, Kunlunxin, MetaX, and Enflame&#8221;. In Table 5 the linear variant converts half the layers (&#8221;even when half of the layers retain full attention&#8221;). MiniMax presents the sparse design as the one behind MiniMax-M3. Table 5 is built on a 9-billion-parameter model: &#8220;We continually train each method on 190B tokens with a 64K context length, maintaining a 1:1 ratio&#8221;. Its Gated DeltaNet row reads 64.00 against 75.28, a gap of 11.28; the &#8220;up to 5.69 points&#8221; in the quoted sentence is the gap of the best sliding-window variant.</p><p>[16] Zhu et al., &#8220;<a href="https://arxiv.org/abs/2609.36322">Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression</a>&#8220;, ByteDance Seed, September 28th, 2026: on DeepSeek-V4.1-Flash at 128,000 tokens, &#8220;a mean accuracy of 92.38% and a best&#8211;worst residue gap of 6.09 percentage points.&#8221; The paper&#8217;s 40-point gap is on a base checkpoint, not on the released model. DeepSeek-AI, &#8220;<a href="https://arxiv.org/abs/2609.19969">DeepSeek-V4.1-Flash</a>&#8220;, September 17th, 2026: &#8220;the newly introduced architectural changes also create robustness boundaries that have yet to be fully characterized.&#8221; Zhu et al., Table 7, DeepSeek-V4.1-Flash by residue: even 95.00, 95.12, 94.69 and 95.31; odd 89.22, 89.69, 90.31 and 89.73; best minus worst 6.09. Author&#8217;s calculation: (95.00 + 95.12 + 94.69 + 95.31) &#247; 4 &#8776; 95.03. Author&#8217;s calculation: (89.22 + 89.69 + 90.31 + 89.73) &#247; 4 &#8776; 89.74. Zhu et al. on positions: &#8220;all four of its even residue groups remain above all four odd ones&#8221;.</p><p>[17] Z.AI, &#8220;<a href="https://z.ai/blog/glm-5.3-flash">GLM-5.3-Flash: Frontier Intelligence, Flash Cost</a>&#8220;, August 26th, 2026: &#8220;we compare the per-token compute and KV cache size of GLM-5.3-Flash against GLM-5.3 and two recent open models: DeepSeek-V4-Flash and Kimi-K3.&#8221; Searched October 1st, 2026: the post contains &#8220;Kimi&#8221; twice and &#8220;DeepSeek&#8221; six times, as comparison models, and &#8220;KDA&#8221; and &#8220;Moonshot&#8221; not at all; it contains &#8220;preserving precise long-context capabilities&#8221; and none of &#8220;RULER&#8221;, &#8220;MRCR&#8221;, &#8220;LongBench&#8221;, &#8220;HELMET&#8221;, &#8220;RepoQA&#8221; or &#8220;needle&#8221;; the model card and API docs were read for the same terms as captured on October 1st. The <a href="https://huggingface.co/zai-org/GLM-5.3-Flash/raw/main/config.json">config</a> lists its linear layers as <code>kda_layers</code> and its sparse layers as <code>deepseek_sparse_attention</code>. Moonshot&#8217;s <a href="https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct">Kimi Linear</a> and <a href="https://github.com/MoonshotAI/FlashKDA">FlashKDA</a> and DeepSeek&#8217;s <a href="https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp">V3.2-Exp</a> are all published under the MIT licence.</p><p>[18] Alibaba, &#8220;<a href="https://raw.githubusercontent.com/QwenLM/Qwen3.8-Flash-Next/main/tech_report.pdf">On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability</a>&#8220;, August 26th, 2026. Table 1, three 28-layer, 25-billion-parameter models with 3 billion active, on nine benchmarks: full attention 49.87, sliding-window hybrid 51.15, Gated DeltaNet hybrid 53.81; &#8220;The GDN hybrid improves over the Transformer on eight of the nine selected benchmarks&#8221;; the ninth is EvalPlus, 49.71 against 51.01 for full attention; &#8220;they do not by themselves isolate which architectural component causes each improvement.&#8221; Table 3, the shipped model with its fourth layer as full attention against sparse attention: eight-needle MRCR at 128K 97.14 and 95.98; at 256K 94.20 and 93.00; at 512K 30.66 and 40.53; at 1M 20.71 and 26.44; RULER from 512K to 1M 90.08 and 93.00. The eight-needle test is OpenAI&#8217;s <a href="https://huggingface.co/datasets/openai/mrcr">MRCR</a>: &#8220;2, 4, or 8 identical asks&#8221;, and &#8220;the model is ultimately prompted to return the i-th instance&#8221;. RULER: Hsieh et al., &#8220;<a href="https://arxiv.org/abs/2404.06654">RULER: What&#8217;s the Real Context Size of Your Long-Context Language Models?</a>&#8220;, 13 tasks, with &#8220;new task categories multi-hop tracing and aggregation to test behaviors beyond searching from context&#8221;. The shipped model: &#8220;Qwen3.8-Flash-Next is a sparse mixture-of-experts model with 125B total parameters, 6B activated per token&#8221;. Table 3, &#8220;Full Attn&#8221; against &#8220;w/ QSA&#8221; on the same model. RULER: 99.84 and 99.89 up to 128K, 99.81 and 99.62 to 256K, 97.65 and 98.95 to 512K, 90.08 and 93.00 to 1M. Eight-needle MRCR: 97.14 and 95.98 at 128K, 94.20 and 93.00 at 256K, 30.66 and 40.53 at 512K, 20.71 and 26.44 at 1M.</p><p>[19] DeepSeek, &#8220;<a href="https://api-docs.deepseek.com/news/news260424">DeepSeek-V4 Preview Release</a>&#8220;, April 24th, 2026: &#8220;Novel Attention: Token-wise compression + DSA (DeepSeek Sparse Attention).&#8221; Moonshot&#8217;s <a href="https://www.kimi.com/blog/kimi-k3">launch post</a> names Kimi Delta Attention.</p><p>[20] The ten comparisons are in notes 14 (MiniMax, sliding windows), 6 (Moonshot), 13 (DeepSeek), 15 (Z.AI three times, and MiniMax, sparse) and 18 (Alibaba, two), plus Ai2: Merrill et al., &#8220;<a href="https://arxiv.org/abs/2604.03444">Olmo Hybrid</a>&#8220;, arXiv 2604.03444, submitted April 3rd, 2026, and announced on Ai2&#8217;s blog on <a href="https://allenai.org/blog/olmohybrid">March 5th, 2026</a>. Table 3, RULER at 4K and at 64K: Olmo 3 with YaRN 95.8 and 70.9, Olmo Hybrid with YaRN 92.8 and 76.9, with DroPE 92.2 and 85.0; the baseline already used sliding-window attention in 75% of its layers (&#8221;we replace the sliding-window attention (SWA) layers (75% of layers) from Olmo 3&#8221;). Ai2&#8217;s own framing: &#8220;a controlled comparison between transformer and modern hybrid LMs is lacking at large scale.&#8221; Ai2 published the weights, code, training logs and data. Searched October 1st, 2026: the <a href="https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf">Kimi K3 technical report</a> contains no comparison with a full-attention model (&#8221;ablat&#8221; matches twice, on the vision tower and on data sampling; &#8220;RULER&#8221; and &#8220;MRCR&#8221; do not appear); the DeepSeek V4 report has no match for &#8220;ablat&#8221; (note 3); the Qwen3.8-2.4T model card has no comparison or ablation; Z.AI&#8217;s GLM-5.3-Flash post has none (note 17).</p><p>[21] Zhang et al., &#8220;<a href="https://arxiv.org/abs/2605.23170">Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks</a>&#8220;, Beijing Jiaotong University, May 22nd, 2026: an audit of DeepSeek-V4-Pro, MiMo-V2.5-Pro, Kimi-K2.6 and GLM-5.1 against 19 benchmarks. Seven long-context benchmarks (NIAH, RULER, LongBench, HELMET, InfiniteBench, BABILong, LOFT) across four releases: 0 of 28 in a main results table. Seven agent and coding benchmarks: 20 of 28. MRCR, which DeepSeek&#8217;s Table 6 does print, is not among the seven.</p><p>[22] Dates from GitHub, accessed October 1st, 2026; the day counts are mine, from Qwen3-Next&#8217;s release on September 11th, 2025 (weights uploaded September 9th, card and config published September 11th, by the Hub&#8217;s commit history), Kimi Linear&#8217;s on October 30th, 2025 (to FlashKDA&#8217;s first commit on April 21st, 2026) and DeepSeek-V3.2-Exp&#8217;s on September 29th, 2025. DeepSeek&#8217;s <a href="https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp">V3.2-Exp model card</a> has a section &#8220;Open-Source Kernels&#8221; pointing to its own CUDA kernels. vLLM <a href="https://github.com/vllm-project/vllm/pull/24518">#24518</a>, September 2025: the kernels &#8220;are adapted from&#8221; the open-source <a href="https://github.com/fla-org/flash-linear-attention">flash-linear-attention</a> library. FlashInfer <a href="https://github.com/flashinfer-ai/flashinfer/pull/2276">#2276</a>, by an NVIDIA engineer, merged January 3rd, 2026: &#8220;implementation for Gated Delta Rule (or Gated Delta Net) on Hopper architecture&#8221;; it covers prefill, and a decode kernel came in a later pull request. <a href="https://github.com/MoonshotAI/FlashKDA">MoonshotAI/FlashKDA</a>, created April 20th, 2026; its H20 benchmark came with the first commit, on April 21st, and was updated April 22nd; its GB200 benchmark was added May 26th (GitHub API, read October 3rd, 2026). vLLM <a href="https://github.com/vllm-project/vllm/pull/26670">#26670</a>, merged November 20th, 2025: &#8220;The PR add Deepseek v3.2 support on ROCm platforms.&#8221;; it reports an accuracy score and no throughput. vLLM, &#8220;<a href="https://vllm.ai/blog/2026-08-12-qwen3.8">Day 0 Support for Qwen3.8-2.4T-A95B on vLLM</a>&#8220;, August 12th, 2026. vLLM on Qwen3.8: &#8220;NVIDIA and Inferact co-developed ultra-fast kernels for Linear Attention (Gated Delta Rule)&#8221;.</p><p>[23] DeepSeek V4 report (note 3), section 3.1. DeepSeek, <a href="https://github.com/deepseek-ai/FlashMLA/blob/main/README.md">FlashMLA README</a>, commit of September 30th, 2026: &#8220;In the 2026.09.30 release, we removed support for the Hopper architecture and for earlier models (including DeepSeek V3 / V3.2 / V4.0)&#8221;; &#8220;We&#8217;ve released sparse attention prefill and decoding kernels for the Huawei Ascend 950 NPU&#8221;; requirements &#8220;NVIDIA SM100 / SM103 GPU&#8221;; Hopper users are told to &#8220;switch to&#8221; an earlier commit. The V4 report says only that one kernel scheme was validated &#8220;on both NVIDIA GPUs and HUAWEI Ascend NPUs platforms&#8221;. The report&#8217;s words: &#8220;some proposals to hardware vendors, in the hope of aiding efficient hardware design and achieving better software-hardware co-design&#8221;; the first proposal is headed &#8220;Computation-Communication Ratio&#8221;. Read October 1st, 2026. NVIDIA lists the H20 among its <a href="https://docs.nvidia.com/ai-enterprise/release-7/latest/infra-software/vgpu/reference/hopper.html">Hopper Architecture vGPU Types</a>, and its <a href="https://developer.nvidia.com/cuda-gpus">CUDA GPUs</a> page gives compute capability 10.0 for the B200 and GB200 and 10.3 for the B300 and GB300, the Blackwell parts (both read October 3rd, 2026).</p><p>[24] Etched, &#8220;<a href="https://web.archive.org/web/20240625142106/https://www.etched.com/announcing-etched">Announcing Etched</a>&#8220;, June 25th, 2024, archived the same day. <a href="https://www.etched.com/">etched.com</a>, read October 1st, 2026: &#8220;We&#8217;re building a new category of AI hardware: frontier inference clusters.&#8221;; the page contains neither &#8220;transformer&#8221; nor &#8220;Sohu&#8221;. &#8220;<a href="https://www.airealist.ai/p/the-model-is-the-machine">The Model Is the Machine</a>&#8220;, The AI Realist, August 9th, 2026. Taalas describes its product as &#8220;<a href="https://taalas.com/the-path-to-ubiquitous-ai/">a hard-wired Llama 3.1 8B</a>&#8220;; AMD <a href="https://newsroom.amd.com/news/amd-acquires-taalas-ai-inference/">announced an agreement</a> to acquire Taalas on August 6th, 2026. Etched, June 2024: &#8220;If transformers are replaced by SSMs, RWKV, or any new architecture, our chips will be useless.&#8221;</p><p>[25] &#8220;<a href="https://www.airealist.ai/p/the-cache-is-the-price">The Cache Is the Price</a>&#8220;, The AI Realist, September 8th, 2026.</p><p>[26] Liu et al., &#8220;<a href="https://arxiv.org/abs/2609.33477">Just Let Linear States Forget the Distant Past: Prefix Caching via SuffixReplay for Hybrid LLMs</a>&#8220;, arXiv 2609.33477, China Telecom, September 27th, 2026: &#8220;in Qwen3.5-4B, one linear-attention state checkpoint is approximately 49 MiB, whereas the KV cache for one token is only about 32 KiB&#8221;. <a href="https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf">Kimi K3 technical report</a>, section 5.4.1. 49 MiB is 1,568 times 32 KiB. Moonshot&#8217;s Kimi K3 technical report: &#8220;At such a coarse granularity caching is nearly useless: requests shorter than one block can never be reused&#8221;.</p><p>[27] vLLM, &#8220;<a href="https://vllm.ai/blog/2026-07-27-k3">Kimi K3 Is Here</a>&#8220;, July 27th, 2026: &#8220;Prefix caching is typically enabled by default in vLLM, but it is currently disabled by default for Kimi K3 while the hybrid-cache design continues to evolve.&#8221; The <a href="https://github.com/vllm-project/vllm/releases/tag/v0.28.0">v0.28.0 release notes</a>, August 26th, 2026: &#8220;prefix caching enabled by default for Mamba models&#8221;. vLLM, &#8220;<a href="https://vllm.ai/blog/2026-09-13-kimi-k3-performance-optimization">Kimi K3 Performance Optimizations in vLLM</a>&#8220;, September 13th, 2026: &#8220;Prefix caching was disabled in both runs because v0.27.1 had a known Kimi K3 prefix-caching issue fixed later&#8221;; the 2.8 times is the concurrency-16 row, 258.3 to 725.0 tokens per second. LMCache, &#8220;<a href="https://docs.lmcache.ai/mp/hybrid_models.html">Hybrid models</a>&#8220;, read October 1st, 2026: &#8220;Generation is not bit-exact between a cached and a fresh run: GDN backends do not support vLLM&#8217;s batch-invariant mode.&#8221; vLLM <a href="https://github.com/vllm-project/vllm/issues/42960">#42960</a>: &#8220;VLLM batch_invariant mode is not supported for GDN_ATTN.&#8221;; still the case in v0.30.0. No hit rate for Kimi K3 appears in either vLLM post, in LMCache&#8217;s documentation or in SuffixReplay (read October 1st, 2026). Independent confirmation that reuse is tied to saved positions: SuffixReplay and Moonshot&#8217;s report, both in note 26. The LMCache sentence in full: &#8220;Generation is not bit-exact between a cached and a fresh run: GDN backends do not support vLLM&#8217;s batch-invariant mode.&#8221;</p><p>[28] GitHub issues, read with their full threads on October 1st, 2026. vLLM <a href="https://github.com/vllm-project/vllm/issues/50249">#50249</a>, July 29th, 2026, Kimi K3 on A100. SGLang <a href="https://github.com/sgl-project/sglang/issues/38846">#38846</a>, September 10th, 2026, on v0.5.19: &#8220;max_running_requests is capped to 9 by the mamba state cache&#8221;. vLLM <a href="https://github.com/vllm-project/vllm/issues/55196">#55196</a>, September 3rd, 2026: one hybrid, Falcon-H1-1.5B, computed from vLLM&#8217;s capacity functions and not measured in a run; &#8220;1.84x at 32k context and 1.00x (no gain) at short context&#8221;. The cases left out include vLLM <a href="https://github.com/vllm-project/vllm/issues/42876">#42876</a>, closed in August 2026, and llama.cpp <a href="https://github.com/ggml-org/llama.cpp/issues/29092">#29092</a>, which reproduces only with the build bundled in Ollama. In vLLM #50249: &#8220;no kernel image is available for execution on the device&#8221;, and the maintainer&#8217;s reply, &#8220;K3 doesn&#8217;t support Ampere yet.&#8221;</p>]]></content:encoded></item><item><title><![CDATA[An Act of God, With Letterhead]]></title><description><![CDATA[Oracle reportedly called its power problem an act of God. Its annual report had already listed it.]]></description><link>https://www.airealist.ai/p/an-act-of-god-with-letterhead</link><guid isPermaLink="false">https://www.airealist.ai/p/an-act-of-god-with-letterhead</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Mon, 28 Sep 2026 16:04:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!86hv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!86hv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!86hv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!86hv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!86hv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!86hv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!86hv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1085560,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/217852706?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!86hv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!86hv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!86hv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!86hv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51c48c6e-bfcd-40a1-ab3b-7eaed5d13571_1456x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On September 24th, 2026, Bloomberg reported that Oracle had sent a force majeure notice to Stack Infrastructure, the Blue Owl company building Project Jupiter, a 2.45-gigawatt Stargate campus in the New Mexico desert [1]. Reuters had its own source and the same reason: &#8220;potential delays in securing power for the project&#8221; [1]. Force majeure is the clause companies invoke, in Reuters&#8217;s words, when &#8220;problems arise beyond their control&#8221;: war, flood, the occasional act of God. I read Oracle&#8217;s filings from this summer. Power was in the risk factors.</p><p style="text-align: justify;">Oracle told CNBC: &#8220;Project Jupiter remains on our planned schedule.&#8221; Blue Owl told CNBC: &#8220;This notice does not change the financial commitments to this multi-year project&#8221; [1]. So a notice nobody will describe, about a delay nobody admits, changes nothing. Officially.</p><h2><strong>God, in chronological order</strong></h2><p style="text-align: justify;">On <strong>March 20th</strong>, New Mexico&#8217;s land commissioner refused to let the campus&#8217;s gas pipeline cross state trust land [2]. On <strong>April 13th</strong>, federal staff protested the pipeline because Transwestern, its builder, had not filed a historic-preservation finding [2]. On <strong>April 27th</strong>, Oracle announced fuel cells for Jupiter and wrote that &#8220;construction continues to move forward on schedule&#8221;, adding that &#8220;Oracle will continue to bear all energy costs for Project Jupiter&#8221; [3]. On <strong>June 22nd</strong>, its annual report told investors: &#8220;<strong>We have faced, and may continue to face, challenges with securing reliable and cost-effective power sources</strong>&#8221; [4]. On <strong>July 14th</strong>, the commissioner said no again: &#8220;Advancing the massive use of gas for a project of this scale is simply not in the best interest of the trust&#8221; [2]. The finding had not been filed by the deadline, and on <strong>July 24th</strong> the federal regulator took the pipeline off its fast track, with other agencies&#8217; decisions under federal authority due by December 3rd [2]. On <strong>September 10th</strong>, asked about potential delays in New Mexico and Wisconsin, Oracle&#8217;s co-chief executive Clay Magouyrk said that a plan relying on &#8220;100% achievement of every one of their deliverables&#8221; is &#8220;called a bad plan&#8221;, and that the data center was &#8220;definitely on track&#8221; [5]. <strong>Two weeks later: force majeure.</strong></p><p style="text-align: justify;">A government saying no is the kind of event many force majeure clauses list, and nobody outside the deal has read this one. But Reuters&#8217;s source adds a detail: &#8220;securing power for the site is Oracle&#8217;s responsibility under the contract&#8221; [6]. If the event is the land office&#8217;s refusal, which no one has confirmed, then the act of God is a letter from a land commissioner about a job that the contract reportedly gives to Oracle, a risk Oracle disclosed in its annual report. God, it turns out, has letterhead.</p><h2><strong>What the notice buys</strong></h2><p style="text-align: justify;">According to the same Reuters source, Blue Owl has about $3 billion of equity in Jupiter, earns 9% on it while the campus is being built and expects about 11% once it is finished, and force majeure &#8220;extends the period during which it pays Blue Owl the lower development-stage rent&#8221; [6]. Oracle, by the same account, is responsible for the debt costs and &#8220;cannot terminate the lease under any circumstances&#8221; [6]. The two points of yield on Blue Owl&#8217;s equity come to about $60 million, deferred rather than lost, for each year of delay, or less than five hours of what Oracle spent on capex last quarter [6][7]. The act of God comes with a payment plan.</p><p style="text-align: justify;">If you lend to data centers, look at the $288 billion in Oracle leases that have not yet started, &#8220;generally expected to commence between the second quarter of fiscal 2027 and fiscal 2029 and for terms of fifteen to nineteen years&#8221; [7]. Nobody outside those deals has read their clauses. The notice shows how this tenant reads one. Two questions for the next lease: does the clause exclude risks the tenant has already disclosed, and who carries the power obligation?</p><h2><strong>We said so, mostly</strong></h2><p style="text-align: justify;">In &#8220;<a href="https://www.airealist.ai/p/cloud-vs-clout">Cloud vs. Clout</a>&#8221; we wrote that &#8220;Oracle is shifting capital risk to its counterparties&#8221; [8]. In &#8220;<a href="https://www.airealist.ai/p/welcome-to-hotel-abilene">Welcome to Hotel Abilene</a>&#8221; we argued that nobody in the Stargate web could check out. We had the direction wrong: the tenant is the one invoking the clause, and the landlord is the one waiting. In May, &#8220;<a href="https://www.airealist.ai/p/the-500-billion-umbrella">The $500 Billion Umbrella</a>&#8221; set a test for any Stargate deal: &#8220;a site under construction, a power purchase agreement in force, and a builder whose balance sheet can finish the job. Anything less is a press release&#8221; [8]. Jupiter has the site. Its power has a fuel-cell announcement from April 27th and a gas line the state refused and then refused to reconsider.</p><p style="text-align: justify;">On February 1st, Oracle said it did &#8220;not expect to issue additional bonds during calendar year 2026&#8221;, a plan that &#8220;reflects Oracle&#8217;s commitment to maintaining an investment-grade rating&#8221; [9]. The same document put about half the money into equity, and Oracle delivered $5 billion in mandatory convertible preferred that week and $19.9 billion in stock over the summer [10]. S&amp;P cut it to BBB-, one notch above junk, on July 9th [11]. The plan was kept. So was the investment-grade rating, by one notch.</p><h2><strong>The bill comes back</strong></h2><p>Oracle&#8217;s pricing sheets show what it paid over Treasuries at each bond sale since January 2025 [9]:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Bdgn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Bdgn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png 424w, https://substackcdn.com/image/fetch/$s_!Bdgn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png 848w, https://substackcdn.com/image/fetch/$s_!Bdgn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png 1272w, https://substackcdn.com/image/fetch/$s_!Bdgn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Bdgn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:311281,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/217852706?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Bdgn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png 424w, https://substackcdn.com/image/fetch/$s_!Bdgn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png 848w, https://substackcdn.com/image/fetch/$s_!Bdgn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png 1272w, https://substackcdn.com/image/fetch/$s_!Bdgn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd9f065a-22ec-40b4-8049-00db97f790b7_3200x1800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">Ten-year: 97 basis points to 145. Thirty-year: 122 to 180. That last bond, the 2056 note, has not held up: on September 28th, Bloomberg data reported by ZeroHedge showed it at 82.4 cents on the dollar and 273 basis points over Treasuries, compared with 180 at pricing [12]. Oracle&#8217;s five-year protection cost about 200 basis points in late July, compared with about 93 for Meta and about 53 for the investment-grade index [13], and reached a record of about 227 in the week of the notice [13]. In fairness, that was when the ten-year Treasury yield hit a 19-year high, and every hyperscaler&#8217;s spreads widened [13]. Oracle just widened from a higher floor.</p><p style="text-align: justify;">Has anyone climbed out of a spread like this? Boeing, whose default swaps hit 488 basis points in March 2020, a price Bloomberg read as 31% odds of default within five years, is still rated investment grade by all three agencies [14]. Others did not: Enron was downgraded to junk status four days before it filed for bankruptcy, and Lehman remained in the A category until it failed [14]. Oracle&#8217;s exposure is in its filings. My read: that is the kind of problem a company pays for rather than dies of, and the payment is the chart above.</p><p style="text-align: justify;">S&amp;P already adds Oracle&#8217;s $260 billion of lease commitments to the debt in its forecast, puts leverage &#8220;in the mid-4x area in fiscal 2027&#8221;, and says it could cut again if it expects Oracle &#8220;to sustain leverage exceeding 4.5x&#8221; [11]. The forecast sits at the edge of the trigger. </p><p style="text-align: justify;">Oracle&#8217;s own annual report says a downgrade could &#8220;affect the terms or availability of certain long-term commitments (including data center leases)&#8221; [15]. One more S&amp;P notch would not, by itself, push Oracle&#8217;s bonds out of the main index, which uses the middle of three ratings: Moody&#8217;s (Baa2, negative) and Fitch (BBB) each sat two notches above junk at Oracle&#8217;s last bond sale [16]. </p><p style="text-align: justify;">But every notch raises the price of the next dollar, and Oracle needs a lot of dollars, as do landlords like Jupiter&#8217;s, who borrow against its name. The Financial Times reports Jupiter&#8217;s loans quoted at 89 to 91 cents on the dollar [17]. At those prices, the roughly $18 billion banks reportedly lent would be marked $1.6 billion to $2 billion below par, by my calculation [17].</p><p style="text-align: justify;">So the notice moves risk. A few tens of millions a year of it is deferred on Blue Owl&#8217;s equity. The rest, over time, lands on Oracle&#8217;s credit: Oracle will pay for it in new-issue spreads when it next borrows, and whoever lends to it bears the risk.</p><h2><strong>Oracle&#8217;s case</strong></h2><p style="text-align: justify;">Michael Egbert, an Oracle vice president, wrote to The Santa Fe New Mexican that &#8220;force-majeure notices are commonplace in developments of this scale and are often used to preserve contractual rights among project partners,&#8221; and that they &#8220;do not, by themselves, establish a project delay or change delivery expectations&#8221; [18]. That is true of many notices: a tenant facing a permit fight it does not control is doing what counsel advises. And no, Larry Ellison&#8217;s pledged shares are a separate matter: Oracle&#8217;s proxy says those loans fund only &#8220;outside personal business ventures&#8221; [19].</p><h2><strong>What would prove me wrong</strong></h2><p style="text-align: justify;">Two dates and a sentence, all public. New Mexico&#8217;s decision on the fuel-cell air permit, due by November 23rd; other agencies&#8217; pipeline decisions under federal authority, due by December 3rd [2]; and Oracle or Blue Owl saying on the record that Jupiter&#8217;s lease has started on time. If all three go Jupiter&#8217;s way, the notice was mostly paperwork, and I will say so.</p><p>Until then, the act of God has a letterhead and a docket number: CP26-80.</p><h3><strong>Notes</strong></h3><p>[1] Bloomberg, &#8220;<a href="https://news.bloomberglaw.com/artificial-intelligence/oracle-cites-force-majeure-to-shield-itself-on-big-data-center">Oracle Cites Force Majeure to Shield Itself on Big Data Center</a>&#8220;, September 24th, 2026, citing &#8220;people familiar with the situation&#8221;. Independent report: Reuters, same day, via <a href="https://www.ksl.com/article/51628131/oracle-cites-force-majeure-to-shield-itself-on-data-center-bloomberg-news-reports">KSL</a>: &#8220;Oracle issued a &#8220;force majeure&#8221; notice to a Blue Owl unit &#8230; citing potential delays in securing power for the project, a person familiar with the matter told Reuters&#8221;. Blue Owl, per Reuters, &#8220;owns the data-center developer Stack Infrastructure&#8221;. CNBC &#8220;confirmed&#8221; the recipient was a Blue Owl unit. Company statements to <a href="https://www.cnbc.com/2026/09/24/oracle-data-center-force-majeure.html">CNBC</a>, September 24th, 2026. Neither company has confirmed the notice on the record; Blue Owl&#8217;s statement refers to &#8220;this notice&#8221;. Reported, not confirmed. Captured: sources/2026-09-24-oracle-force-majeure-project-jupiter-blue-owl.md.</p><p>[2] New Mexico State Land Office, <a href="https://www.nmstatelands.org/wp-content/uploads/2026/07/2026-07-14-Letter-re-Informal-Request-for-Reconsideration_Final.pdf">letter of July 14th, 2026</a> denying reconsideration (first denial by letters of March 20th, 2026), and <a href="https://www.nmstatelands.org/2026/07/15/commissioner-garcia-richard-again-denies-request-to-run-portion-of-project-jupiter-pipeline-through-state-lands/">press release of July 15th, 2026</a>. Federal Energy Regulatory Commission, docket CP26-80-000, <a href="https://www.federalregister.gov/documents/full_text/text/2026/07/29/2026-15286.txt">Federal Register, July 29th, 2026</a>: the prior notice request &#8220;will proceed as an application for case-specific authorization&#8221;; &#8220;90-day Federal Authorization Decision Deadline--December 3, 2026&#8221;. Same notice: &#8220;On April 13, 2026, Commission staff protested the prior notice request because Transwestern did not provide a copy of a finding by the New Mexico State Historic Preservation Office&#8221;; &#8220;Transwestern did not file the documentation by the deadline (or subsequently)&#8221;. NMED air permit deadline of November 23rd per <a href="https://techcrunch.com/2026/09/24/oracle-sends-force-majeure-notice-on-its-new-mexico-stargate-data-center/">TechCrunch</a>.</p><p>[3] Oracle, &#8220;<a href="https://www.oracle.com/news/announcement/oracle-borderplex-and-bloom-energy-to-power-project-jupiter-with-fuel-cell-technology-2026-04-27/">Oracle, BorderPlex and Bloom Energy to Power Project Jupiter with Cleaner, Water-Efficient Fuel Cell Technology</a>&#8220;, April 27th, 2026: &#8220;Construction continues to move forward on schedule.&#8221;; &#8220;Oracle will continue to bear all energy costs for Project Jupiter&#8221;.</p><p>[4] Oracle Corporation, <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312526277521/orcl-20260531.htm">Form 10-K for fiscal 2026</a>, filed June 22nd, 2026, risk factors: &#8220;we depend on third parties to develop and deliver certain data center capacity and related infrastructure, and their inability to obtain financing, complete construction on schedule or manage construction cost overruns could delay the availability of data center space or increase our costs. We have faced, and may continue to face, challenges with securing reliable and cost-effective power sources&#8221;. Neither the 10-K nor the <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312526389274/orcl-20260831.htm">10-Q</a> mentions force majeure, Jupiter or New Mexico.</p><p>[5] Oracle co-chief executive Clay Magouyrk on the September 10th, 2026 earnings call, as reported by the <a href="https://www.abqjournal.com/business/oracle-invokes-force-majeure-on-project-jupiter-but-says-data-center-still-on-track/3127907">Albuquerque Journal</a>, September 24th, 2026: &#8220;anyone that&#8217;s been in the business of doing construction or large-scale infrastructure development, if their plan relies on 100% achievement of every one of their deliverables, we have a term for that: It&#8217;s called a bad plan.&#8221; &#8220;We&#8217;re making very good progress in terms of construction, (the) data center is definitely on track.&#8221; The Journal: &#8220;Pressed by an analyst about potential data center delays in New Mexico and Wisconsin&#8221;. Same passage in the <a href="https://www.fool.com/earnings/call-transcripts/2026/09/11/oracle-orcl-q1-2027-earnings-call-transcript/">call transcript</a> (Motley Fool, machine-style): &#8220;we do not assume 100% of everything is going to work all the time. And we have backup options for those things&#8221;; &#8220;We are going through the process of acquiring our air permit.&#8221;</p><p>[6] Reuters via <a href="https://www.ksl.com/article/51628131/oracle-cites-force-majeure-to-shield-itself-on-data-center-bloomberg-news-reports">KSL</a>, September 24th, 2026, one person familiar with the matter: &#8220;The source said that securing power for the site is Oracle&#8217;s responsibility under the contract. Oracle cannot terminate the lease under any circumstances&#8221;; &#8220;with Oracle responsible for paying the debt costs&#8221;; &#8220;During the development stage, Blue Owl earns a 9% yield on its equity &#8230; upon completion of the project, the levered yield is expected to be around 11%. By invoking force majeure, Oracle extends the period during which it pays Blue Owl the lower development-stage rent.&#8221; &#8220;Blue Owl will still receive the higher rent for the originally planned duration, but its start will be delayed.&#8221; Force majeure, per Reuters, is invoked &#8220;when problems arise beyond their control&#8221;. William Blair, in the same Reuters report: &#8220;fiscal 2027 should not be affected, since Jupiter contributes no revenue this year.&#8221; Reported, not confirmed. Two points of yield on the equity, in billions of dollars a year. Author&#8217;s calculation: 0.02 &#215; 3 = 0.06. Hours of last quarter&#8217;s capex (quarter of 92 days). Author&#8217;s calculation: 0.06 &#247; 28.499 &#215; 92 &#215; 24 &#8776; 4.6.</p><p>[7] Oracle Corporation, <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312526389274/orcl-20260831.htm">Form 10-Q for the quarter ended August 31st, 2026</a>, filed September 11th, 2026: &#8220;As of August 31, 2026, we had $288 billion of additional lease commitments, substantially all related to data center arrangements, that are generally expected to commence between the second quarter of fiscal 2027 and fiscal 2029 and for terms of fifteen to nineteen years that were not reflected on our condensed consolidated balance sheets&#8221;. Capital expenditures of $28.5 billion in the quarter. The same commitments stood at $260 billion three months earlier: Oracle <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312526277521/orcl-20260531.htm">10-K for fiscal 2026</a>, &#8220;As of May 31, 2026, we had $260 billion of additional lease commitments&#8221;. Captured: sources/2026-09-11-oracle-q1-fy27-10q-288bn-uncommenced-leases.md.</p><p>[8] The AI Realist: &#8220;<a href="https://www.airealist.ai/p/cloud-vs-clout">Cloud vs. Clout</a>&#8220;, March 11th, 2026; &#8220;<a href="https://www.airealist.ai/p/welcome-to-hotel-abilene">Welcome to Hotel Abilene</a>&#8220;, March 8th, 2026; &#8220;<a href="https://www.airealist.ai/p/the-500-billion-umbrella">The $500 Billion Umbrella</a>&#8220;, May 6th, 2026.</p><p>[9] Oracle pricing term sheets on EDGAR: <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312525017365/d901407dfwp.htm">January 30th, 2025</a> (5.500% notes due 2035 at +97 basis points; 6.000% due 2055 at +122), <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312525215807/d49252dfwp.htm">September 24th, 2025</a> (5.200% due 2035 at +105; 5.950% due 2055 at +125), <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312526033882/d44245dfwp.htm">February 2nd, 2026</a> (5.700% due 2036 at +145; 6.700% due 2056 at +180). Chart data: data/oracle-issue-spreads.csv. Funding plan, <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312526032650/d36096dfwp.htm">free writing prospectus of February 1st, 2026</a>: &#8220;Oracle does not expect to issue additional bonds during calendar year 2026 beyond this transaction.&#8221;; &#8220;This funding plan reflects Oracle&#8217;s commitment to maintaining an investment-grade rating&#8221;; &#8220;On the equity side, Oracle plans to raise approximately half of its 2026 funding through a combination of equity-linked and common equity issuances.&#8221;</p><p>[10] Oracle 10-Q (note 7): &#8220;During the first quarter ended August 31, 2026, we fully utilized the ATM Program and issued 141 million shares of common stock under the ATM Program for net proceeds of $19.9 billion.&#8221; Mandatory convertible preferred: <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312526034351/d24013dfwp.htm">free writing prospectus</a>, February 2026, 100 million depositary shares at $50.00.</p><p>[11] S&amp;P Global Ratings, &#8220;<a href="https://www.spglobal.com/ratings/en/regulatory/article/-/view/sourceId/101695609">Oracle Corp. Downgraded To &#8216;BBB-/A-3&#8217; From &#8216;BBB/A-2&#8217; On Rising Business Risk And Weaker Cash Flow; Outlook Stable</a>&#8220;, July 9th, 2026 (read in full by the main session, 2026-09-28): &#8220;We adjust the debt in our forecast to include Oracle&#8217;s $260 billion of additional lease commitments (which are expected to commence between fiscal 2027 and 2029)&#8221;; &#8220;We forecast S&amp;P Global Ratings-adjusted leverage will reach the mid-4x area in fiscal 2027&#8221;; &#8220;We could lower the rating again if we: Expect Oracle to sustain leverage exceeding 4.5x&#8221;. Independent report of the same trigger: Reuters, &#8220;<a href="https://www.investing.com/news/economy-news/analysisoracle-corp-goes-for-highstakes-ratings-gamble-in-ai-strategy-4835408">Oracle Corp goes for high-stakes ratings gamble in AI strategy</a>&#8220;, August 4th, 2026, S&amp;P analyst Andrew Chang: &#8220;We could downgrade Oracle if Oracle sustains leverage exceeding 4.5 times.&#8221;</p><p>[12] Bloomberg terminal data for &#8220;ORCL 6.7 02/04/56&#8221; (Oracle&#8217;s 6.700% notes due 2056, the February 2nd, 2026 tranche in note 9), as posted by <a href="https://x.com/zerohedge/status/2104589989430329358">ZeroHedge on X</a>, September 28th, 2026, 15:12 UTC: &#8220;Last Price 82.386&#8221;; &#8220;Yield To Maturity 8.310&#8221;; &#8220;G-Spread 273.315&#8221;; same screen, &#8220;ORCLCP 5 Year CDS 237.065&#8221;. Screenshot saved as data/zh-2026-09-28.jpg. Market data relayed by a secondary outlet; reported, not independently confirmed by the desk.</p><p>[13] Late July: S&amp;P Global Market Intelligence data cited by Reuters, relayed by GuruFocus on <a href="https://finance.yahoo.com/markets/stocks/articles/oracle-stock-drops-2-3-185300368.html">Yahoo Finance</a>, July 29th, 2026: Oracle &#8220;near 200 basis points&#8221;; Meta about 93; investment-grade index about 53. Week of the notice: Seeking Alpha via <a href="https://www.tradingview.com/news/seekingalpha:62ed2825e094b:0-hyperscaler-cds-spreads-widen-with-oracle-hitting-record-high/">TradingView</a>: &#8220;Its 5-year CDS spread rose 16.2% from the prior week across both bid and ask quotes, reaching a record 227.15 bps&#8221;; the same report on &#8220;the 30-year yield climbing to its highest level since 2004 and the 10-year yield hitting a fresh 19-year high&#8221;.</p><p>[14] Enron: US Senate Committee on Governmental Affairs, <a href="https://www.govinfo.gov/content/pkg/CPRT-107SPRT82147/pdf/CPRT-107SPRT82147.pdf">staff report</a>, October 8th, 2002: &#8220;On November 28, 2001, Enron&#8217;s credit rating was reduced from investment grade to junk&#8230; On December 2, Enron declared bankruptcy.&#8221; Lehman: Flannery, Houston and Partnoy, &#8220;<a href="https://www.law.upenn.edu/live/files/22-flannery158upalrev20852010pdf">CDS Spreads as Viable Substitutes for Credit Ratings</a>&#8220;, University of Pennsylvania Law Review 158 (2010): &#8220;Bear Stearns and Lehman Brothers remained in the A category through the sample period until both firms failed in March and September 2008, respectively.&#8221; Boeing: Bloomberg via <a href="https://fortune.com/2020/03/18/coronavirus-boeing-stock-ba-share-price-bailout-request/">Fortune</a>, March 18th, 2020: &#8220;Five-year contracts climbed to 465 after rising as high as 488, a record. At that rate, the market is effectively pricing in 31% odds that the planemaker would default&#8221;; Boeing&#8217;s <a href="https://www.sec.gov/Archives/edgar/data/0000012927/000162828026050038/ba-20260630.htm">10-Q for the quarter ended June 30th, 2026</a>: &#8220;We currently maintain investment grade credit ratings across all three credit rating agencies.&#8221; A market-implied default probability is risk-neutral and runs above realised default rates: S&amp;P&#8217;s historical five-year cumulative default rate for BBB- issuers is 2.40% (S&amp;P, 2024 Annual Global Corporate Default And Rating Transition Study, table 26). Related reading on Oracle&#8217;s September filings: Pete Cleary, &#8220;<a href="https://filinginsights.substack.com/p/two-sec-filings-twelve-days-reading">Two SEC Filings, Twelve Days</a>&#8220;, Filing Insights, September 26th, 2026.</p><p>[15] Oracle <a href="https://www.sec.gov/Archives/edgar/data/1341439/000119312526277521/orcl-20260531.htm">10-K for fiscal 2026</a>, risk factors: &#8220;A downgrade could also reduce our access to, or increase the cost of, commercial paper or other short-term financing, affect the terms or availability of certain long-term commitments (including data center leases), limit eligibility to contract with certain customers, and increase collateral, letter of credit or other credit support requirements under certain contractual arrangements.&#8221;</p><p>[16] Bloomberg, <a href="https://assets.bbhub.io/professional/sites/27/US-Aggregate-Index.pdf">US Aggregate Index methodology</a>: &#8220;Securities must be rated investment grade (Baa3/BBB-/BBB- or higher) using the middle rating of Moody&#8217;s, S&amp;P and Fitch&#8221;. Moody&#8217;s Baa2 (negative) and Fitch BBB (stable) as printed on Oracle&#8217;s February 2nd, 2026 term sheet (note 9) and reported since.</p><p>[17] The Financial Times, as reported by <a href="https://searchlightnm.org/oracle-sends-force-majeure-notice-to-project-jupiter-developers-to-protect-itself-from-liability/">Searchlight New Mexico</a>, September 25th, 2026: &#8220;being quoted at just 89 to 91 cents on the dollar, according to the Financial Times&#8221;. The FT article was not read by the desk. Reported, not confirmed. The roughly $18 billion bank loan: <a href="https://www.bloomberg.com/news/articles/2025-11-07/banks-lend-18-billion-for-oracle-tied-data-center-in-new-mexico">Bloomberg</a>, November 7th, 2025, citing people familiar with the deal; Searchlight&#8217;s text: &#8220;Oracle&#8217;s $18 billion in loans tied to the data center&#8221;. Reported, not confirmed. In billions of dollars, at 91 cents. Author&#8217;s calculation: 18 &#215; (1 &#8722; 0.91) = 1.62. At 89 cents. Author&#8217;s calculation: 18 &#215; (1 &#8722; 0.89) = 1.98.</p><p>[18] Michael Egbert, Oracle vice president, email to The Santa Fe New Mexican, as published by <a href="https://searchlightnm.org/oracle-sends-force-majeure-notice-to-project-jupiter-developers-to-protect-itself-from-liability/">Searchlight New Mexico</a>, September 25th, 2026: &#8220;Force-majeure notices are commonplace in developments of this scale and are often used to preserve contractual rights among project partners. They do not, by themselves, establish a project delay or change delivery expectations.&#8221; TechCrunch reported that &#8220;Neither Oracle nor Blue Owl immediately responded to TechCrunch&#8217;s requests for comment.&#8221;</p><p>[19] Oracle Corporation, <a href="https://www.sec.gov/Archives/edgar/data/0001341439/000119312526402816/orcl-20260925.htm">Definitive Proxy Statement</a>, filed September 25th, 2026: &#8220;As of September 21, 2026, Mr. Ellison &#8230; had pledged 413 million shares of Oracle common stock as collateral to secure certain personal indebtedness.&#8221;; &#8220;The pledged shares secure personal term loans only used to fund outside personal business ventures.&#8221;</p>]]></content:encoded></item><item><title><![CDATA[Vent Mauvais]]></title><description><![CDATA[Verlaine&#8217;s ill wind carries the leaf where it will. States bought Mistral without a tender; where free choice can be measured, it goes elsewhere.]]></description><link>https://www.airealist.ai/p/vent-mauvais</link><guid isPermaLink="false">https://www.airealist.ai/p/vent-mauvais</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Thu, 24 Sep 2026 10:35:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!E4B7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!E4B7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!E4B7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png 424w, https://substackcdn.com/image/fetch/$s_!E4B7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png 848w, https://substackcdn.com/image/fetch/$s_!E4B7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png 1272w, https://substackcdn.com/image/fetch/$s_!E4B7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!E4B7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1505698,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/217210377?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!E4B7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png 424w, https://substackcdn.com/image/fetch/$s_!E4B7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png 848w, https://substackcdn.com/image/fetch/$s_!E4B7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png 1272w, https://substackcdn.com/image/fetch/$s_!E4B7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe72ec2a9-37b0-4db2-ab74-03c8b482aaed_1456x720.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On June 16th, 2026, Arthur Mensch promised &#8220;a very exciting model to come this summer&#8221;, open-weight, with early access in July for key partners. Summer ended on September 23rd with no public release.[1] In its last days, on September 21st, Mistral&#8217;s own product account posted that GLM-5.3, a model trained by Z.ai in Beijing, was now in Vibe, the product Le Chat became, and &#8220;the default in the web app&#8221;.[2]</p><p style="text-align: justify;">In &#8220;Chanson d&#8217;automne&#8221;, Verlaine wrote that turn of the season: the violins of autumn sob, and the poet goes off &#8220;au vent mauvais / Qui m&#8217;emporte / De&#231;&#224;, del&#224;, / Pareil &#224; la / Feuille morte&#8221;, on the ill wind that carries him here and there like a dead leaf. A mistral is a wind, the cold north wind of the French Mediterranean.[3] But in this story, Mistral is the leaf, and it goes where the money blows: a frontier lab for its investors, a sovereign supplier for the states that buy it, and, since August, the host of Chinese models, one of them now the default.</p><p style="text-align: justify;">Mistral has two kinds of customers. The states whose purchases you can read, Luxembourg and Hesse, bought it without a competitive tender and published the price to the euro. Companies buy it too, Airbus and HSBC among them; they announce the deal, and rarely the price, and no instrument outside their walls can see what they run.[4] Where a free choice can be measured at all, in a survey of German firms, on the Hugging Face download counter and on one developer marketplace, it goes to other models. And the valuation, more than &#8364;21 billion as of September 8th, 2026, rests on numbers nobody outside the company and its investors can read.[5]</p><p style="text-align: justify;">Les Echos has spent September on a three-part &#8220;saga Mistral&#8221;, and on the 17th called the company &#8220;le champion qui pourrait abandonner la course &#224; l&#8217;IA&#8221;, the champion that might drop out of the AI race.[6] The story has been told. What follows is the ledger: who pays for Mistral, how much, and who picks it when nobody makes them.</p><h2><strong>Who bought</strong></h2><p style="text-align: justify;">On June 17th, 2025, Luxembourg&#8217;s state IT center, the CTIE, signed a contract with Mistral AI carrying a published value of &#8364;39,500,000. The notice on the EU&#8217;s tender platform records one tender received, no duration, and a single award criterion, weighted at 100 percent: &#8220;Absence de concurrents pour des raisons techniques&#8221;, no competitors for technical reasons. The justification lists five conditions the supplier had to meet simultaneously; taken together, they describe one company.[7]</p><p style="text-align: justify;">Six weeks later, on July 28th, Luxembourg&#8217;s Direction de la D&#233;fense recorded a second award to Mistral, with a published value of &#8364;4,880,000, also without a call for competition, coded &#8220;technical&#8221;, and its only written justification is a single article of the Grand Duchy&#8217;s defense-procurement law.[8]</p><p style="text-align: justify;">The third is German. On July 2nd, 2026, the Land of Hesse awarded &#8364;7,800,000 to Mistral after a documented market exploration and five mandatory requirements, including that the models the vendor trained itself were developed within the European Economic Area or a country with an EU adequacy decision. Hesse also published a voluntary transparency notice 45 days in advance, which is the step that allows a rival to challenge the award. None is on the record. The notice says the absence of competition is not the result of an artificial narrowing of the requirements but of the need as lawfully defined and the market as it stood.[9]</p><p style="text-align: justify;">Fourteen months after its first purchase, the Grand Duchy joined Mistral&#8217;s cap table in the September 2026 Series D.[10] Nothing in the notices or the funding announcement connects the two, and three different arms of the state were involved.[11]</p><p style="text-align: justify;">So that is the buying side: &#8364;52.18 million of published contract value, all of it from governments, none of it competed. On the French side, the state that treats Mistral as its national champion has published no award at all: I searched three registers for the company every year. What Les Echos reported on September 2nd, and what the ministry confirmed to Capital, is a &#8364;6 million, non-exclusive contract for 2026 and 2027, said to follow a public-procurement procedure; on September 22nd, it was in neither.[12]</p><p style="text-align: justify;">What France has published is a defense framework with no amount,[13] and one costing from its own inspectors: about &#8364;300,000 a year for 10,000 active users on the state&#8217;s Mistral-powered assistant, and about &#8364;3 million a year if it went to one or two million, with Mistral named as the current supplier. French defense contracts can be withheld from publication. But the Luxembourg defense award was itself a negotiated defense contract, and Luxembourg published it, so the silence is a choice, or a contract carried under a reseller&#8217;s name; the registers cannot say which.</p><p style="text-align: justify;">Mensch has put a split on the record, once, from memory. Asked under oath on May 12th what share of revenue public orders represent, he said 20 percent of software revenue, 10 percent for the French state, &#8220;de t&#234;te&#8221;, with compute carved out, and named Luxembourg unprompted as a significant framework customer.[14] No audited figure supports the 80 percent that leaves for private buyers.</p><h2><strong>Who chose</strong></h2><p style="text-align: justify;">Bitkom, the German digital industry association, published a representative survey on September 9th, 2026, the day after Mistral announced its &#8364;3 billion round. Its research arm interviewed 603 German companies with 20 or more employees; 343 of them use AI, and those 343 form the basis of what follows. Asked which AI applications they use, 76 percent said ChatGPT, 35 percent Microsoft Copilot, 28 percent Gemini, 6 percent SAP&#8217;s Joule, 3 percent Claude. Mistral: none, and none the year before either, when the same question was put. The release says so in one sentence: no surveyed company reported using models from &#8220;des f&#252;hrenden europ&#228;ischen Anbieters Mistral&#8221;, the leading European provider Mistral.[15]</p><p style="text-align: justify;">This instrument measures brand names, not spend, which is why Claude scores 3 percent despite being a large enterprise API business. SAP&#8217;s Joule, at 6 percent, is not the clean comparator it looks like either: SAP has resold Mistral since 2024, and a German firm running Mistral through SAP answers &#8220;SAP&#8221;.[16] Le Chat, now Vibe, is a named end-user assistant, the kind of product a brand question sees best, and it read zero in both waves. And Mistral&#8217;s own &#8220;125+ global enterprises&#8221; is compatible with zero of 343 German firms: a hundred and twenty-five customers worldwide is a rounding error in a national sample.</p><p style="text-align: justify;">The second instrument is Hugging Face, where the people who run models pull them. On September 22nd, 2026, I read the platform&#8217;s public counter. Mistral&#8217;s Apache-2.0 flagship, Mistral Large 3, a 675-billion-parameter mixture-of-experts model released in December 2025, had been downloaded 2,097 times in the past 30 days and 13,457 times in its lifetime. Its four sibling repositories add 11,033 more; counted together, the gap to the comparator below is still 85 to one.[17] The fair comparison is another open model of the same size. DeepSeek-V3, 671 billion parameters, also a mixture of experts, released eleven months earlier in December 2024 and long past any launch spike, had been downloaded 1,118,995 times in the same thirty days, and 21 million times in its life.[18] Same size, both permitting commercial use, same window, same counter. A factor of about 530 on the flagship repository alone.</p><p style="text-align: justify;">The counter counts file fetches, so a model pulled once into a company mirror counts once, and it does not see API usage at all. The counters tell you something narrower and, for a company with &#8220;open-weight&#8221; in the headline of its funding announcement, more awkward: the open flagship is the one almost nobody downloads. The smaller open-weight model, Medium 3.5, released in April under a modified MIT license, drew 89,194 pulls in the same window, forty times the flagship and a thirteenth of DeepSeek-V3.[17]</p><p style="text-align: justify;">The third instrument is OpenRouter, the marketplace through which developers route requests to whichever model they choose; it does not see Mistral&#8217;s own API. Its chart of token share by the model lab&#8217;s country, posted by OpenRouter&#8217;s Peter Walker on September 21st, has France holding most of the platform&#8217;s tokens in early 2024, when Mistral&#8217;s small open models were the thing to run, 5 percent by the end of 2025, and no band wide enough to label since. On the same site&#8217;s live ranking of model authors by share of requests, Mistral stands eighth, at 2.6 percent, behind DeepSeek, Google, OpenAI, Z.ai, Qwen, Tencent and Anthropic. &#8220;Mistral really did have a moment there,&#8221; the post says. Developers with a choice are where the moment was.[19]</p><h2><strong>The first hour</strong></h2><p style="text-align: justify;">The fourth instrument is my own hands. On August 17th, 2026, I installed Mistral Vibe, the company&#8217;s command-line coding agent, to build a small tool: pull the audio of a YouTube video with yt-dlp, then transcribe it with a Mistral speech model.</p><p style="text-align: justify;">I asked Vibe for the transcription flow.[20] It was running Mistral Medium 3.5, the company&#8217;s newest general model, in Vibe since May 22nd, 2026. It told me that Mistral has no speech-to-text model, and steered me elsewhere. Voxtral has been public since July 15th, 2025, with two more speech models since then. The company&#8217;s own agent, on the company&#8217;s own three-month-old model, denied that it existed, and a developer who trusted the answer would have built on a competitor&#8217;s API. Training cutoff is not the excuse. The documentation is online, the company publishes an <code>llms.txt</code> index of it for exactly this purpose, and the agent had a web-search tool in that session: it used it four times later in the same session, and not once before telling me the product did not exist. I re-ran the question on September 22nd on the same model, still the CLI&#8217;s default, in a single turn with web tools on, and the model gave the same answer: Mistral &#8220;does not currently offer a dedicated speech-to-text model&#8221;.</p><p>The fault is a broken window, not a model-quality problem, and the fix is cheap: an official Mistral documentation server that Vibe consults before it answers a question about Mistral.</p><h2><strong>What the founding memo said</strong></h2><p style="text-align: justify;">Mistral&#8217;s founding strategic memo, seven pages, from before its first round in the spring of 2023, was published in full by Sifted on June 21st, 2023.[21]</p><p style="text-align: justify;">The memo says where the founders thought the money was: &#8220;We believe that most of the value in the emerging generative AI market will be located in the hard-to-make technology, i.e., the generative models themselves.&#8221; It says what they promised: to become &#8220;a research leader in the generative AI field, eventually offering the best technology within 4 years&#8221;, which puts the deadline in April 2027.</p><p style="text-align: justify;">And it says how they planned to compete: &#8220;Training a competitive model requires at least an exa-scale cluster for a few months. We intend to rent such capacity for a full year.&#8221; So this was never a plan to win on scale, and the founders said so on page five. It was a plan to win on efficiency: &#8220;know-hows that will allow us to gain a factor 10-100 in training efficiency compared to public methods&#8221;, so that the team could &#8220;train the strongest model for a given computational budget&#8221;.[22]</p><p style="text-align: justify;">The open-source promise was hedged from the start: &#8220;We will balance our open-source strategy with our economic interests, keeping the strongest and most specialized models reserved for negotiated access.&#8221; And the route to market was a network, stated plainly: the early investors &#8220;will open all necessary doors for acquiring high-quality datasets&#8221;, and &#8220;the founding team is already organizing business exploration with major French and European industrial actors&#8221;.[23] Doors and introductions, in April 2023. Not a drift.</p><p style="text-align: justify;">On June 16th, 2026, Arthur Mensch posted on LinkedIn that Mistral had &#8220;spent a few percent of the deployment of some of our competitors, in a domain that is fundamentally governed by flops budget&#8221;, and that &#8220;today, we do not yet own the best language models&#8221;.[1] The founding bet was that efficiency would beat a budget; in 2026 the founder wrote that budget governs. I won&#8217;t call that a contradiction; a founder is allowed to learn in three years.</p><h2><strong>What he told whom</strong></h2><p style="text-align: justify;">On January 14th, 2026, on an American podcast, Mensch agreed with his host that frontier models were commoditizing and gave the reason: &#8220;the knowledge unit to actually train a model is fairly short, so because it&#8217;s short, it actually circulates.&#8221; The word was his host&#8217;s; the reasoning was his.[24]</p><p style="text-align: justify;">On May 12th, 2026, under oath before a French parliamentary commission of inquiry, he described the value chain the other way up. The high-margin layer, the one that pays for research, &#8220;c&#8217;est l&#8217;intelligence artificielle&#8221;, is AI; the rest of cloud services &#8220;sont surtout des commodit&#233;s&#8221;, are mostly commodities. In 10,547 words of testimony, the word &#8220;commodit&#233;&#8221;, commodity, appears once, and it points at the cloud, not at the models.[14] His own bridge between the two is on the same podcast, at 23:01: &#8220;to get to value, they need to have great models. So the two things are extremely linked together.&#8221;[24]</p><p style="text-align: justify;">On September 8th, 2026, the funding announcement that raised &#8364;3 billion described Mistral as &#8220;the only AI company in the world building the full stack&#8221;. It is not in any of the nine archived snapshots of the company&#8217;s home page from May 2023 to September 2026.[25]</p><p style="text-align: justify;">The September 21st post does not say which tiers get the Chinese model; the free tier&#8217;s command-line client did not offer it when I looked the next day.[2] Mistral&#8217;s newsroom, pricing page, and product page say nothing about it; the model&#8217;s own documentation page, dated September 15th, says it is &#8220;served without Mistral modifications&#8221; and prices it at Z.ai&#8217;s own list for input and output.[26]</p><p style="text-align: justify;">On September 18th, three days before that post, Juan Branco, a left-wing lawyer and agitator, posted that Mistral lived on public orders pushed by the &#201;lys&#233;e and served its clients &#8220;des mod&#232;les chinois, librement t&#233;l&#233;chargeables, d&#8217;avant derni&#232;re g&#233;n&#233;ration&#8221;, freely downloadable Chinese models of the previous generation. Mistral had announced the Chinese models itself, on August 11th. Its own documentation had made &#8220;previous generation&#8221; stale on September 15th, and Branco&#8217;s rebuttal narrowed it to GLM 5.2 &#8220;pendant tout un &#233;t&#233;&#8221;, all summer long, which was true of the model if not of the whole summer.[27] Mensch answered twice on September 20th: &#8220;Nos travaux d&#233;plaisent &#224; certaines petites frappes&#8221;, our work displeases certain small-time thugs, and &#8220;c&#8217;est un obsessionnel&#8221;, the man is obsessive. He contested no figure and no fact, and by September 22nd had deleted both posts. A buyer at Airbus or HSBC reads the same feed, and sees the founder of a &#8364;21 billion company, accused in public of living on state orders, answer with an insult rather than the 20 percent he had given the Assembl&#233;e under oath in May.</p><h2><strong>The rolodex, counted</strong></h2><p style="text-align: justify;">My May piece argued that Mistral was converting a perishable contact list into permanent law. Since then, I have counted the list, and the honest result is smaller than the phrase suggests.</p><p style="text-align: justify;">Mistral names 23 investors in its Series D, and CMA CGM&#8217;s own release adds a 24th, from 2023.[28] Its customer pages name 42 companies, and partner, compute, and September&#8217;s announcements add eleven, for 53 commercial names against a claimed &#8220;125+&#8221;.[29] Seven names are on both lists: ASML, CMA CGM, Luxembourg, BNP Paribas, Belfius, Nvidia, and Samsung, which led the Series D and announced the next day that it would deploy Mistral Large on premises across its chip operations.[5] Five of the seven are on the customer pages; the other 37 customers there, Airbus, HSBC, Capgemini, and on down the list, have no equity link in any document I read.[30]</p><p style="text-align: justify;">So the cap table does not explain the customer list. It explains four of the seven, the ones where a stake is large, or a purchase has a price: Samsung, which led the round, and these three. ASML led the Series C with &#8364;1.3 billion for about 11 percent; it is a named customer and a compute anchor; its own audited filing carries the stake at &#8364;1,320.7 million, and its announcement prices the commercial side at nothing: an agreement to &#8220;explore the use of AI models&#8221;.[31] CMA CGM invested in 2023, anchors the compute plan, and published a figure Mistral never has: a five-year partnership &#8220;supported by a &#8364;100 million investment&#8221;, which the group&#8217;s releases never split between its own program and what it pays Mistral.[32] And Luxembourg bought twice, without competition, then invested.</p><p style="text-align: justify;">Of 56 relationships announced by Mistral or its counterparties, ten have a published value, with eleven figures in total because Microsoft appears twice, and three of the eleven are words rather than numbers. Three are the procurement notices, exact to the euro. Two are equity: ASML&#8217;s audited stake and Microsoft&#8217;s 2024 convertible bonds, a figure neither party announced, reported by TechCrunch the next day and later printed by two competition authorities. Two are acquisition prices: Mistral&#8217;s own &#8220;several million euros&#8221; for Pimento, and a press report on Emmi AI that neither party confirmed. And four are private commercial deals whose size a party chose to publish: CMA CGM&#8217;s &#8364;100 million, by the buyer; TotalEnergies&#8217; &#8220;more than &#8364;100 million&#8221; over three years, by the buyer, on September 15th; HUMAIN&#8217;s, &#8220;in the hundreds of millions of Euros&#8221;; and Microsoft&#8217;s compute agreement of July 2026, &#8220;a multibillion dollar commitment&#8221; in Microsoft&#8217;s own release, with no figure, no term, and no mention in the 10-K Microsoft filed eight days later.[33] So the only exact, audited euro figures associated with Mistral are the amounts others paid for its shares. What anyone pays for its product is a procurement notice, a buyer&#8217;s press release, or an adjective. And no revenue figure is filed or audited anywhere: Mistral AI SAS has never deposited annual accounts in France, and no counterparty&#8217;s filing states a contract value.[4]</p><h2><strong>What actually exists</strong></h2><p style="text-align: justify;">The commoditization diagnosis was right, and Mensch made it in January, before his own numbers made it visible. His company&#8217;s speech model, Voxtral-Mini-4B-Realtime, is a genuine world-class franchise: 1,746,931 downloads in thirty days and 11.4 million since January 2026, the only post-2025 model in the company&#8217;s top three.[34] Its 2023 and 2024 small models had a publisher&#8217;s distribution, with 64 million and 46 million lifetime pulls for the two 7B versions.[35] Mistral has had users who signed nothing, in the millions. They came for the small models and for speech, and the 2025 flagship did not inherit them. And there is a documented repeat buyer: BNP Paribas, also an investor, renewed for three years in May 2026.[4]</p><p style="text-align: justify;">The flagship line reopened under Apache 2.0 with Mistral Large 3, 21 months after the closure of Mistral Large in February 2024.[36] And the domain he named in his June post, &#8220;voice, vision and document processing&#8221;, is where the company&#8217;s measured strength actually is. He knows which of his products work. His funding announcement sells the other one.</p><h2><strong>The label, written down</strong></h2><p style="text-align: justify;">The institutions that designated the champion wrote the label down before they measured the product. Thirty days after Mistral shipped its flagship closed, API-only, and first on Microsoft Azure, in February 2024, the European Parliament&#8217;s research service listed Mistral AI as &#8220;Open source&#8221;, on the strength of Mistral 7B, the only model it named.[37][38] France&#8217;s competition authority did look: its opinion of June 28th, 2024 recorded that Mistral had not released Mistral Large as open weights, set out the test for control, and did not apply it to Microsoft&#8217;s deal.[39] And on September 2nd, 2026, the economy minister told AFP: &#8220;Si l&#8217;IA europ&#233;enne se r&#233;sume &#224; Mistral, alors nous sommes fichus&#8221;, if European AI comes down to Mistral, we are done for.[40]</p><h2><strong>What would have to break</strong></h2><p style="text-align: justify;">In March I wrote that Mistral existed despite the state.[41] In May I relayed the company&#8217;s reported $400 million of annual recurring revenue, with a billion targeted for the year.[42] Four months on, nobody outside the company and its investors can confirm either figure. Aleph Alpha, the company that Germany&#8217;s press and ministers treated as its champion, filed its accounts for 2022, 2023 and 2024 before signing this month to merge with Cohere. [43] Mistral AI SAS has filed none, ever.[4] The French state holds shares through Bpifrance,[5] so it can read the accounts Mistral has not filed; the public cannot.</p><p>The argument fails if any of three things happens:</p><ol><li><p>If Mistral files audited accounts in France, the valuation no longer rests on numbers nobody outside can read; and if they put private customers near the 80 percent of software revenue that Mensch&#8217;s sworn 20 percent leaves them.</p></li><li><p>If any Apache-2.0 Mistral flagship passes 100,000 thirty-day downloads on the Hugging Face counter, or the DeepSeek-V3 gap falls below ten to one, the choosing side is wrong. </p></li><li><p>If a general-purpose Mistral model ships before the memo&#8217;s April 2027 deadline and scores above, say, 30 on Artificial Analysis&#8217;s intelligence index, against 14 for its best today and 45 for the Chinese model it now serves, the efficiency bet is back on.</p></li></ol><p style="text-align: justify;">Until then, the picture is the one the documents draw: states buying without a tender, seven investors doubling as customers or partners, a free choice going elsewhere wherever it can be measured, and private contracts nobody outside can price. The four-year clock the founders set themselves runs out in seven months, with no general-purpose model released since April 28th, and, by its own product account, a Chinese one as the web app&#8217;s default. De&#231;&#224;, del&#224;.</p><h3><strong>Notes</strong></h3><p>[1] Arthur Mensch, <a href="https://www.linkedin.com/posts/arthur-mensch_we-somehow-got-put-in-the-spotlight-the-last-share-7472689447629729792-eCno/">LinkedIn post</a>, June 16th, 2026, 16:40 UTC, 495 words, read in full; the date is decoded from the post&#8217;s activity identifier. Independent confirmation of the quoted sentences: <a href="https://sifted.eu/articles/mistral-arthur-mensch-open-source-anthropic">Sifted</a>, June 17th, 2026; Data Center Dynamics, June 18th, 2026; TechCrunch, July 4th, 2026. Sifted anglicises the spelling; the post reads &#8220;organization&#8221;. Same post, re-read September 24th, 2026: &#8220;We have a very exciting model to come this summer &#8211; it will be open-weight, and we&#8217;re opening early access to it in July.&#8221; Who the early access is for is in his <a href="https://x.com/arthurmensch/status/2066913356548542827">X post</a>of the same day, 15:58 UTC, read on X&#8217;s syndication endpoint September 24th, 2026: &#8220;First, we have a nice model coming this summer &#8211; we hope it will delight and surprise in a few capabilities. This will be the start of a new family of models, fat indeed, but sparse. We&#8217;re opening up an early access program in July for key partners in research, government and the&#8221; (truncated in the endpoint text). Mistral&#8217;s news listing from June 23rd to September 22nd, 2026 carries Leanstral 1.5, Robostral Navigate, Shieldstral and OCR releases and no such model. The end of summer: US Naval Observatory, <a href="https://aa.usno.navy.mil/api/seasons?year=2026&amp;tz=0">Earth&#8217;s seasons, 2026</a>, September equinox on September 23rd, 2026, 00:05 UTC, 02:05 in Paris; re-run the no-release check on the day of publication. Mistral held a &#8220;Building Frontier AI&#8221; slot at AI Engineer Paris on September 23rd, 2026, 18:40 to 19:05 CEST (L&#233;lio Renard Lavaud, programme at ai.engineer/paris/2026). Checked at 19:15 and 22:54 CEST on the 23rd and 07:30 CEST on the 24th: no post on mistral.ai/news, no changelog entry after August 31st, no new repository under mistralai on Hugging Face since July 16th, no new Mistral entry on Artificial Analysis, no launch item in Google News in French or English.</p><p>[2] The post, <a href="https://x.com/mistralvibe/status/2102056993531871446">@mistralvibe on X</a>, September 21st, 2026, 15:26:56 UTC, verbatim: &#8220;GLM 5.3, hosted on Mistral, is now in Vibe Code, and the default in the web app. 1M context window and generous limits included. Happy &#128755;&#65039;!&#8221; Read via two independent endpoints; the Internet Archive refused every save attempt that day, so the corpus capture is the record copy. That Vibe and Le Chat are one product: <code>mistral.ai/products/vibe</code> and <code>mistral.ai/products/le-chat</code> served the identical 446,309-byte page on September 22nd, 2026. The web app is login-gated and Mistral&#8217;s documentation names no default model; the claim &#8220;the default in the web app&#8221; rests on this post alone. Author&#8217;s check, September 22nd, 2026, Mistral Vibe 2.25.7, command-line client, plan shown as &#8220;Free&#8221;, <code>/model</code> selector: &#8220;Default (currently Mistral Medium 3.5)&#8221;, &#8220;Mistral Medium 3.5&#8221;, &#8220;Devstral (local)&#8221;; &#8220;2 models&#8221; in the session header; GLM-5.3 not offered. The web app&#8217;s paid-tier selector was not seen. A French text on the same launch, posted to Reddit r/MistralAI by u/isidor_n on the evening of September 21st, reads &#8220;GLM-5.3 est maintenant disponible dans Mistral Vibe Code pour les utilisateurs Pro, Team et Enterprise. H&#233;berg&#233; et servi par Mistral AI dans l&#8217;UE. Limites d&#8217;utilisation g&#233;n&#233;reuses. Disponible avec jusqu&#8217;&#224; 1M de tokens de contexte dans Vibe Code.&#8221; It does not use the word &#8220;default&#8221;; the poster&#8217;s affiliation is unverified, so it corroborates the tier restriction and is not cited for it.</p><p>[3] Paul Verlaine, <a href="https://fr.wikisource.org/wiki/Po%C3%A8mes_saturniens_%281866%29/Chanson_d%E2%80%99automne">&#8220;Chanson d&#8217;automne&#8221;</a>, <em>Po&#232;mes saturniens</em>, Paris, Lemerre, 1866, third stanza verbatim: &#8220;Et je m&#8217;en vais / Au vent mauvais / Qui m&#8217;emporte / De&#231;&#224;, del&#224;, / Pareil &#224; la / Feuille morte.&#8221; Read September 24th, 2026. The wind: <a href="https://www.larousse.fr/dictionnaires/francais/mistral/51794">Larousse</a>, &#8220;mistral&#8221;, read September 24th, 2026: &#8220;Vent violent, froid, turbulent et sec, qui souffle du secteur nord, sur la France m&#233;diterran&#233;enne, entre les m&#233;ridiens de S&#232;te et de Toulon.&#8221; Nothing here says why the company chose its name.</p><p>[4] The <a href="https://bodacc-datadila.opendatasoft.com/api/explore/v2.1/catalog/datasets/annonces-commerciales/records?where=registre%20like%20%22952418325%22">BODACC open data</a>, registre 952418325, read September 22nd, 2026: 12 notices, none of family &#8220;d&#233;p&#244;t&#8221;; re-read September 24th, 2026: 13, the newest a &#8220;Modifications diverses&#8221; notice of that day, families &#8220;Cr&#233;ations&#8221; and &#8220;Modifications diverses&#8221; only; the government&#8217;s <a href="https://recherche-entreprises.api.gouv.fr/search?q=952418325">company register API</a> returns <code>finances: null</code> for the same SIREN; INPI accounts-ratio dataset, zero rows. SEC full-text search for &#8220;Mistral AI&#8221; across every 10-K, 20-F and 40-F filed from June 1st, 2025 to September 22nd, 2026: five registrants (ASML, SAP, TotalEnergies, Elastic, ProCap Financial), only ASML attaching a number, and that number equity; Microsoft&#8217;s 10-K and 10-Qs, zero occurrences; CoreWeave&#8217;s 424B4 prospectus of March 31st, 2025, 23 mentions and no contract value. Counterparty releases read with the contract acknowledged and the value absent: Airbus, May 28th, 2026, licences for &#8220;the full Mistral AI product suite&#8221;; HSBC, December 1st, 2025, &#8220;multi-year&#8221;; BNP Paribas, May 26th, 2026, renewed &#8220;for a three-year period&#8221;; AFP, January 16th, 2025, &#8220;multi-year&#8221;, and AFP&#8217;s 2025 results release of April 16th, 2026, zero mentions. SNCF&#8217;s Num&#233;rique &amp; IA director, on the group&#8217;s own <a href="https://www.groupe-sncf.com/fr/innovation/technologies-emergentes/ia-generative">page</a> updated August 5th, 2025, names its suppliers as &#8220;OpenIA, Anthropic et Mistral qui est le mod&#232;le fran&#231;ais. Nous combinons ces diff&#233;rents LLM en fonction du besoin, du mod&#232;le &#233;conomique et de leur consommation &#233;nerg&#233;tique.&#8221; &#8220;Mistral&#8217;s roughly $400M in annual recurring revenue&#8221; in my May piece, note 42, was relayed from press reporting and is not filed anywhere.</p><p>[5] Mistral AI, <a href="https://mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/">&#8220;Mistral raises &#8364;3B to make sovereign, open-weight AI the technology frontier&#8221;</a>, September 8th, 2026: &#8220;a post-money valuation of more than &#8364;21 billion&#8221;. Same announcement: &#8220;Samsung Electronics led the round, joined by co-leads Scaleup Europe Fund, managed by EQT, and existing investor PSG Equity.&#8221; Samsung Electronics, <a href="https://news.samsung.com/global/samsung-and-mistral-ai-announce-strategic-partnership-for-intelligence-driven-semiconductor-infrastructure">release</a>, September 9th, 2026: &#8220;Samsung has also led Mistral AI&#8217;s Series D funding round, securing a strategic equity stake&#8221;; &#8220;Samsung will integrate Mistral&#8217;s AI services and solutions &#8212; including its flagship large language model, Mistral Large &#8212; across its semiconductor operations to develop customized on-premises AI models&#8221;. Same announcement: &#8220;Existing investors a16z, ASML, Belfius, BNP Paribas CIB, Bpifrance, Carmignac, DST Global, Eurazeo, General Catalyst, Headline, Hillspire, Index Ventures, Korelya Capital, Lightspeed, NVIDIA, Phoenix Court&#8217;s Solar fund (home to LocalGlobe, Latitude and Solar) and Salesforce Ventures participated in the round.&#8221; Confirmed by <a href="https://www.theregister.com/ai-and-ml/2026/09/08/mistral-bags-3b-to-build-europes-sovereign-ai-champion/5294941">The Register</a>, September 8th, 2026.</p><p>[6] Les Echos, <a href="https://www.lesechos.fr/tech-medias/intelligence-artificielle/la-saga-mistral-le-champion-qui-pourrait-abandonner-la-course-a-lia-2251972">&#8220;La saga Mistral : le champion qui pourrait abandonner la course &#224; l&#8217;IA&#8221;</a>, September 17th, 2026, 06:17 CEST, part 2 of 3; standfirst: &#8220;La promesse de la start-up &#233;tait de d&#233;velopper des mod&#232;les capables de concurrencer les meilleurs du march&#233;. Trois ans plus tard, le projet a largement pivot&#233;.&#8221; Paywalled; the author read the headline, the standfirst and the passage quoting the founding memo (note 21). Part 3, &#8220;La saga Mistral : la souverainet&#233; &#224; g&#233;om&#233;trie variable d&#8217;un champion europ&#233;en&#8221;, ran September 24th, 2026.</p><p>[7] The notice, <a href="https://ted.europa.eu/en/notice/-/detail/415226-2025">TED 415226-2025</a>, OJ S 121/2025, published June 27th, 2025; contract reference 2501587 signed June 17th, 2025; winner identified by SIREN 952418325; procedure &#8220;negotiated without prior call for competition&#8221;, justification code &#8220;technical&#8221;; one tender received; award criterion &#8220;Absence de concurrents pour des raisons techniques&#8221; at weight 100. The five conditions in the justification: an already-built high-performing model, an open-source offering alongside commercial products, deployability on state infrastructure, first-party provider status, and a security guarantee. Wayback <a href="https://web.archive.org/web/20260922065031/https://ted.europa.eu/en/notice/-/detail/415226-2025">20260922065031</a>. The value is the notice&#8217;s total; no duration is published and it cannot be annualised.</p><p>[8] The notice, <a href="https://ted.europa.eu/en/notice/-/detail/492208-2025">TED 492208-2025</a>, OJ S 142/2025, published July 28th, 2025; buyer &#8220;Minist&#232;re des Affaires &#233;trang&#232;res et europ&#233;ennes, de la D&#233;fense, de la Coop&#233;ration et du Commerce ext&#233;rieur &#8211; Direction de la D&#233;fense&#8221;; justification, in full, &#8220;Article 28 1. e) de la loi du 26 d&#233;cembre 2012 sur les march&#233;s publics de la d&#233;fense&#8221;; winner record filled with placeholder data, company number FR00000000, town &#8220;Halles&#8221;, telephone 0033, email info@info.fr, read from the eForms XML on September 22nd, 2026. No archive snapshot could be made.</p><p>[9] The notice, <a href="https://ted.europa.eu/en/notice/-/detail/453996-2026">TED 453996-2026</a>, OJ S 125/2026, published July 2nd, 2026; buyer Land Hessen via HZD for the Hessian finance ministry; legal basis &#167;14 Abs. 4 Nr. 2 lit. b and &#167;14 Abs. 6 VgV; 24 months plus one 24-month extension; one tender. Two of the five requirements, verbatim: &#8220;Modellherkunft und -kontrolle: Der Auftragnehmer muss die bereitgestellten textverarbeitenden KI-Modelle eigenst&#228;ndig trainieren und nachweisbare Kontrolle &#252;ber die verwendeten Trainingsdaten und Trainingsprozesse aus&#252;ben.&#8221; and &#8220;Entwicklungsstandort: Die verwendeten KI-Modelle m&#252;ssen innerhalb des Europ&#228;ischen Wirtschaftsraums oder in L&#228;ndern mit angemessenem Datenschutzniveau gem&#228;&#223; Art. 45 DSGVO entwickelt worden sein.&#8221; The award justification: &#8220;lediglich Mistral AI SAS s&#228;mtliche zwingenden Mindestanforderungen kumulativ erf&#252;llen kann&#8221;. The voluntary ex-ante transparency notice is <a href="https://ted.europa.eu/en/notice/-/detail/335597-2026">TED 335597-2026</a>, May 18th, 2026. All three notices carry the eForms placeholder award date of January 1st, 2000.</p><p>[10] Mistral AI, the <a href="https://mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/">Series D announcement</a>, September 8th, 2026, as note 5: &#8220;Advent, funds and accounts managed by BlackRock as well as the Grand Duchy of Luxembourg joined the round as new investors.&#8221;</p><p>[11] The CTIE reports to Luxembourg&#8217;s digitalisation ministry; the Direction de la D&#233;fense sits inside the foreign affairs ministry; the Series D names &#8220;the Grand Duchy&#8221; with no vehicle. Nothing in notices 415226-2025 or 492208-2025 refers to any equity arrangement, and the funding announcement does not refer to any contract.</p><p>[12] Author&#8217;s search, September 22nd, 2026: BOAMP open data (1,708,130 records), SIREN 952418325 in any field, 0; titulaire &#8220;Mistral AI&#8221;, 0. DECP datasets <code>decp-v3-marches-valides</code> (702,901 rows) and <code>decp_augmente</code> (994,123 rows), SIRET prefix 952418325, 0. TED, winner name &#8220;Mistral AI&#8221;, 3, the three notices above. The French ministries contract: Les Echos, September 2nd, 2026, confirmed by the minist&#232;re de l&#8217;Action et des Comptes publics to Capital, as relayed by <a href="https://www.actuia.com/actualite/mistral-ai-dans-les-ministeres-ce-que-letat-achete-pour-6-meur/">ActuIA</a>, September 7th, 2026: &#8220;six millions d&#8217;euros sur la p&#233;riode 2026-2027, un contrat sans caract&#232;re exclusif, des ing&#233;nieurs de l&#8217;entreprise d&#233;tach&#233;s au sein des administrations&#8221;; &#8220;L&#8217;accord a &#233;t&#233; conclu &#224; l&#8217;issue d&#8217;une proc&#233;dure de march&#233; public&#8221;, the ministry&#8217;s detail as reported by Generation-NT, per the same ActuIA article. No notice in BOAMP, DECP or TED on September 22nd, 2026.</p><p>[13] Minist&#232;re des Arm&#233;es et des Anciens combattants, <a href="https://www.defense.gouv.fr/sites/default/files/ministere-armees/Communiqu%C3%A9_A_le%20minist%C3%A8re%20des%20Arm%C3%A9es%20et%20des%20Anciens%20combattants%20notifie%20un%20accord-cadre%20%C3%A0%20Mistral%20AI%20pour%20renforcer%20la%20souverainet%C3%A9%20technologique%20de%20la%20d%C3%A9fense.pdf">communiqu&#233;</a> of January 8th, 2026, on the accord-cadre between AMIAD and Mistral AI notified December 16th, 2025: no montant, duration or ceiling stated; no BOAMP or TED notice, the defence exemption applying. Inspection g&#233;n&#233;rale des finances, IGAS and IGA, joint report on AI deployment in the state, IGF 2025-E-081-03, April 2026, published July 2nd, 2026, at <a href="https://www.vie-publique.fr/rapport/303927">vie-publique</a>: annexe XII, tableau 3, DINUM commitments 2023 to 2025, &#8220;Assistant IA 338 840 &#8364;&#8221;; footnote 48 of the main report, verbatim, &#8220;De l&#8217;ordre de 300 k&#8364; pour 8 GPU et 10 000 &#233;quivalents utilisateurs actifs (pour un taux d&#8217;utilisation observ&#233; de 10 %, le contrat ouvre un acc&#232;s &#224; pr&#232;s de 100 000 agents) sur 12 mois. En cas de g&#233;n&#233;ralisation de l&#8217;exp&#233;rimentation, le passage &#224; 1 &#224; 2 millions d&#8217;utilisateurs [...] correspondrait &#224; un co&#251;t de l&#8217;ordre de 3 M&#8364; par an &#8211; hors ren&#233;gociations [...] pour le fournisseur actuel (MistralAI)&#8221;. These are the inspectors&#8217; estimates, not a contract price. The often-repeated &#8220;two million agents&#8221; is the upper bound of that estimate, not a headcount in any signed agreement; DINUM&#8217;s own <a href="https://ia.numerique.gouv.fr/actualit%C3%A9s/lancement-de-lexp%C3%A9rimentation-mistral-ai-dans-lassistant-ia-interminist%C3%A9riel/">page</a>, published November 17th, 2025, on the launch of October 22nd, 2025, gives 10,000 agents across eight ministries on Mistral Medium 3, hosted by Outscale, with no amount.</p><p>[14] Assembl&#233;e nationale, commission d&#8217;enqu&#234;te sur les d&#233;pendances structurelles et les vuln&#233;rabilit&#233;s syst&#233;miques dans le secteur du num&#233;rique, <a href="https://www.assemblee-nationale.fr/dyn/opendata/CRCANR5L17S2026PO878304N042.html">compte rendu n&#176; 42</a>, May 12th, 2026, hearing of Arthur Mensch under oath, 10,547 words. Verbatim: &#8220;les services &#224; haute valeur ajout&#233;e, c&#8217;est-&#224;-dire ceux qui offrent une forte marge donc permettent de financer de la R&amp;D, c&#8217;est l&#8217;intelligence artificielle. &#192; partir du moment o&#249; on d&#233;veloppe l&#8217;intelligence artificielle, on peut construire tout le reste des services du cloud, qui sont surtout des commodit&#233;s.&#8221; Occurrences of &#8220;commodit&#8221; in the transcript: one. The revenue split, same hearing, verbatim: &#8220;De t&#234;te, je dirais que la commande publique globale repr&#233;sente 20 % de notre chiffre d&#8217;affaires, sur la partie logiciels, et la commande publique fran&#231;aise 10 %. Nous ne cherchons pas &#224; la faire beaucoup augmenter, mais tout cro&#238;t. Nous avons des contrats-cadres significatifs avec le Luxembourg par exemple, o&#249; nous faisons du d&#233;ploiement en administration centrale.&#8221; The verbatim annex is reproduced in <a href="https://www.assemblee-nationale.fr/dyn/opendata/RAPPANR5L17B3054-t2.html">rapport n&#176; 3054</a>, tome 2, July 8th, 2026.</p><p>[15] Bitkom, <a href="https://www.bitkom.org/Presse/Presseinformation/KI-aus-China-fuer-deutsche-Wirtschaft-kein-Thema">&#8220;KI aus China? F&#252;r die deutsche Wirtschaft praktisch kein Thema&#8221;</a>, September 9th, 2026. Base sentence verbatim: &#8220;Basis der genannten Werte sind Unternehmen, die KI einsetzen.&#8221; Question: &#8220;Welche der folgenden KI-Anwendungen nutzen Sie im Unternehmen?&#8221;, multiple answers. Rows not in the body: Llama 8 per cent, DeepSeek 2, Qwen 1. Bitkom&#8217;s president, Ralf Wintergerst, in the same release: &#8220;Wer digitale Souver&#228;nit&#228;t will, muss europ&#228;ischen Anbietern eine Chance geben.&#8221; The n of 343 AI-using firms is in the <a href="https://www.bitkom.org/sites/main/files/2026-09/bitkom-praesentation-kuenstliche-intelligenz-in-der-wirtschaft-2026.pdf">conference presentation</a> of September 14th, 2026. Methodology in the companion release <a href="https://www.bitkom.org/Presse/Presseinformation/Erstmals-nutzt-Mehrheit-Unternehmen-KI">&#8220;Erstmals nutzt die Mehrheit der Unternehmen KI&#8221;</a>: 603 firms of 20+ employees, telephone, calendar weeks 28 to 33 of 2026, representative. Independent confirmation: <a href="https://www.heise.de/news/US-KI-dominiert-in-Firmen-China-Modelle-kaum-genutzt-EU-KI-spielt-keine-Rolle-11446802.html">heise</a>, same day. Bitkom&#8217;s chart prints the 2025 value beside each 2026 row; Mistral was on the prompted list in 2025 and read 0 in both years. Aleph Alpha is not on the eleven-name prompted list in either year, so it was not asked about rather than scored.</p><p>[16] SAP as a Mistral channel: SAP News, <a href="https://news.sap.com/2025/11/sap-mistral-ai-new-alliance-european-sovereign-ai/">&#8220;SAP and Mistral AI announce new alliance for European sovereign AI&#8221;</a>, November 18th, 2025: &#8220;The next phase will see SAP providing Mistral AI&#8217;s frontier models and products through a sovereign AI foundation on SAP Business Technology Platform (SAP BTP)&#8221;; Mistral&#8217;s own <a href="https://mistral.ai/customers/sap">customer page for SAP</a>: &#8220;the SAP team built a multilingual chatbot leveraging Mistral Large hosted entirely on the SAP operated infrastructure&#8221;. The &#8220;125+ global enterprises&#8221; figure is Mistral&#8217;s own, in the announcement of note 5. The companion <a href="https://www.bitkom.org/Presse/Presseinformation/Erstmals-nutzt-Mehrheit-Unternehmen-KI">Bitkom release</a> gives token consumption as 8 per cent of AI cost among users, against 51 per cent for infrastructure and compute, with Wintergerst&#8217;s line &#8220;Bei KI geht es nicht in erster Linie um das eingesetzte Modell.&#8221;</p><p>[17] The Hugging Face Hub API record for <a href="https://huggingface.co/api/models/mistralai/Mistral-Large-3-675B-Instruct-2512">Mistral-Large-3-675B-Instruct-2512</a>, read September 22nd, 2026, 05:23 UTC: <code>downloads</code> 2,097 (trailing thirty days), <code>downloadsAllTime</code> 13,457, created November 28th, 2025, licence Apache 2.0. Sibling repositories in the same call, thirty-day <code>downloads</code>: Mistral-Large-3-675B-Instruct-2512-NVFP4 10,798; -Base-2512 75; -Instruct-2512-BF16 41; -Instruct-2512-Eagle 119; family total 13,130. Author&#8217;s calculation: 1,118,995 &#247; 13,130 &#8776; 85.2. The <code>downloads</code> field is a rolling thirty-day count; a re-read at 17:03 UTC the same day gave 2,125 and 13,504, DeepSeek-V3 1,148,200, ratios 540 and 84.9, so the 05:23 figures are one dated measurement and the conclusions do not move. Mistral-Medium-3.5-128B, same counter, 89,194 thirty-day pulls, created March 31st, 2026, Hub licence field &#8220;other&#8221;; Mistral&#8217;s <a href="https://docs.mistral.ai/getting-started/changelog">changelog</a>: &#8220;April 28 We released Mistral Medium 3.5 (mistral-medium-3-5) [...] Released as open weights under a Modified MIT license.&#8221; Archived org page: <a href="http://web.archive.org/web/20260921182241/https://huggingface.co/mistralai">Wayback 20260921182241</a>.</p><p>[18] The same API record for <a href="https://huggingface.co/api/models/deepseek-ai/DeepSeek-V3">DeepSeek-V3</a>, same call: <code>downloads</code> 1,118,995, <code>downloadsAllTime</code>21,044,555, created December 25th, 2024. The model&#8217;s size and licence class are from DeepSeek&#8217;s own <a href="https://arxiv.org/abs/2412.19437">technical report</a>, December 26th, 2024: 671 billion total parameters, 37 billion active, mixture of experts. Licence from the repository, not the paper: README section 7, &#8220;The use of DeepSeek-V3 Base/Chat models is subject to the Model License. DeepSeek-V3 series (including Base and Chat) supports commercial use&#8221;; the DeepSeek Model License carries use-based restrictions and is not Apache-class. Author&#8217;s calculation: 1,118,995 &#247; 2,097 &#8776; 534x.</p><p>[19] Peter Walker, OpenRouter, <a href="https://x.com/PeterJ_Walker/status/2102185973367116025">post on X</a>, September 21st, 2026, 23:59 UTC: &#8220;I know it feels forever ago, but Mistral really did have a moment there.&#8221;, with the chart &#8220;Token share by model lab country&#8221;, &#8220;Share of tokens used on OpenRouter by week&#8221;; the bars run from January 2024 to September 2026 although the chart&#8217;s own caption says &#8220;Jan 1 2025-Sep 20 2026&#8221;. France, read from the chart&#8217;s labels: 62, 68, 57, 46, 36, 23, 14, 15, 10, 10, 7 and 7 per cent for the twelve months of 2024; 5 to 6 per cent in the first months of 2025; 5 per cent in December 2025; no label since. September 2026: China 66 per cent, USA 32. OpenRouter&#8217;s own data, read from the image; the platform counts tokens routed through one marketplace and does not see Mistral&#8217;s direct API or on-premises deployments. Live <a href="https://openrouter.ai/rankings">ranking</a>, read September 23rd, 2026, 05:50 UTC, &#8220;Model authors ranked by their share of text requests on OpenRouter in the most recent complete week&#8221;: deepseek 25.4, google 18.6, openai 17.0, z-ai 9.4, qwen 6.7, tencent 6.4, anthropic 2.7, mistralai 2.6, xiaomi 1.9 per cent. Desk check, September 23rd, 2026: OpenRouter&#8217;s own chart endpoint (<code>openrouter.ai/api/frontend/v1/rankings/model-rankings-chart</code>) returns 52 weeks of the nine top models by tokens plus an &#8220;Others&#8221; bucket, September 29th, 2025 to September 21st, 2026; no Mistral model appears in any week, and Chinese labs&#8217; share of the named tokens rises from 2 to 6 per cent in October and November 2025 to 55 per cent in the week of September 21st, 2026, consistent with the chart&#8217;s right-hand side. The 2024 bars could not be reproduced: the endpoint covers one year, the Internet Archive was offline and archive.today&#8217;s first snapshot of the rankings page is from March 2025. The 2024 figures are the chart&#8217;s alone. Payload kept.</p><p>[20] Author&#8217;s test, August 17th, 2026, 15:48 local time, on Mistral Vibe (the <code>mistral-vibe</code> command-line agent), model Mistral Medium 3.5; Vibe&#8217;s own session log, kept locally, is the record. Verbatim from that log: &#8220;Mistral <strong>does not offer a speech-to-text (STT) model</strong> &#8212; their API is exclusively for text generation (LLMs). Mistral has no audio processing capabilities.&#8221; When told of Voxtral in the next turn, the agent replied &#8220;You&#8217;re absolutely right &#8212; <strong>Voxtral</strong> is Mistral&#8217;s speech-to-text model&#8221; and wrote the pipeline against <code>voxtral-mini-latest</code>. The session&#8217;s tool list included <code>web_search</code>, which the agent called four times later in the same session and not once before this answer. Re-run September 22nd, 2026, 15:23 UTC, Vibe 2.25.7 in programmatic mode, one turn, no web-search or web-fetch tool (the catalogue in the kept journal offers file, shell, code and skill tools only), active model <code>mistral-medium-3.5</code> per the session journal; answer verbatim: &#8220;Mistral AI does not currently offer a dedicated speech&#8209;to&#8209;text model. Their public API serves text generation via endpoints like <code>/v1/chat/completions</code>. For transcription, use a separate ASR service such as Whisper or a cloud provider.&#8221; A request for GLM-5.3 from the free tier timed out; my own check found it not offered there. Two smaller findings from the same session: the console detected a French browser and served the interface in French with documentation links to the English docs; the free tier&#8217;s limits were hard to find. Mistral Medium 3.5: earlier dates circulate for the model&#8217;s first appearance; the piece uses Mistral&#8217;s own post putting it in Vibe on May 22nd, 2026. On August 17th the documented install script answered 429 to the first request. Install path as documented at <a href="https://docs.mistral.ai/vibe">docs.mistral.ai/vibe</a>; re-measured September 22nd, 2026, 13:17 UTC: <code>mistral.ai/vibe/install.sh</code> answers 301 to <a href="https://raw.githubusercontent.com/mistralai/mistral-vibe/refs/heads/main/scripts/install.sh">the script on GitHub</a>, which answers 200. Product dates from <code>datePublished</code> on Mistral&#8217;s own pages: <a href="https://mistral.ai/news/voxtral">Voxtral</a>, July 15th, 2025; <a href="https://mistral.ai/news/voxtral-transcribe-2">Voxtral Transcribe 2</a>, February 4th, 2026; <a href="https://mistral.ai/news/voxtral-tts">Voxtral TTS</a>, March 23rd, 2026; <a href="https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5">&#8220;Remote agents in Vibe. Powered by Mistral Medium 3.5&#8221;</a>, May 22nd, 2026; GLM-5.2 hosting in <a href="https://mistral.ai/news/regional-inference-open-models-new-compute">&#8220;Regional inference for open models on new compute&#8221;</a>, August 11th, 2026. <code>docs.mistral.ai/llms.txt</code> exists and indexes the documentation. The August answer is my testimony, with Vibe&#8217;s session log as its record. The web app&#8217;s paid-tier default I have not seen; the &#8220;default in the web app&#8221; claim of note 2 rests on the product account alone.</p><p>[21] Tim Smith and Amy Lewin, <a href="https://sifted.eu/articles/pitch-deck-mistral">&#8220;See the pitch memo that raised &#8364;105m for four-week-old startup Mistral&#8221;</a>, Sifted, June 21st, 2023. The document, titled &#8220;mistral.ai strategic memo&#8221;, seven pages, is embedded and <a href="https://drive.google.com/uc?export=download&amp;id=1gquqRqiT-2Be85p_5w0izGQGgHvVzncQ">downloadable</a>; read in full on September 22nd, 2026. Les Echos, &#8220;La saga Mistral (2/3)&#8221;, September 17th, 2026, quotes the same two sentences from a copy &#8220;que &#171; Les Echos &#187; se sont procur&#233;&#8221; and renders the passage in its own translation.</p><p>[22] The <a href="https://drive.google.com/uc?export=download&amp;id=1gquqRqiT-2Be85p_5w0izGQGgHvVzncQ">memo</a>, page 5, &#8220;Infrastructure and data sources&#8221;, verbatim as quoted; published and described by <a href="https://sifted.eu/articles/pitch-deck-mistral">Sifted</a> on June 21st, 2023, which is the independent record that this is the document the founders circulated. The starting cluster, same page: 1,536 H100s from September 2023 &#8220;with a summer ramp up&#8221;. The efficiency sentence in full: &#8220;Having trained models at large-scale before has provided us know-hows that will allow us to gain a factor 10-100 in training efficiency compared to public methods &#8212; our founders and early employees know exactly what to do to train the strongest model for a given computational budget.&#8221;</p><p>[23] Memo, page 3 for the open-source hedge and page 5 for the two sentences on investors&#8217; doors and industrial actors, verbatim as quoted.</p><p>[24] Big Technology Podcast, <a href="https://www.youtube.com/watch?v=xxUTdyEDpbU">&#8220;Who Wins if AI Models Commoditize? &#8212; With Mistral CEO Arthur Mensch&#8221;</a>, audio published January 14th, 2026, YouTube upload January 16th. Timestamps on the YouTube version, whose duration matches the podcast feed&#8217;s ad-free audio at 3,268 seconds: host at 01:58, &#8220;we&#8217;re just hitting commoditization of the foundational model much faster than I thought it would be&#8221;; Mensch at 02:26 onward as quoted; Mensch at 17:09, &#8220;If AI effectively becomes a commodity, which is what&#8217;s happening&#8221;; Mensch at 23:01, &#8220;to get to value, they need to have great models [...] So the two things are extremely linked together.&#8221; Transcribed by the author with whisper.cpp large-v3-turbo on September 21st, 2026, because no human caption track or official transcript exists. Note that Mensch adopts the host&#8217;s word and supplies the reasoning; the episode title is the host&#8217;s. A week later at Davos he told Axios that Mistral&#8217;s policy assumed &#8220;the models are essentially free&#8221;: Axios, <a href="https://www.youtube.com/watch?v=iIOuzEb_2kE">&#8220;Mistral AI&#8217;s Arthur Mensch &amp; Axios&#8217; Ina Fried&#8221;</a>, recorded at Axios House, Davos, during the week of January 19th, 2026, uploaded January 22nd; at 02:22, &#8220;if you have a policy where you assume that the models is not where the value accrues but that the models are essentially free&#8221;; at 03:35, &#8220;what we&#8217;ve essentially been saying from the very beginning which is that the model layer is a free layer&#8221;. Same transcription method. Axios&#8217;s own <a href="https://www.axios.com/2026/01/26/axios-house-davos-2026-ai-adoption-business-mistral-ceo">write-up</a> of January 26th quotes the passage from 03:06 word for word, which confirms the transcription in that window.</p><p>[25] Author&#8217;s sweep of Wayback snapshots of <code>mistral.ai</code> on May 5th and June 13th, 2023; October 1st and December 28th, 2023; March 1st, 2024; June 1st, 2025; and January 1st, June 1st and September 1st, 2026. Occurrences of &#8220;full stack&#8221; in any snapshot: zero. The phrase appears in the announcement of note 5. The word &#8220;sovereign&#8221; first appears on the home page in the September 1st, 2025 snapshot; &#8220;compute&#8221; between the May 1st and July 1st, 2025 snapshots.</p><p>[26] Mistral documentation, <code>zai-glm-5-3</code>, dated September 15th, 2026, badges &#8220;Public Preview&#8221;, &#8220;Third-party&#8221;, &#8220;Open&#8221;; &#8220;hosted by Mistral&#8221;; &#8220;served without Mistral modifications&#8221;; $1.40 input, $0.14 cached input, $4.40 output per million tokens; EU and global endpoints, no US endpoint. Z.ai&#8217;s own <a href="https://docs.z.ai/guides/overview/pricing">pricing</a>: $1.40 input, $0.26 cached, $4.40 output. <code>mistral.ai/news</code> (about 1.26 MB), <code>mistral.ai/pricing</code> (487 KB) and <code>mistral.ai/products/vibe</code> (446 KB) read September 22nd, 2026: zero occurrences of &#8220;GLM&#8221;, &#8220;Z.ai&#8221; or &#8220;Zhipu&#8221; in the pages&#8217; text (one image filename on the news page carries &#8220;GLM&#8221;). Artificial Analysis index version 4.3.2, GLM-5.3 scored at maximum reasoning, Mistral Large 3 without a reasoning mode. Artificial Analysis, <a href="https://artificialanalysis.ai/providers/mistral">provider page</a>, read September 21st, 2026: GLM-5.2 (max) 34, Mistral Medium 3.5 14, Mistral Large 3 9 (non-reasoning); GLM-5.3 (max) 45 on the leaderboard the same minute; frontier top 53.</p><p>[27] The exchange, seven posts in French, six distinct texts (five by Branco, two by Mensch), September 18th to 20th, 2026, read September 21st: Branco&#8217;s anchor, <a href="https://x.com/anatolium/status/2101017370261414207">@anatolium</a>, September 18th, 18:35:50 UTC, itself a quote of an anonymous account&#8217;s anecdote about Mistral&#8217;s teams; Mensch, <a href="https://x.com/arthurmensch/status/2101703285291770068">@arthurmensch</a>, September 20th, 16:01:25 UTC, verbatim &#8220;Nos travaux d&#233;plaisent &#224; certaines petites frappes. Celui-l&#224; est candidat &#224; l&#8217;&#233;lection pr&#233;sidentielle &#128517;&#8221;; Branco&#8217;s reply 43 seconds later; Mensch, <a href="https://x.com/arthurmensch/status/2101706859069509740">second post</a>, 16:15:37 UTC, verbatim &#8220;1 minute pour me r&#233;pondre, c&#8217;est un obsessionnel&#8221;. Both of Mensch&#8217;s posts were deleted between that reading and September 22nd, 2026, 16:20 UTC, when X&#8217;s syndication endpoint (<code>cdn.syndication.twimg.com/tweet-result?id=&lt;id&gt;&amp;token=a</code>) returned &#8220;This Post was deleted by the Post author&#8221; for both identifiers and a live post for Branco&#8217;s anchor; the two links above are dead. No archive ever captured them: the Internet Archive refused every save on September 21st, and the copy taken that day is the only record. View counters at the September 21st reading, frozen by the deletion: Mensch&#8217;s two posts 1,541,771 and 399,869, total 1,941,640; Branco&#8217;s anchor 1,310,275; Branco&#8217;s five posts together 1,929,477. Branco&#8217;s rebuttal names &#8220;des resuc&#233;es de GLM 5.2 pendant tout un &#233;t&#233;&#8221;; GLM-5.2 hosting is Mistral&#8217;s own announcement of August 11th, 2026; GLM-5.3&#8217;s documentation page is dated September 15th (note 26), and GLM-5.3 in Vibe Code is note 2. Branco&#8217;s descriptors: French Wikipedia&#8217;s <a href="https://fr.wikipedia.org/wiki/Juan_Branco">lead</a>, read September 22nd, 2026, &#8220;avocat, essayiste, militant politique d&#8217;extr&#234;me gauche et r&#233;alisateur franco-espagnol&#8221;, sworn in as a lawyer in April 2017, and his announcement in December 2025 of his 2027 candidacy, as summarised there. Cr&#233;puscule: published online December 2018, then Paris, Au diable vauvert, March 21st, 2019, 312 pages, ISBN 979-10-307-0260-6. C&#233;dric O&#8217;s shareholding through his company Nopeunteo is on the parliamentary record (compte rendu n&#176; 45, May 13th, 2026, the chair&#8217;s opening words); its size is not. The &#8364;90 million, the under-&#8364;200 entry, EPFL&#8217;s cost, the engineering headcount and the share of public orders carry no document and are not used here.</p><p>[28] CMA CGM, <a href="https://www.cmacgm-group.com/en/news-media/cma-cgm-group-adopts-custom-designed-ai-solutions-mistral-ai">&#8220;CMA CGM Group adopts custom-designed AI solutions with Mistral AI&#8221;</a>, April 6th, 2025: &#8220;From the moment Mistral AI secured its first round of funding in June 2023, the CMA CGM Group took the initiative to invest in this pioneering startup&#8221;. Series D investor list as note 5.</p><p>[29] <code>mistral.ai/customers/</code>, 42 named organisations, pages dated April 24th to July 28th, 2026, read September 22nd, 2026; plus Amadeus, Caisse des D&#233;p&#244;ts, Microsoft, Nvidia, HUMAIN, Cloudera, Mozilla, AFP and Land Hessen from Mistral&#8217;s own announcements, Samsung Electronics from its own release of September 9th, 2026 (note 5), and the European Space Agency, <a href="https://www.esa.int/Newsroom/Press_Releases/ESA_and_Mistral_strengthen_cooperation_on_artificial_intelligence">press release</a>, numbered N&#176; 57-2026 in ESA&#8217;s site search, September 23rd, 2026: &#8220;The European Space Agency (ESA) and Mistral have signed a Letter of Intent establishing a strategic framework to deepen their collaboration on artificial intelligence&#8221;, &#8220;signed in Paris on 16 September by ESA Director General Josef Aschbacher and Mistral CEO Arthur Mensch&#8221;; no amount, no term. Not on mistral.ai/news on September 24th, 2026, 07:30 CEST.</p><p>[30] Author&#8217;s count from notes 28 and 29: cap table 24, commercial 53; both lists 7 (ASML, CMA CGM, Grand Duchy of Luxembourg, BNP Paribas CIB, Belfius, Nvidia, Samsung Electronics); on the customers page and the cap table both, 5 of 42 (ASML, Belfius, BNP Paribas, CMA CGM, Government of Luxembourg); customers with no equity link in any document read, 37 of 42. Documents read for the five largest: Airbus Board report FY2025 p.43 and financial statements; HSBC Annual Report and Accounts 2025; Capgemini 2025 URD; Amadeus Global Report 2025, FY2025 accounts and H1 2026 results, zero mentions of Mistral; Caisse des D&#233;p&#244;ts Rapport d&#8217;activit&#233; 2025, &#8220;investissements strat&#233;giques (Mistral IA)&#8221;, and RAFI 2025, zero mentions. SNCF: own <a href="https://www.groupe-sncf.com/medias-publics/2025-06/groupe-sncf-et-mistral-ai.pdf">release</a> of June 12th, 2025, and Rapport annuel int&#233;gr&#233; 2025, one mention each, no amount; Rapport financier annuel 2025, zero mentions. Caisse des D&#233;p&#244;ts&#8217;s dedicated <a href="https://www.caissedesdepots.fr/sites/cdc.fr/files/2026-05/2026_05_04_CP_Le_groupe_Caisse_des_D%C3%A9p%C3%B4ts_sengage_avec_Mistral_AI.pdf">release</a> of May 4th, 2026 describes a two-lot framework agreement, generative AI and GPU compute, &#8220;pr&#232;s de 40000 licences&#8221; at the start and up to 100,000 users, 19 subsidiaries in the buying group, and no price. Were CDC confirmed as a shareholder the overlap would be 8 of 25, and the non-overlap would still be the larger number. Salesforce Ventures is an existing investor per note 5; no Salesforce entity appears on any customer or partner page.</p><p>[31] ASML, <a href="https://www.asml.com/en/news/press-releases/2025/asml-mistral-ai-enter-strategic-partnership">press release</a>, September 9th, 2025: &#8220;ASML is investing 1.3 billion EUR in Mistral AI&#8217;s Series C funding round as lead investor&#8221;; &#8220;approximately 11 percent share on a fully diluted basis&#8221;; &#8220;a seat on the Strategic Committee of Mistral AI [...] an advisory role&#8221;; &#8220;a long-term collaboration agreement to explore the use of AI models&#8221;. ASML Form 20-F for 2025, filed February 25th, 2026, <code>asml-20251231.htm</code>, XBRL member <code>asml:MistralAIMember</code>: additions &#8364;1,302.2 million, carrying value &#8364;1,320.7 million, 11.1 per cent fully diluted. Anchor tenant per Mistral&#8217;s <a href="https://mistral.ai/news/regional-inference-open-models-new-compute/">August 11th, 2026 announcement</a>.</p><p>[32] CMA CGM, April 6th, 2025, as note 28: &#8220;a five-year strategic partnership&#8221; and &#8220;Supported by a &#8364;100 million investment, this partnership marks a significant milestone&#8221;. Repeated in CMA CGM&#8217;s <a href="https://www.cmacgm-group.com/en/news-media/solid-group-continuing-its-transformation-unstable-market-environment">Q1 2025 results</a>, May 16th, 2025, &#8220;coupled with a &#8364;100 million investment&#8221;, and in its <a href="https://www.cmacgm-group.com/en/news-media/annual-financial-results-2025">FY2025 results release</a>, March 6th, 2026; the &#8220;around twenty Mistral AI engineers&#8221; in Marseille is also CMA CGM&#8217;s own wording, AI Now Summit release, May 27th, 2026. The group&#8217;s audited consolidated statements were not read; its site serves them through a client-side endpoint. Occurrences of the figure on <code>mistral.ai</code>: zero. The group&#8217;s total AI commitment, &#8220;EUR 500 million&#8221;, names Google, Perplexity, Poolside and Dataiku alongside Mistral.</p><p>[33] Author&#8217;s count, September 22nd, 2026, recounted September 24th, 2026 with Samsung and the ESA letter of intent: 56 announced relationships, swept from Mistral&#8217;s 137 sitemap pages, SEC full-text search, counterparty releases and French registries, counting commercial deals, equity and partnerships together, so not the sum of note 29&#8217;s 53 names and the cap table; with any published value, 10, eleven figures, Microsoft counted under bonds and under compute; three of the eleven are words rather than numbers (Pimento&#8217;s &#8220;plusieurs millions&#8221;, HUMAIN&#8217;s &#8220;in the hundreds of millions&#8221;, Microsoft&#8217;s &#8220;multibillion dollar commitment&#8221;). Procurement, exact: the three notices of notes 7 to 9. Equity, exact: ASML&#8217;s stake of note 31; Microsoft&#8217;s &#8364;15,000,000 of convertible bonds, &#8220;part of a bond raise of &#8364;120 million&#8221;, with Microsoft to &#8220;own less than 1%&#8221;, in the UK Competition and Markets Authority&#8217;s <a href="https://assets.publishing.service.gov.uk/media/664c6cfd993111924d9d389f/Full_text_decision.pdf">Phase 1 decision</a> of May 17th, 2024, which prints the parties&#8217; submission; neither party&#8217;s own release carries a figure; TechCrunch reported it the day after the announcement, <a href="https://techcrunch.com/2024/02/27/microsoft-made-a-16-million-investment-in-mistral-ai/">&#8220;Microsoft made a $16M investment in Mistral AI&#8221;</a>, February 27th, 2024: &#8220;Microsoft is investing &#8364;15 million ($16.3 million at today&#8217;s exchange rate)&#8221;, and the CMA decision cites that article. Private commercial, party-published: TotalEnergies, own <a href="https://totalenergies.com/newsroom/totalenergies-annonce-un-partenariat-avec-mistral-en-vue-de-developper-des-modeles-de-frontiere-dintelligence-artificielle-dedies-a-lexploration-et-a-lingenierie-des-reservoirs-498514">release</a>, Paris, September 15th, 2026: &#8220;a 3-year joint program representing an investment of more than &#8364;100 million aimed at developing a new generation of frontier AI models&#8221;; CMA CGM of note 32; HUMAIN, &#8220;in the hundreds of millions of Euros&#8221;, Mistral&#8217;s <a href="https://mistral.ai/news/mistral-x-humain/">own announcement</a>, August 24th, 2026; Microsoft, <a href="https://news.microsoft.com/source/2026/07/21/microsoft-and-mistral-expand-strategic-partnership-to-give-enterprises-and-regulated-industries-frontier-ai-they-can-control/">&#8220;a multibillion dollar commitment&#8221;</a>, Microsoft release of July 21st, 2026, Mensch quoted, no figure or term; Microsoft&#8217;s FY2026 10-K, filed July 29th, 2026, contains no occurrence of &#8220;Mistral&#8221;. Acquisitions: Pimento, Mistral&#8217;s own words to AFP, <a href="https://fr.tradingview.com/news/afp:c37d027917a52:0/">dispatch</a> of September 21st, 2026: Mistral &#8220;va racheter &#8220;pour plusieurs millions d&#8217;euros&#8221; la start-up Pimento [...], a-t-elle indiqu&#233; lundi &#224; l&#8217;AFP, confirmant une information du m&#233;dia L&#8217;Inform&#233;&#8221;; press only, reported and not confirmed: Emmi AI, &#8220;up to &#8364;330M&#8221; per Sifted, June 11th, 2026, from leaked internal documents.</p><p>[34] The Hub API record for <a href="https://huggingface.co/api/models/mistralai/Voxtral-Mini-4B-Realtime-2602">Voxtral-Mini-4B-Realtime-2602</a>, same call as note 17: <code>downloads</code> 1,746,931, <code>downloadsAllTime</code> 11,430,456, created January 21st, 2026; third in the account by thirty-day pulls after the two 7B instruct models of note 35.</p><p>[35] The same call, records for <a href="https://huggingface.co/api/models/mistralai/Mistral-7B-Instruct-v0.3">Mistral-7B-Instruct-v0.3</a> and v0.2: <code>downloads</code> 2,357,073, <code>downloadsAllTime</code> 46,426,509, created May 22nd, 2024; <code>mistralai/Mistral-7B-Instruct-v0.2</code>, <code>downloadsAllTime</code> 64,192,688.</p><p>[36] Mensch, June 16th, 2026, as note 1: &#8220;In domains that are less compute bound, e.g. voice, vision and document processing, we have state-of-the-art solutions.&#8221; Mistral Large 3, Apache 2.0, created on the Hub November 28th, 2025, note 17; the closed Mistral Large shipped February 26th, 2024, note 37. Earlier Mistral-Large-Instruct-2407 and -2411 shipped under Mistral&#8217;s research licence, Hub <code>license: other</code>, July and November 2024.</p><p>[37] Mistral AI, <a href="https://mistral.ai/news/mistral-large">&#8220;Au Large&#8221;</a>, February 26th, 2024, Wayback <a href="https://web.archive.org/web/20240226152230/https://mistral.ai/news/mistral-large/">20240226152230</a>: &#8220;Mistral Large is available through la Plateforme. We are also making it available through Azure, our first distribution partner&#8221;; Mistral Small described as &#8220;a refined intermediary solution between our open-weight offering and our flagship model&#8221;. Microsoft, <a href="https://azure.microsoft.com/en-us/blog/microsoft-and-mistral-ai-announce-new-partnership-to-accelerate-ai-innovation-and-introduce-mistral-large-first-on-azure/">Azure blog</a>, same day.</p><p>[38] The European Parliamentary Research Service <a href="https://www.europarl.europa.eu/thinktank/en/document/EPRS_ATA%282024%29760392">at-a-glance briefing PE 760.392</a>, March 27th, 2024, table 2: &#8220;Mistral AI | Mistral 7B | France | Open source&#8221;. Microsoft, Azure and Mistral Large do not appear in it.</p><p>[39] The Autorit&#233; de la concurrence&#8217;s <a href="https://www.autoritedelaconcurrence.fr/en/opinion/competitive-functioning-generative-artificial-intelligence-sector">avis n&#176; 24-A-05</a>, June 28th, 2024, paragraph 182: &#8220;Mistral AI has released several of its models as open-weights, but not Mistral Large, its most powerful model&#8221; (English version); the opinion records &#8220;the &#8364;15 million investment by Microsoft in Mistral AI in 2024 through a bond convertible into shares&#8221;, sets out the article 3(1) control test and a 1991 precedent, and applies neither to this deal.</p><p>[40] Roland Lescure, September 2nd, 2026, San Francisco, to AFP, as relayed by <a href="https://www.actuia.com/actualite/mistral-ai-dans-les-ministeres-ce-que-letat-achete-pour-6-meur/">ActuIA</a>, September 7th, 2026: &#8220;Si l&#8217;IA europ&#233;enne se r&#233;sume &#224; Mistral, alors nous sommes fichus. On ne peut pas mettre tous nos &#339;ufs dans le m&#234;me panier : c&#8217;est tout un &#233;cosyst&#232;me qu&#8217;il faut d&#233;velopper.&#8221;</p><p>[41] Julien Simon, <a href="https://www.airealist.ai/p/mistral-succeeded-frances-ai-strategy">&#8220;Mistral Succeeded. France&#8217;s AI Strategy Didn&#8217;t.&#8221;</a>, The AI Realist, March 10th, 2026.</p><p>[42] Julien Simon, <a href="https://www.airealist.ai/p/lobby-levy-legislate">&#8220;Lobby, Levy, Legislate&#8221;</a>, The AI Realist, May 22nd, 2026: &#8220;Mistral&#8217;s roughly $400M in annual recurring revenue, with &#8364;1bn targeted for the year&#8221;.</p><p>[43] Cohere&#8217;s own <a href="https://www.prnewswire.com/news-releases/cohere-and-aleph-alpha-sign-agreement-to-become-the-first-transatlantic-sovereign-ai-solution-302880732.html">announcement</a>, PR Newswire, September 16th, 2026: definitive business combination agreement, the combined company operating as Cohere with dual headquarters in Berlin and Toronto, subject to regulatory approvals; the two ministers, Evan Solomon and Karsten Wildberger, are named in the photo caption. Aleph Alpha GmbH, Amtsgericht Mannheim HRB 732804, <a href="https://www.unternehmensregister.de/">Unternehmensregister</a>, read September 22nd, 2026: Jahresabschl&#252;sse for fiscal 2022 (published October 27th, 2023), 2023 (January 30th, 2025) and 2024 (February 2nd, 2026); their contents sit behind a human-verification step and were not read by the author. Aleph Alpha is not among the options in Bitkom&#8217;s questionnaire of note 15.</p>]]></content:encoded></item><item><title><![CDATA[At Least 45 Days]]></title><description><![CDATA[Almost every model on Bedrock gets six months' notice before it can be pulled. The one Washington just accused gets 45 days, and no stated reason.]]></description><link>https://www.airealist.ai/p/at-least-45-days</link><guid isPermaLink="false">https://www.airealist.ai/p/at-least-45-days</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Sun, 20 Sep 2026 10:40:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rcUa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rcUa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rcUa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!rcUa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!rcUa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!rcUa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rcUa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:526443,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/216561969?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rcUa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!rcUa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!rcUa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!rcUa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e672d72-a0ff-4597-a4dc-e9b40cd12be7_1456x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On September 18th, 2026, AWS added Moonshot&#8217;s Kimi K3 to Bedrock at Moonshot&#8217;s own price to the cent: $3 per million input tokens and $15 per million output tokens.[1] Ten days earlier, a joint advisory from the NSA, CISA, and the FBI said that &#8220;Moonshot AI extracted significant Claude Fable 5 data to train its Kimi-K3 model.&#8221;[2] The launch post doesn&#8217;t mention it. I don&#8217;t think AWS did anything wrong by listing the model. The line I care about isn&#8217;t in the launch post. It&#8217;s on the model card, and if you pick models on a managed cloud, it&#8217;s the line you should be reading first.</p><p style="text-align: justify;">Here it is: </p><blockquote><p style="text-align: justify;">&#8220;<em>EOL no sooner than: Not Applicable, at least 45 day EOL Notice will be provided</em>.&#8221;[3] </p></blockquote><p style="text-align: justify;">In plain English: AWS promises to keep K3 running for no minimum period, and to warn you 45 days before it switches it off.</p><h2><strong>What did AWS promise before?</strong></h2><p style="text-align: justify;">A lot more. On September 3rd, Bedrock&#8217;s lifecycle page said that &#8220;once a model launches on Amazon Bedrock, it will remain on Amazon Bedrock for at least 12 months before the EOL date&#8221;, with at least six months&#8217; notice before retirement.[4] The Chinese models already on the shelf got exactly that: Kimi K2.5, launched in January, is promised until January 2027.[5]</p><p style="text-align: justify;">Some time between September 3rd and September 9th, the page split in two. The old text moved to a &#8220;Legacy&#8221; page covering models launched before September 7th. A new page covers everything after. Each model card now carries its own floor, and one of two notice periods: &#8220;There are two Legacy periods: 6 months and 45 days. Most models have a 6-month Legacy period.&#8221;[6] I can find no announcement: nothing in the Bedrock document history. And the new page doesn&#8217;t say who gets 45 days, or why.</p><p style="text-align: justify;">Two models have launched under the new rules. On September 19th, I checked all 128 model cards on Bedrock.[5] OpenAI&#8217;s GPT-6 Astra, launched September 8th, is promised until September 8th, 2027, with six months&#8217; notice.[7] Kimi K3 got no floor and 45 days. It is the only card on the entire shelf with a 45-day notice.</p><p style="text-align: justify;">That matters because K3 is not the only accused model AWS sells. The advisory named six companies, and five of them are already on Bedrock: DeepSeek, Alibaba&#8217;s Qwen, MiniMax, Z.AI, and Moonshot itself. That is nineteen cards, K3 included. The other eighteen all promise a year on the shelf and six months&#8217; notice.[5] They arrived before the rewrite. K3 arrived after it.</p><p style="text-align: justify;">The advisory came out on September 8th, inside that September 3rd to 9th window. The web archive can&#8217;t tell me whether AWS rewrote the page before or after it, and I&#8217;m not going to pretend otherwise.</p><h2><strong>Why write a clause you don&#8217;t need?</strong></h2><p style="text-align: justify;">AWS doesn&#8217;t need a 45-day clause to obey the law. If Moonshot lands on the Entity List, or Congress bans the model for some purpose, the contract gives way to the statute, and AWS pulls the model on whatever schedule the law sets. Congress has already done this once: the Defense Authorization Act signed in December 2025 gave the Pentagon 30 days to remove DeepSeek from its systems, and the Senate&#8217;s bill for next year would add Moonshot to the same list.[8]</p><p style="text-align: justify;">So the 45 days aren&#8217;t about the law. They&#8217;re about everything short of it. A letter from a committee chairman. A big customer whose own contracts rule out the model. Under the old rules, AWS would owe K3 customers at least a year on the shelf, with six months&#8217; warning before the end. Under the new ones, it can leave in about six weeks without breaking a promise.[3][4]</p><p style="text-align: justify;">I don&#8217;t know that AWS set the term with the advisory in mind. But if it did, this is what it would look like: sell the model, keep the exit short, and don&#8217;t explain the rule.</p><p style="text-align: justify;">K3&#8217;s license says that a company running it as a paid service with more than $20 million in revenue over twelve months &#8220;must enter into a separate agreement with Moonshot AI&#8221;.[9] Bedrock is well past $20 million. Either AWS has that agreement, or it falls under the license&#8217;s exemption for Moonshot&#8217;s &#8220;certified inference partners&#8221;. Both are relationships with Moonshot, and neither AWS page specifies which.</p><h2><strong>AWS&#8217;s side of it</strong></h2><p style="text-align: justify;">AWS has hosted Chinese open-weight models since DeepSeek-R1 in January 2025 under standard terms.[10] No US rule forbids a company from deploying K3 in its own business.[11] And AWS says the model runs inside its own walls: &#8220;Your data is processed within the AWS data boundary, is not shared with the model provider, and is not used to train the underlying model.&#8221;[1]</p><p style="text-align: justify;">On September 10th, Anthropic, which has committed more than $100bn to AWS over ten years, reported that Moonshot had &#8220;silently forwarded customer requests to Claude, instead of processing them using Kimi&#8221;.[12] On Bedrock, by AWS&#8217;s account, that can&#8217;t happen: the requests stay inside AWS. For a customer, self-hosting is AWS&#8217;s unaudited answer to the forwarding problem, and it only covers Bedrock: AWS Marketplace also sells Moonshot&#8217;s own Kimi API Platform, run by Moonshot. Moonshot&#8217;s only public statement since the report, on Weibo on September 12th, called rumors about its founder and staff &#8220;pure fabrication&#8221;; I can&#8217;t find one that answers the advisory or the forwarding claim.[12]</p><p style="text-align: justify;">It&#8217;s also possible that 45 days is simply how AWS will treat open-weight models going forward, since the provider isn&#8217;t there to promise support. That&#8217;s a fair reading, and the evidence can&#8217;t rule it out yet. Nor can it rule out a third: Astra arrived under AWS&#8217;s partnership with OpenAI, and K3 under nothing comparable.[7]</p><p style="text-align: justify;">What makes me doubt the open-weight explanation is the comparison. Microsoft runs the same kind of shelf and prints the terms in a table. Its other outside-lab models run for twelve months; the notice on a generally available model is at least 60 days; and the exit for trouble is written out, with its reason: if a model &#8220;is found to have compliance or security issues, Microsoft reserves the right to invoke an emergency retirement with shortened notice&#8221;.[13] Microsoft sells Kimi as well, through Fireworks, and that tier is harsher than AWS&#8217;s, not softer: 15 days&#8217; notice, with one retirement date printed on every Fireworks model page, July 1st, 2027.[14] </p><p style="text-align: justify;">Microsoft prints the rule, the date and the reason before you build. AWS gave one model a shorter exit than anything else on its shelf and published none of the three. Amazon&#8217;s 2024 letter to shareholders ends: &#8220;It remains Day One.&#8221;[15] When Microsoft is the one saying things plainly, is it still day one?</p><h2><strong>The meter, again</strong></h2><p style="text-align: justify;">In July I argued that AWS has decided to own the meter, not the model: it sells Claude at Anthropic&#8217;s price and makes its money underneath.[16] K3 at Moonshot&#8217;s price is the same move. What&#8217;s new is that the landlord has started writing different leases for different tenants, and setting the shorter one on itself, in advance, on a page most customers never open.</p><h2><strong>What to do about it</strong></h2><p style="text-align: justify;">If you&#8217;re building on Bedrock, add one line to your model selection: open the card, read &#8220;EOL no sooner than&#8221; and &#8220;Legacy period&#8221;, and write both into your design document next to the price. If you hold an Enterprise Discount Program or private pricing deal, check whether it overrides this page. If the answer is 45 days, ask yourself one question: could I move this workload in six weeks?</p><p style="text-align: justify;">For K3, the answer can be yes, because the weights are public. You can download them today, test them on a second host (your own GPUs, another provider), and keep that path warm. Keeping it warm isn&#8217;t free: you pay for a host you may never use. The 45 days is then a migration window, not a cliff. For a closed model on a 45-day term, you&#8217;d have no such option.</p><p style="text-align: justify;">What would prove me wrong? The next Western open-weight model to launch on Bedrock showing up with 45 days. Then this is a rule about open weights, not about China. Or the next Chinese model getting six months.</p><h2><strong>The house style</strong></h2><p style="text-align: justify;">On March 1st, an AWS availability zone in the Emirates caught fire. The incident report said it had been &#8220;impacted by objects that struck the data center, creating sparks and fire&#8221;. Objects. About a day and a half later, with the attack on the UAE leading the news, AWS named them: two of its buildings there had been &#8220;directly struck by drones&#8221;.[17] Then it went quiet again. Three weeks on, when it waived the region&#8217;s bills, the email gave no cause at all.[18]</p><p style="text-align: justify;">This story is written in the same hand. The policy changed with no announcement. The short tier has no published criteria. AWS has said nothing about what the advisory means for Moonshot. None of it is a lie. None of it is a statement either. Everything is written down precisely in Seattle, in a place where nobody will read it.</p><p>So read it. It&#8217;s one line on the model card.</p><h3><strong>Notes</strong></h3><p>[1] AWS, &#8220;<a href="https://aws.amazon.com/blogs/machine-learning/introducing-kimi-k3-on-amazon-bedrock/">Introducing Kimi K3 on Amazon Bedrock</a>&#8220;, September 18th, 2026, accessed September 19th, 2026: &#8220;Your data is processed within the AWS data boundary, is not shared with the model provider, and is not used to train the underlying model.&#8221; Prices from the Bedrock <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k3.html">Kimi K3 model card</a>, accessed September 19th, 2026, Standard tier, Global cross-Region: $3.00 input, $15.00 output, $0.30 cache read per million tokens; the US cross-Region profile (&#8221;US CRIS&#8221; in AWS&#8217;s price table) is 1.10x on every cell. Moonshot&#8217;s own <a href="https://platform.kimi.ai/docs/pricing">pricing page</a>, accessed September 19th, 2026, lists kimi-k3 at $3.00 input (cache miss), $15.00 output, $0.30 cached input. The document is the AWS model card; the independent confirmation of the price match is Moonshot&#8217;s page.</p><p>[2] NSA, CISA and FBI, <a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a">Cybersecurity Advisory AA26-251A</a>, September 8th, 2026, section on Moonshot AI, accessed September 19th, 2026: &#8220;Notably, Moonshot AI extracted significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model.&#8221; The record copy is the PDF on media.defense.gov, which refuses automated requests; an <a href="https://web.archive.org/web/20260909111728/https://media.defense.gov/2026/Sep/08/2003992823/-1/-1/0/CSA_CHINA_BASED_AI_COMPANIES_MALICIOUS_DISTILLATION_AGAINST_US.PDF">archived copy</a> dates from September 9th, 2026. The piece reports what the advisory says, not that the allegation is proven. The advisory&#8217;s reference list cites Anthropic&#8217;s earlier post &#8220;Detecting and preventing distillation attacks&#8221;, so Anthropic&#8217;s September 10th report, which came two days after the advisory, is not an independent confirmation of the allegation.</p><p>[3] Bedrock <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k3.html">Kimi K3 model card</a>, accessed September 19th, 2026: &#8220;Model launch date: 18th Sept 2026 EOL no sooner than: Not Applicable, at least 45 day EOL Notice will be provided Legacy period: at least 45 days&#8221;. The new <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html">Model lifecycle page</a>, accessed September 19th, 2026, defines the two fields: &#8220;An EOL no sooner than date: the model will not reach EOL before this date&#8221; and &#8220;The Legacy period: the notice period before EOL.&#8221; Archived copy of the card: <a href="https://web.archive.org/web/20260919175108/https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k3.html">20260919175108</a>, which carries the same line. The launch post does not state the retirement terms. The document is the live card; the record copy is the archive; what the fields mean comes from the lifecycle page.</p><p>[4] Bedrock lifecycle page as archived on September 3rd, 2026, <a href="https://web.archive.org/web/20260903155829/https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html">Wayback copy 20260903155829</a>: &#8220;Once a model launches on Amazon Bedrock, it will remain on Amazon Bedrock for at least 12 months before the EOL date.&#8221; and &#8220;A model will be in the Legacy state for at least 6 months before the EOL date.&#8221; The same sentences now sit on the <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle-legacy.html">Model lifecycle (Legacy) page</a>, accessed September 19th, 2026, under the note &#8220;This page applies to all models launched on Amazon Bedrock before September 7, 2026.&#8221; The document is the archived copy; the independent confirmation is the live Legacy page.</p><p>[5] Author&#8217;s census of every model card linked from the Bedrock <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards.html">model cards index</a>, fetched September 19th, 2026: 128 cards linked, 127 print lifecycle fields, 126 give &#8220;Legacy period: at least 6 months&#8221;, one (Kimi K3) gives 45 days; only GPT-6 Astra and Kimi K3 carry a launch date on or after September 7th, 2026. Examples: the <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k2-5.html">Kimi K2.5 card</a>, &#8220;Model launch date: Jan 27, 2026 EOL no sooner than: Jan 27, 2027 Legacy period: at least 6 months&#8221;; the <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-deepseek-deepseek-v3-2.html">DeepSeek V3.2 card</a>, &#8220;Dec 01, 2025 &#8230; Dec 01, 2026 &#8230; at least 6 months&#8221;; the <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-qwen-qwen3-coder-next.html">Qwen3 Coder Next card</a>, &#8220;Feb 04, 2026 &#8230; Feb 04, 2027 &#8230; at least 6 months&#8221;. Of the six companies named in the advisory, five have models on Bedrock: DeepSeek (3 cards), Qwen (7), MiniMax (3), Z.AI (3) and Moonshot (3, including K3); StepFun has none. All eighteen cards other than K3 give a 12-month floor and &#8220;at least 6 months&#8221;. Spot-checked live on September 20th, 2026: the <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-deepseek-deepseek-v3-2.html">DeepSeek V3.2 card</a> (&#8221;Dec 01, 2025 &#8230; Dec 01, 2026 &#8230; at least 6 months&#8221;), the <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-zai-glm-5.html">GLM-5 card</a> (&#8221;Feb 11, 2026 &#8230; Feb 11, 2027&#8221;) and the <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-minimax-minimax-m2-5.html">MiniMax M2.5 card</a> (&#8221;Feb 12, 2026 &#8230; Feb 12, 2027&#8221;). Raw lines are kept with the piece&#8217;s working files. The index stood at 128 cards on September 19th and 127 on September 20th, after Claude Mythos Preview was delisted; no new card had appeared by September 20th.</p><p>[6] Bedrock <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html">Model lifecycle page</a>, accessed September 19th, 2026, archived at <a href="https://web.archive.org/web/20260919173232/https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html">20260919173232</a>: &#8220;This page describes the model lifecycle policy for models launched on Amazon Bedrock on or after September 7, 2026.&#8221; and &#8220;There are two Legacy periods: 6 months and 45 days. Most models have a 6-month Legacy period.&#8221; Dating: the <a href="https://web.archive.org/cdx/search/cdx?url=docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html&amp;output=json&amp;fl=timestamp,statuscode&amp;from=202608">Wayback index for the page</a> lists copies on August 2nd, September 3rd and September 19th, 2026; the Legacy page&#8217;s first archived copy is September 9th, 2026, 22:35 UTC, and the Astra card in the new format was archived September 9th, 2026, 09:36 UTC. The Bedrock <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/doc-history.html">document history</a>, accessed September 19th, 2026, has no entry for the change.</p><p>[7] Bedrock <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-openai-gpt-6-astra.html">GPT-6 Astra model card</a>, accessed September 19th, 2026: &#8220;Model launch date: September 8, 2026 EOL no sooner than: September 8, 2027 Legacy period: at least 6 months&#8221;. Archived on September 9th, 2026, at <a href="https://web.archive.org/web/20260909093622/https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-openai-gpt-6-astra.html">20260909093622</a>, which carries the same fields. The document is the live card; the independent confirmation is the archived copy. The OpenAI models came to Bedrock under the agreement Amazon announced on <a href="https://www.aboutamazon.com/news/aws/bedrock-openai-models">April 28th, 2026</a>: &#8220;We are excited to announce that starting today, the latest OpenAI models will be available on Amazon Bedrock.&#8221; No comparable announcement names Moonshot; its listed tie to AWS is a Marketplace seller page (note [12]).</p><p>[8] The statute is <a href="https://www.congress.gov/119/plaws/publ60/PLAW-119publ60.pdf">Public Law 119-60</a>, the National Defense Authorization Act for Fiscal Year 2026, enacted December 18th, 2025, Sec. 1532: &#8220;Not later than 30 days after the date of the enactment of this Act, the Secretary of Defense shall require the exclusion and removal of covered artificial intelligence from the systems and devices of the Department of Defense.&#8221; The Senate&#8217;s FY2027 bill, <a href="https://www.govinfo.gov/content/pkg/BILLS-119s4784rs/pdf/BILLS-119s4784rs.pdf">S. 4784 as reported</a>, Sec. 1651, would amend Sec. 1532 to add AI developed by Moonshot AI among others; as reported, not enacted. Its cloture vote on the motion to proceed failed 50&#8211;46 on July 14th, 2026 (<a href="https://web.archive.org/web/2026/https://www.senate.gov/legislative/LIS/roll_call_votes/vote1192/vote_119_2_00195.htm">Senate roll call 195</a>, accessed September 19th, 2026). Both are discussed, with their texts, in <a href="https://www.airealist.ai/p/selective-availability">Selective Availability</a>.</p><p>[9] Moonshot AI, <a href="https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE">Kimi K3 licence</a>, section 2, accessed September 19th, 2026: &#8220;the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose&#8221;, applying to a Model-as-a-Service operator whose aggregate revenue with affiliates exceeds $20 million over any consecutive 12 months; Moonshot&#8217;s &#8220;certified inference partners&#8221; are exempt. The same text is in Moonshot&#8217;s <a href="https://raw.githubusercontent.com/MoonshotAI/Kimi-K3/main/LICENSE">GitHub copy of the licence</a>, accessed September 19th, 2026, identical apart from whitespace. Neither AWS page states whether an agreement exists.</p><p>[10] AWS News Blog, &#8220;<a href="https://aws.amazon.com/blogs/aws/deepseek-r1-models-now-available-on-aws">DeepSeek-R1 models now available on AWS</a>&#8220;, January 30th, 2025: &#8220;We highly recommend integrating your deployments of the DeepSeek-R1 models with Amazon Bedrock Guardrails to add a layer of protection for your generative AI applications.&#8221; Its Bedrock <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-deepseek-deepseek-r1.html">card</a>, accessed September 19th, 2026: &#8220;Model launch date: Jan 20, 2025 EOL no sooner than: Jan 20, 2026 Legacy period: at least 6 months&#8221;.</p><p>[11] The negative search is the author&#8217;s, rerun September 19th, 2026. The <a href="https://www.federalregister.gov/api/v1/documents.json?conditions%5Bterm%5D=%22Moonshot%20AI%22&amp;conditions%5Bpublication_date%5D%5Bgte%5D=2026-01-01">Federal Register documents API</a>, publication dates January 1st to September 19th, 2026, returns 0 documents for &#8220;Moonshot AI&#8221; (the bare word &#8220;Moonshot&#8221; returns Defense Production Act notices that use the word, not the company). The <a href="https://www.ecfr.gov/api/search/v1/count?query=Moonshot&amp;hierarchy%5Btitle%5D=15">eCFR search API</a> over Title 15, which holds the Entity List, returns 0 for &#8220;Moonshot&#8221;. The one statute that bans a Chinese model, Public Law 119-60, names DeepSeek and covers Pentagon systems and contract work (note [8]), and S. 4784 is not law. A registry search can only see rules; the same search was first run for <a href="https://www.airealist.ai/p/selective-availability">Selective Availability</a> through September 10th, 2026.</p><p>[12] Anthropic, <a href="https://www.anthropic.com/threat-intelligence-report-september-2026">threat intelligence report, September 2026</a>, published September 10th, 2026, accessed September 19th, 2026: &#8220;We discovered that Moonshot AI, the company that produces the Kimi family of models, silently forwarded customer requests to Claude, instead of processing them using Kimi.&#8221; The commitment is Anthropic&#8217;s own, in &#8220;<a href="https://www.anthropic.com/news/anthropic-amazon-compute">Anthropic and Amazon expand collaboration for up to 5 gigawatts of new compute</a>&#8220;, April 20th, 2026, accessed September 19th, 2026: &#8220;We are committing more than $100 billion over the next ten years to AWS technologies, securing up to 5GW of new capacity to train and run Claude.&#8221; Amazon&#8217;s <a href="https://www.aboutamazon.com/news/company-news/amazon-invests-additional-5-billion-anthropic-ai">post of the same day</a> gives the same figure. AWS Marketplace lists the &#8220;<a href="https://aws.amazon.com/marketplace/pp/prodview-rfjb2elzc5jp4">Kimi API Platform</a>&#8220;, seller &#8220;MOONSHOT AI&#8221;, delivered as software as a service through private offers and covering K3 among other Kimi models, accessed September 19th, 2026; the page&#8217;s embedded data gives a creation date of June 3rd, 2026. It is Moonshot&#8217;s service, not Bedrock&#8217;s. Moonshot&#8217;s statement of September 12th, 2026, on its official Weibo account, as quoted by <a href="https://www.kocpc.com.tw/archives/668740">Kocpc</a> on September 13th, 2026: &#8220;&#36817;&#26085;&#32178;&#20659;&#38364;&#26044;&#20844;&#21496;&#21109;&#36774;&#20154;&#21450;&#21729;&#24037;&#30340;&#30456;&#38364;&#36039;&#35338;&#32020;&#23660;&#34395;&#27083;&#65292;&#20418;&#24801;&#24847;&#36896;&#35616;&#12290;&#26376;&#20043;&#26263;&#38754;&#24050;&#31532;&#19968;&#26178;&#38291;&#21521;&#20844;&#23433;&#27231;&#38364;&#22577;&#26696;&#8221; (author&#8217;s translation: information circulating online about the company&#8217;s founder and employees is pure fabrication and malicious rumour; Moonshot has reported it to the police). The statement does not address the advisory or Anthropic&#8217;s allegation; searches in English and Chinese on September 19th, 2026 found no Moonshot statement that does.</p><p>[13] Microsoft, &#8220;<a href="https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirements">Foundry Models lifecycle and support policy</a>&#8220;, updated July 24th, 2026, accessed September 19th, 2026: &#8220;Generally available models from Anthropic, DeepSeek, Fireworks, and Mistral AI follow a 12-month lifecycle instead of the standard 18-month lifecycle.&#8221; and &#8220;If a model is found to have compliance or security issues, Microsoft reserves the right to invoke an emergency retirement with shortened notice.&#8221; GA retirement notice: &#8220;At least 60 days before retirement&#8221;. The catalogue entry is <a href="https://ai.azure.com/catalog/models/FW-Kimi-K3">FW-Kimi-K3</a>, accessed September 19th, 2026: &#8220;Generally available&#8221;, publisher Fireworks, created by Moonshot AI.</p><p>[14] Microsoft, &#8220;<a href="https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule">Model retirement schedule</a>&#8220;, updated September 14th, 2026, accessed September 20th, 2026, under &#8220;Foundry Models from partners and community&#8221;, Fireworks: &#8220;Fireworks models on Standard (Per-Token) inference offerings are subject to a 15-day notice period prior to model retirement.&#8221; Every Fireworks row on that page carries the same retirement date, 2027-07-01, including FW-Kimi-K2.5, FW-DeepSeek-V3.2 and FW-GLM-5; FW-Kimi-K3 is in the catalogue but has no row on the schedule yet. This page is more recent and more specific than the policy page in note 13, and it governs: the 12-month lifecycle in that policy is not what a Fireworks-served model gets on notice.</p><p>[15] Andy Jassy, <a href="https://www.aboutamazon.com/news/company-news/amazon-ceo-andy-jassy-2024-letter-to-shareholders">2024 letter to shareholders</a>, accessed September 20th, 2026, closing line: &#8220;It remains Day One.&#8221; Jassy&#8217;s <a href="https://www.aboutamazon.com/news/company-news/amazon-ceo-andy-jassy-2025-letter-to-shareholders">2025 letter</a>, published April 2026, reprints the 1997 letter in full (&#8221;this is Day 1 for the Internet and, if we execute well, for Amazon.com&#8221;).</p><p>[16] Our earlier piece: <a href="https://www.airealist.ai/p/amazon-priced-the-frontier-and-declined">Amazon Priced the Frontier and Declined It</a>, July 29th, 2026: &#8220;the durable position is always the meter&#8221;, and &#8220;AWS doesn&#8217;t mark Claude up.&#8221;</p><p>[17] The wording is AWS&#8217;s own, posted to its Health Dashboard on March 1st, 2026 and reported on the wire that morning; our piece <a href="https://www.airealist.ai/p/objects-that-struck-the-data-center">Objects That Struck the Data Center</a>, March 2nd, 2026, quotes it and sets it against that week&#8217;s IRGC barrage on the UAE. AWS named the cause in a dashboard post the following day: two UAE facilities &#8220;directly struck by drones&#8221;, a third in Bahrain hit by &#8220;a drone strike in close proximity&#8221;, as reported by <a href="https://www.datacenterdynamics.com/en/news/amazon-confirms-two-uae-data-centers-hit-by-drone-strikes-third-in-bahrain-damaged/">DataCenterDynamics</a> on March 3rd, 2026 (read September 20th, 2026 in the Wayback copy of May 28th, 2026, the site refusing automated requests) and by <a href="https://www.cbsnews.com/news/amazon-drone-strike-aws-data-center-uae-bahrain-iran/">CBS News</a> on March 3rd, 2026, accessed September 20th, 2026: &#8220;Drones directly struck two Amazon Web Services facilities in the United Arab Emirates, and a drone strike near an Amazon data center in Bahrain also damaged that facility, the company said in a post on Monday on AWS&#8217;s health dashboard.&#8221; AWS has not attributed the strikes to Iran; our March piece calls the correlation &#8220;strong but circumstantial&#8221; and notes the site may have been hit by intercepted debris rather than directly. That piece dates the update at &#8220;seventy-two hours&#8221;; the statement it describes came the evening of March 2nd, about a day and a half after the first report.</p><p>[18] Corey Quinn, &#8220;<a href="https://www.theregister.com/2026/03/26/aws_would_prefer_to_forget/">AWS would prefer to forget March ever happened in its UAE region</a>&#8220;, The Register, March 26th, 2026, accessed September 20th, 2026, on the billing-waiver email: &#8220;No explanation. No mention of the Iranian drone strikes that physically destroyed two of three availability zones in the region on March 1st.&#8221;</p>]]></content:encoded></item><item><title><![CDATA[Selective Availability]]></title><description><![CDATA[Three agencies advised American AI providers to degrade answers for accounts they suspect of malicious distillation, subtly enough to &#8220;avoid triggering obvious alerts&#8221;.]]></description><link>https://www.airealist.ai/p/selective-availability</link><guid isPermaLink="false">https://www.airealist.ai/p/selective-availability</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Fri, 11 Sep 2026 14:58:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!cbeV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cbeV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cbeV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cbeV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cbeV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cbeV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cbeV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:530864,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/215230556?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cbeV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cbeV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cbeV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cbeV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc46cbd2-1e25-499d-931c-8eb52ed6a5e6_1456x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On September 8th, 2026, the NSA, CISA, and the FBI published a joint advisory about Chinese AI companies extracting capabilities from American models.[1] It runs to eighteen pages, names six companies, and is mostly what you&#8217;d expect. One recommendation isn&#8217;t, and I think it&#8217;s the most important thing in it. If you buy tokens from an American AI lab, it&#8217;s about you.</p><p style="text-align: justify;">The second of the three things the advisory asks American AI companies to do right away is this: </p><blockquote><p style="text-align: justify;">&#8220;Subtly alter responses for suspected malicious distillation attempts to attenuate the payoffs to companies conducting industrial-scale distillation campaigns.&#8221;[2] </p></blockquote><p style="text-align: justify;">Page fourteen says what that looks like:</p><blockquote><p>&#8220;Reducing reasoning depth, presenting correct information with different reasoning, or stylistic inconsistencies may evade detection while reducing training usefulness.&#8221;</p></blockquote><p>And right after that, it says who shouldn&#8217;t be told:</p><blockquote><p>&#8220;Avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model.&#8221;[3]</p></blockquote><p style="text-align: justify;">Why should you care if you&#8217;re not a Chinese lab? Because of how a provider is supposed to spot one. The advisory tells providers to monitor &#8220;immediate maximum usage from new accounts, and enterprise-scale throughput patterns&#8221;, and lists &#8220;usage optimized for cache maximization versus task diversity&#8221;.[4] That describes a lot of ordinary production traffic, for example, a batch job moved to a new account that runs one cached prompt over a large corpus at full volume from day one. Vendors tell you to cache: OpenAI&#8217;s guide says &#8220;Keep the prefix stable.&#8221;[5] And distillers often come in on somebody else&#8217;s account: the proxy services they route through, Anthropic says, &#8220;often use stolen API credentials belonging to legitimate companies or individuals&#8221;.[6] Degrade that traffic, and the company paying for the key gets the worst answers, with nothing in the advisory saying anyone should tell it.</p><p style="text-align: justify;">Google and Anthropic have both publicly said that they make what distillers get less useful. What matters is whether they tell the account. Anthropic&#8217;s help page for its newest models says it does, at least when a safeguard blocks a request, distillation attempts included: in its apps &#8220;you&#8217;ll see a notice explaining that the model switched, and the response will be labeled with the model that answered&#8221;, and on the API the response comes back with a stop reason.[7] Google&#8217;s Gemini API documentation, which never mentions distillation, reserves the right to change &#8220;which model answers a specific request&#8221; and says Google &#8220;may reach out to you through email&#8221;, not that it will.[8] So ask your account manager two questions. Do you alter answers for accounts you suspect? Would you tell me if I were one of them?</p><h2><strong>Who says so?</strong></h2><p style="text-align: justify;">The advisory&#8217;s central claim is that distillation &#8220;is not a supplement to these companies&#8217; AI model development, but the critical core of it.&#8221;[9] It names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, and says they have processed &#8220;billions of tokens across millions of exchanges/requests&#8221; since at least late 2024.</p><p style="text-align: justify;">So what&#8217;s the evidence? The eight references are all public: three come from companies that say they were targeted (Anthropic, Google and OpenAI), and the rest are a NIST taxonomy, a trade article, two White House memoranda and a post on X.[10] Nothing in the list is the agencies&#8217; own work (for example, telemetry, a seizure, or an indictment), and the White House memo on distillation says the government &#8220;has information&#8221; without showing any of it. StepFun, Z.AI, and GPT-oss-20b appear in none of them.</p><p style="text-align: justify;">Is the distillation happening? It may well be. Anthropic&#8217;s February post is specific in a way the advisory isn&#8217;t: about 24,000 fraudulent accounts and more than 16 million exchanges, attributed to three labs.[11] Compared with that, the advisory adds three more company names, longer lists of the models it says were used, detection signals, mitigations, and one instruction.</p><p style="text-align: justify;">And most of what it recommends only works while the model sits behind an API, which it never says.</p><h2><strong>Twelve of the seventeen measures only work at the front door</strong></h2><p style="text-align: justify;">The advisory makes seventeen recommendations: eleven from a MITRE catalog, three in its own prose, and three from a NIST taxonomy.</p><p style="text-align: justify;">Twelve of them do nothing unless the attacker&#8217;s requests hit infrastructure the provider runs, i.e., rate limiting, authenticated access, logging, output obfuscation, ensembles, query detection, and input sanitization.[12]</p><p style="text-align: justify;">Picture the same model downloaded and running on somebody else&#8217;s hardware. What&#8217;s left? There&#8217;s no quota to enforce, no log to write, no answer to degrade, and no way to know whose inference it is. Two measures survive the download, and they&#8217;re the same idea filed under two frameworks: safety training baked into the weights before release.[13] Anyone who holds the weights and is willing to tweak could strip that out, and the advisory doesn&#8217;t discuss it.</p><p style="text-align: justify;">On the weights themselves, there&#8217;s exactly one line, AML.M0001: </p><blockquote><p style="text-align: justify;">&#8220;Limit Model Artifact Release: Limit release of data, algorithms, architectures, and model checkpoints.&#8221;[14]</p></blockquote><p style="text-align: justify;">Search the eighteen pages for &#8220;weights&#8221;, &#8220;open-weight&#8221;, &#8220;open source&#8221; or &#8220;publicly available,&#8221; and you get nothing; &#8220;checkpoint&#8221; shows up once, inside that one line.[15]</p><p style="text-align: justify;">MITRE&#8217;s own catalog is less shy. One of its entries, AML.M0017, spells the assumption out: </p><blockquote><p style="text-align: justify;">&#8220;Deploying AI models to edge devices can increase the attack surface of the system. Consider serving models in the cloud to reduce the level of access the adversary has to the model.&#8221;[16] </p></blockquote><p style="text-align: justify;">The advisory copied ten identifiers from that catalog and skipped this one. Oops.</p><h2><strong>Exhibit A was a free download</strong></h2><p style="text-align: justify;">The Moonshot section lists eighteen models the advisory says the company used &#8220;to distill SFT optimization, reinforcement learning (RL), software engineering, and math capabilities&#8221;. Most are what you&#8217;d expect (six Claude variants, several GPT and Gemini releases, Grok Code Fast-1), and one of them is GPT-oss-20b.[17]</p><p style="text-align: justify;">GPT-oss-20b is OpenAI&#8217;s own open-weight model, released in August 2025 under the Apache 2.0 license. It isn&#8217;t gated; it was downloaded about 6.5 million times in the last thirty days, and its license permits the use described in the advisory.[18]</p><p style="text-align: justify;">So what in the document reaches that row? Almost nothing. That leaves AML.M0001 (limit what you release), and OpenAI decided the other way thirteen months ago.</p><h2><strong>February&#8217;s version had a clause for customers</strong></h2><p style="text-align: justify;">The idea wasn&#8217;t the agencies&#8217;. On February 12th, Google wrote that its defenses &#8220;can degrade student model performance&#8221; [19], and on February 23rd, Anthropic described what it was building: &#8220;Product, API and model-level safeguards designed to reduce the efficacy of model outputs for illicit distillation, without degrading the experience for legitimate customers&#8221;[20]</p><p>That last clause is as close as either post comes to a promise: legitimate customers aren&#8217;t meant to pay for the defense.</p><p style="text-align: justify;">Both posts are in the advisory&#8217;s reference list [10], and the agencies changed two things: the promise shrinks to &#8220;lower-to-no legitimate user risk&#8221;, and next to it sits an instruction not to tell. The degradation was publicly proposed by vendors and defended on its merits. The instruction not to tell is the agency&#8217;s own.</p><p style="text-align: justify;">The agencies do give a reason for it in the next sentence: &#8220;Informing malicious distillers would enable them to improve their defense evasions and indicate when to roll back training.&#8221;[3] The distillers know the defenses exist, if not when one hits them. Google and Anthropic said so in February, and in April the White House warned that, as defenses improve, distillers &#8220;should have little confidence in the integrity and reliability of the models they produce&#8221;.[21] And page 10 says China-based entities already run test pipelines that separate &#8220;service issues from defensive data degradation&#8221;, which is why page 14 suggests changes subtle enough to get past them. A change built to fool a lab&#8217;s test pipeline would fool a customer&#8217;s tests too. So the silence keeps two groups in the dark: the distillers and the paying customers.</p><p style="text-align: justify;">Two sentences on, the advisory concedes what this costs: &#8220;In contrast, AI safety researchers and third-party evaluators should be informed of model changes while still applying strong distillation mitigations.&#8221;[3] Why tell the evaluators? Because the defense changes what they would measure, even if the carve-out only means to mark researchers as friends.</p><p style="text-align: justify;">What about you, as a buyer? The notice covers &#8220;AI safety researchers and third-party evaluators&#8221;. Not you. And I wouldn&#8217;t count on your own tests to cover it: in a survey LangChain ran at the end of 2025, 22.8% of the respondents with agents already in production said they weren&#8217;t evaluating them.[22]</p><p style="text-align: justify;">Who does the section cover? It&#8217;s headed &#8220;Response alteration for suspected distillation activity&#8221;, it opens on &#8220;high-confidence malicious distillation requests&#8221;, and its implementation paragraph starts &#8220;When suspecting a malicious distillation campaign&#8221;. None of the three names a nationality. The instruction not to tell comes twice: first for China-based users, then, right after the agencies&#8217; reason, &#8220;Instead, alter responses to users confirmed to be querying frontier models specifically for malicious knowledge distillation campaigns without informing them.&#8221;[23] Read with the sentence before it, which may be intended only for Chinese users. Read on its own, it covers anyone the provider has &#8220;confirmed&#8221;, by a standard the advisory never gives.</p><p style="text-align: justify;">The advisory knows this can catch the wrong people. In its section on sharing indicators, it says correlated activity justifies degradation &#8220;with lower-to-no legitimate user risk&#8221;.[24] It gives no false-positive rate anywhere.[25] The providers&#8217; own safety filters suggest it wouldn&#8217;t be zero: Anthropic says of its improved biology safeguards that &#8220;there will inevitably remain false positives&#8221;, and OpenAI says its agent&#8217;s safeguards &#8220;will sometimes accidentally prevent safe uses of the product&#8221;.[26] Those aren&#8217;t distillation detectors, but Anthropic calls the blocking safeguards that catch distillation attacks &#8220;intentionally broad&#8221;.[7] Some paying customers would likely get worse answers by mistake. The agencies know it, and what they offer those customers is a better guess about who they are, not a warning.</p><h2><strong>The last time Washington degraded a service, it said so</strong></h2><p style="text-align: justify;">The United States deliberately degraded the civilian GPS signal until midnight on May 1st, 2000, and the point is how it stopped. The President announced it: &#8220;The United States will stop the intentional degradation of the Global Positioning System (GPS) signals available to the public beginning at midnight tonight. We call this degradation feature Selective Availability (SA).&#8221;[27]</p><p style="text-align: justify;">So the degradation had a name, a directive behind it, a 2006 deadline, and a President announcing it would end early. And Washington kept the capability and said so in the same statement: &#8220;We have demonstrated the capability to selectively deny GPS signals on a regional basis when our national security is threatened.&#8221;</p><p style="text-align: justify;">Two things in Washington&#8217;s favor: it ran GPS itself, whereas here it only advises private firms about their own products, and Selective Availability was fixable with differential GPS, so candor cost less than it looks. Still, the GPS error was global, bounded, published, and correctable, and an altered answer is none of those.</p><p style="text-align: justify;">I&#8217;ve mapped the off switches in the AI stack before, as three: chips, cloud, and models.[28] What the three have in common is that you can see them. A revoked license, a suspended tenant, or a blocked region each leaves a record that someone can point to.</p><p style="text-align: justify;">Degrading answers for accounts a provider suspects is a fourth lever, underneath the other three, and it leaves no record the customer can see. GPS shows the government understood that people whose service it degrades should be told. The 2026 recommendation pulls the same lever and says to tell only the evaluators.</p><h2><strong>What the agencies would say, and where they&#8217;re right</strong></h2><p>The agencies have a good answer, and it starts with what kind of document this is.</p><p style="text-align: justify;">A joint cybersecurity advisory is written for defenders about controls. This one opens by granting that distillation is &#8220;recognized as a legitimate and useful technique in AI research&#8221;, and it recommends instead of requiring.[29] Focusing on the serving path is the genre doing its job. So is the sourcing: an unclassified product cites open material by convention, so the reference list shows what convention produces, not what the agencies know. And the threat it documents really does run through APIs.</p><p style="text-align: justify;">Nor do the two White House memoranda it cites go after the weights: the one on distillation says the United States &#8220;will continue to foster a vibrant open-source ecosystem built on firm foundations&#8221;, and saves its fire for models &#8220;derived from acts of malicious exploitation&#8221;.[30] And degrading answers isn&#8217;t obviously wrong either. On the advisory&#8217;s own account, these are fraudulent accounts, opened in breach of the terms of service, and a provider that makes its product less useful to someone abusing it is enforcing a contract.</p><p style="text-align: justify;">Washington has also told private firms to keep quiet about suspects before: a bank that reports a suspicious transaction may not tell the customer. But the bank hides the report, not what happens to you, because a blocked wire or a closed account is hard to miss. And on September 2nd, six days before this advisory, five federal agencies said a bank may tell a customer that a restriction or closure &#8220;may be related to suspected fraud or other suspicious activity&#8221;.[31] This advisory wants its changes to &#8220;avoid triggering obvious alerts&#8221;.</p><p style="text-align: justify;">Anyone who has run an API bill will raise the strongest objection of all: you could never verify these outputs anyway. Models get swapped behind a stable name, quantization changes, requests get routed by load, and safety filters move. The vendors admit some of it: OpenAI didn&#8217;t proactively announce the GPT-4o update it pulled from ChatGPT in April 2025, &#8220;because we expected this to be a fairly subtle update&#8221;, and Anthropic&#8217;s docs say the weights under a model ID are fixed, but &#8220;the serving infrastructure around the model can change over time&#8221;.[32]</p><p style="text-align: justify;">Fair point, and I wouldn&#8217;t promise anyone the same answer twice from a hosted model. But a change that hits every customer is one thing: somebody notices, and OpenAI began rolling that update back three days later. A change that is deliberate, aimed at particular accounts, built to be hard to spot (for example, shallower reasoning that still lands on the right answer), and recommended by three agencies, with a note not to mention it is another, and a buyer shouldn&#8217;t treat the two alike.</p><p style="text-align: justify;">On September 9th, China&#8217;s commerce ministry said the accusations have &#8220;no basis in fact or law&#8221;, that distillation is &#8220;a neutral technical means&#8221; used by model makers worldwide, American ones included, and that if Washington moves against Chinese AI companies &#8220;in the name of fighting distillation&#8221;, China &#8220;will resolutely take countermeasures&#8221;.[33]</p><h2><strong>A rule for the Pentagon, advice for everyone else</strong></h2><p style="text-align: justify;">I&#8217;ve argued before that rules bind the people who were never the threat.[28] This advisory is what the next step looks like.</p><p style="text-align: justify;">Inside the administration, the direct route has reportedly been tried and has gone nowhere. Axios reported in July 2026, citing sources close to the administration, that during 2025 it considered an executive order, an Entity List designation, an agency advisory, and Commerce rules to restrict Chinese models, and killed every one; the same piece reports the momentum coming back.[34] In August, Axios reported that an unpublished White House framework says nothing in it should be read as restricting open models once they&#8217;re released.[35]</p><p style="text-align: justify;">Executive Order 14409 of June 2nd, 2026, may appear to close that route. It says nothing in its section &#8220;shall be construed to authorize&#8221; a mandatory licensing, preclearance, or permitting requirement for new models. That&#8217;s a rule about how to read one section, not a ban on rulemaking, and the subsection above it orders agencies to design a voluntary framework in which developers could give the government up to thirty days with a covered frontier model before releasing it to other trusted partners.[36] So machinery aimed at the model itself has been ordered, just not made mandatory.</p><p style="text-align: justify;">Congress went further, but only for government systems and Pentagon contractors. The Defense Authorization Act signed in December 2025 orders DeepSeek&#8217;s AI out of Pentagon systems, bars Pentagon contractors from using it on their contract work, and bans the DeepSeek app from intelligence-agency systems. The Senate&#8217;s bill for next year would add Zhipu, Moonshot, MiniMax and Alibaba, among others.[37] For a company deploying these models in its own business, I can&#8217;t find a rule: the Federal Register for 2026 has no document naming DeepSeek, Moonshot, MiniMax, StepFun or Zhipu,[38] and Senator Hawley&#8217;s S. 321, which would ban importing AI &#8220;developed or produced in the People&#8217;s Republic of China&#8221;, hasn&#8217;t moved since it went to committee.[39]</p><p style="text-align: justify;">The Entity List binds too. In January 2025, under the previous administration, the Commerce Department put ten companies on it, seven of them carrying the Zhipu name, on the basis that they advance Chinese military modernization through AI research; the advisory&#8217;s own table gives Z.AI the Chinese name of the first of them.[40] But that&#8217;s an export control: it requires a license for items subject to US export rules when going to Zhipu, and does nothing to address an American company downloading GLM weights and serving them in its own business.</p><p style="text-align: justify;">So, beyond those, no rule applies to the weights. What arrived instead is a recommendation, in a document that binds nobody, that providers quietly degrade answers for accounts they suspect. That&#8217;s slower than a rule, and it leaves no public record.</p><h2><strong>If I were buying tokens today</strong></h2><p>I&#8217;d do two things.</p><p style="text-align: justify;">The first is a test. You can&#8217;t prove from outside that you get the model the vendor sells, but you can look for a weaker one: take a fixed set of prompts from your real traffic, hold the parameters, and run each prompt several times a day from your production account and from one registered and paid for separately and ramped up slowly (new accounts at full use and shared payment details are both on the advisory&#8217;s watch list). Compare scores, not single answers. Single answers change for innocent reasons (OpenAI&#8217;s caching guide says &#8220;identical requests are not guaranteed to produce identical outputs&#8221;), while a weaker model should show up in the scores. Shallower reasoning that still lands on the right answer may not, but it may show up as a lower average in the reasoning tokens that OpenAI, Anthropic, and Google report.[42] A change aimed at one account can appear as a gap between the two; a model update appears in both. A 2024 academic test found that 11 of 31 commercial Llama endpoints produced outputs that didn&#8217;t match Meta&#8217;s released weights. I&#8217;d keep it small because accounts sending identical prompts are also on that list.</p><p style="text-align: justify;">The second is a clause. OpenAI&#8217;s standard agreement already promises an email and a right to walk away if an update &#8220;materially reduces the Services&#8217; functionality&#8221;. Limiting your account falls under another section, which asks only for &#8220;reasonable efforts to notify&#8221; and allows it &#8220;without prior notice to the extent reasonably necessary&#8221;.[43] Ask for the first promise to cover the second case. If that sounds like a big ask, it&#8217;s what Washington demands for itself: its national security AI memo tells the national security enterprise to make sure, &#8220;through contractual clauses or other means&#8221;, that no vendor can &#8220;disable or degrade, or materially modify without Federal Government knowledge and approval&#8221; the AI its people depend on.[30]</p><p style="text-align: justify;">And there&#8217;s one setup where the question never comes up: a model you run yourself, the simplest arrangement in which you can show nothing was altered, because you served it. That used to be an argument about cost and control; now it&#8217;s also about proving what your own system did.</p><p style="text-align: justify;">In June, I wrote that &#8220;the tighter Washington shuts its door, the more of the world&#8217;s usage walks out the back.&#8221;[44] That was about Washington gating its own frontier models; quietly degraded answers could send buyers the same way.</p><p style="text-align: justify;">Leave the API alone, and the extraction carries on on the agencies&#8217; own account. Act on the API as this advisory recommends, and the instruction not to tell turns the American product into one a buyer can&#8217;t easily check from the outside, even when applied narrowly. Meanwhile, a model already published as open weights, like GPT-oss-20b, is untouched because there&#8217;s no interface to access it.</p><p style="text-align: justify;">Most of what Washington advises against distillation works only at the American API, and its advice there is to degrade answers quietly for accounts the provider suspects. In 2000, the President announced when the degradation would stop. If this one starts on your account, it&#8217;s built not to be noticed.</p><h3><strong>Notes</strong></h3><p>[1] NSA, CISA and FBI, joint cybersecurity advisory AA26-251A, &#8220;<a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a">China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies</a>&#8220;, released September 8th, 2026. The record copy is the 18-page PDF, U/OO/6059854-26, TLP:CLEAR, marked Ver 1.0; media.defense.gov returns HTTP 403 to automated fetch, and the copy read for this piece is the <a href="https://web.archive.org/web/2026id_/https://media.defense.gov/2026/Sep/08/2003992823/-1/-1/0/CSA_CHINA_BASED_AI_COMPANIES_MALICIOUS_DISTILLATION_AGAINST_US.PDF">Wayback Machine capture</a> of September 9th, 2026. PDF and HTML were compared in full and agree on every passage quoted here.</p><p>[2] The <a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a">advisory</a>, Executive summary, page 2: &#8220;The authoring agencies recommend U.S. AI companies take three immediate actions&#8221;, of which the second is &#8220;Deploy targeted response changes: Subtly alter responses for suspected malicious distillation attempts to attenuate the payoffs to companies conducting industrial-scale distillation campaigns.&#8221; The other two are detection and mitigation, and cross-organization intelligence sharing.</p><p>[3] The <a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a">advisory</a>, &#8220;Implementation strategies&#8221;, a subsection of &#8220;Response alteration for suspected distillation activity&#8221;, pages 13 and 14. Two sentences separate the passages quoted here: &#8220;Informing malicious distillers would enable them to improve their defense evasions and indicate when to roll back training.&#8221; and &#8220;Instead, alter responses to users confirmed to be querying frontier models specifically for malicious knowledge distillation campaigns without informing them.&#8221;</p><p>[4] The <a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a">advisory</a>, executive summary, page 2, the first of the three immediate actions: &#8220;monitor subscription-to-usage ratios, immediate maximum usage from new accounts, and enterprise-scale throughput patterns&#8221;. And &#8220;Novel TTP 4: Systematic quota and cost optimization&#8221;, which starts on page 12 and whose detection indicators on page 13 include &#8220;usage optimized for cache maximization versus task diversity&#8221;. Quoted verbatim from the record-copy PDF, September 10th, 2026; pages rechecked September 11th, 2026.</p><p>[5] OpenAI, &#8220;<a href="https://developers.openai.com/api/docs/guides/prompt-caching">Prompt caching</a>&#8220; guide, retrieved September 11th, 2026: &#8220;Keep the prefix stable.&#8221; Anthropic and Google publish equivalent guidance.</p><p>[6] Anthropic, &#8220;<a href="https://www.anthropic.com/threat-intelligence-report-september-2026">Detecting and countering misuse of AI: September 2026</a>&#8220;, dated September 10th, 2026 in the <a href="https://www-cdn.anthropic.com/e50be2e51e7695dc4b1366a37a245a597377d3b5/Anthropic-Detecting-and-countering-091026.pdf">PDF</a>, section &#8220;Illicit distillation and scaled abuse&#8221;, retrieved September 11th, 2026: &#8220;These labs generally access Anthropic&#8217;s models by routing requests through proxy services, also known as &#8220;transfer stations.&#8221;&#8220; Of those proxy services: &#8220;They will often use stolen API credentials belonging to legitimate companies or individuals to give unauthorized entities access to US frontier models.&#8221; On its response: &#8220;When we are confident that a set of requests are associated with an illicit distillation campaign or other unauthorized use of Claude, we block the request and ban the associated accounts&#8221;, and &#8220;when we detect signals of potential abuse, like the unauthorized resale of Claude or accounts operating from unsupported countries like China, Russia, and Iran, our systems can require users to verify their identity to retain access.&#8221; The report does not mention the advisory, and gives no false-positive rate. The same section: &#8220;Zhipu eventually gave up trying to target Fable after Anthropic&#8217;s cyber safeguards degraded Zhipu&#8217;s attacks.&#8221;</p><p>[7] Anthropic Help Center, &#8220;<a href="https://support.claude.com/en/articles/15363606-why-claude-switched-models-in-your-conversation-with-fable-5-or-fable-5-1">Why Claude switched models in your conversation with Fable 5 or Fable 5.1</a>&#8220;, undated, retrieved September 11th, 2026. The blocking triggers listed include &#8220;Distillation attacks on Fable 5 and Fable 5.1, including attempts to extract the model&#8217;s summarized thinking&#8221;, and &#8220;These blocking safeguards are intentionally broad&#8221;. On the API, &#8220;Until fallbacks are configured, the model will return a 200 response with a stop reason&#8221;.</p><p>[8] Google, Gemini API &#8220;<a href="https://ai.google.dev/gemini-api/docs/abuse-monitoring">Abuse monitoring</a>&#8220;, last updated June 9th, 2026, retrieved September 11th, 2026: &#8220;Temporary usage limits: We may limit your access to the Gemini API by adjusting rate limits or changing which model answers a specific request, for example.&#8221; The page promises an appeal link for suspension or account closure: &#8220;If we reach out to you regarding a suspension or account closure, we will also provide a link where you can appeal.&#8221; Google Threat Intelligence Group, &#8220;<a href="https://cloud.google.com/blog/topics/threat-intelligence/from-prompting-to-autonomy-the-evolution-of-adversarial-ai/">GTIG AI Threat Tracker: From Prompting to Autonomy &#8211; The Evolution of Adversarial AI</a>&#8220;, September 8th, 2026, retrieved September 11th, 2026: &#8220;we have deployed real-time defenses designed to degrade the performance of unauthorized &#8220;student&#8221; models and detect attempts to clone proprietary logic&#8221;, &#8220;we have developed and successfully deployed numerous methods to both lower the utility of these campaigns, and block the accounts responsible&#8221;, and &#8220;we have developed techniques to identify Gemini-distilled models, enabling us to trace the provenance of models derived from our technology&#8221;. The post says nothing about notifying accounts and does not mention the advisory. The abuse-monitoring page lists, before usage limits, &#8220;Get in touch: We may reach out to you through email to understand your use case and explore ways to bring your usage into compliance.&#8221; It does not mention distillation.</p><p>[9] Advisory, Attribution section, page 3. The second quotation, &#8220;billions of tokens across millions of exchanges/requests&#8221;, is in the executive summary on page 1, in a sentence that opens &#8220;Likely with Chinese government awareness&#8221;.</p><p>[10] The <a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a">advisory</a>, References, retrieved September 10th, 2026. The eight listed sources are Anthropic&#8217;s distillation post, Google&#8217;s GTIG AI Threat Tracker, NIST AI 100-2e2025, OpenAI&#8217;s &#8220;RE: Updated Stakes for American-Led, Democratic AI&#8221;, an article in The Decoder on the gray market in Claude tokens, National Security Presidential Memorandum 11, National Science and Technology Memorandum 4 (&#8221;Adversarial Distillation of American AI Models&#8221;), and a post on X by the Office of Science and Technology Policy. The OpenAI item is a memo, not a letter to an agency: its header reads &#8220;To: US House Select Committee on Strategic Competition between the United States and the Chinese Communist Party / From: OpenAI / Date: February 12, 2026 / Re: Updated Stakes for American-Led, Democratic AI&#8221; (<a href="https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf">OpenAI CDN PDF</a>, retrieved September 10th, 2026; <a href="https://web.archive.org/web/20260608055441id_/https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf">Wayback capture</a>), and the receiving committee&#8217;s own hearing record of <a href="https://docs.house.gov/meetings/ZS/ZS00/20260416/119165/HHRG-119-ZS00-Wstate-MahmoodY-20260416.pdf">April 16th, 2026</a> confirms it: &#8220;In February 2026, OpenAI delivered a memo to this Committee&#8221;. The advisory&#8217;s hyperlink for the item points not at OpenAI but at a <a href="https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0">Bloomberg-hosted copy</a>, whose first page is identical. Read in full for this piece: the Anthropic and Google posts, the OpenAI memo, <a href="https://the-decoder.com/how-chinas-gray-market-sells-claude-tokens-at-a-fraction-of-the-price/">The Decoder&#8217;s article</a>, NSPM-11, NSTM-4 and the post on X, which is from the account <a href="https://x.com/mkratsios47/status/2079933645888880708">@mkratsios47</a>, July 22nd, 2026, read through a mirror of X&#8217;s API: &#8220;We have information that Moonshot AI distilled Anthropic&#8217;s Fable for the development of its K3 model.&#8221; NIST AI 100-2e2025 was read by the desk&#8217;s fact-check the same day. Search of all eight, September 11th, 2026: StepFun, Z.AI, Zhipu and gpt-oss appear in none of them. The account @mkratsios47 is that of Michael Kratsios, Director of the Office of Science and Technology Policy. <a href="https://www.whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf">NSTM-4</a>, April 23rd, 2026, is a scanned PDF, read by optical character recognition: &#8220;the United States government has information indicating that foreign entities, principally based in China, are engaged in deliberate, industrial-scale campaigns to distill U.S. frontier AI systems.&#8221; OpenAI&#8217;s memo on its own practice: &#8220;We proactively remove users who appear to be attempting to distill our models to develop competitive models to OpenAI.&#8221;</p><p>[11] Anthropic, &#8220;<a href="https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks">Detecting and preventing distillation attacks</a>&#8220;, February 23rd, 2026: &#8220;These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts&#8221;. The three labs named there are DeepSeek, Moonshot and MiniMax; the advisory of September names six.</p><p>[12] Author&#8217;s count from the advisory, verified independently by a second pass over the same document. The seventeen are eleven MITRE ATLAS bullet entries (ten unique identifiers, AML.M0015 printed twice with different descriptions), three prose subsections under &#8220;Mitigations&#8221;, and three NIST AI 100-2e2025 items. The three prose subsections are the third-level headings &#8220;Behavioral detection and monitoring&#8221;, &#8220;Response alteration for suspected distillation activity&#8221; and &#8220;Cross-organization information sharing and ecosystem coordination&#8221;; &#8220;Implementation strategies&#8221; is a fourth-level heading inside the second of them, not a section of its own, and the three map one-to-one onto the three immediate actions on page 2. Twelve require the request to reach provider-operated infrastructure: AML.M0015 twice, M0004, M0019, M0024, M0002, M0006, all three prose subsections, and the NIST differential-privacy and prompt-formatting items. Counting only unique identifiers gives 11 of 16; every defensible recount leaves a majority.</p><p>[13] Advisory, AML.M0003 (&#8221;Use adversarial training and defensive distillation to increase jailbreak difficulty&#8221;) and the NIST pre- and post-training interventions. Both describe safety properties trained into the weights, which travel with the artifact and can be removed by fine-tuning; the advisory does not address removal. OpenAI&#8217;s <a href="https://arxiv.org/abs/2508.10925">gpt-oss model card</a>: &#8220;Once they are released, determined attackers could fine-tune them to bypass safety refusals or directly optimize for harm without the possibility for OpenAI to implement additional mitigations or to revoke access.&#8221;</p><p>[14] Advisory, MITRE ATLAS mitigations, page 15.</p><p>[15] Author&#8217;s search of the record-copy PDF text and, independently, of the CISA HTML rendering, case-insensitively, September 10th, 2026. Counts identical in both: &#8220;weight&#8221; as a substring 0, &#8220;open source&#8221; 0, &#8220;open-source&#8221; 0, &#8220;publicly available&#8221; 0, &#8220;checkpoint&#8221; as a substring 1. The word &#8220;publicly&#8221; does appear twice, both as &#8220;publicly quoted&#8221;, about DeepSeek&#8217;s training cost.</p><p>[16] MITRE ATLAS mitigation AML.M0017, &#8220;AI Model Distribution Methods&#8221;, read from MITRE&#8217;s <a href="https://raw.githubusercontent.com/mitre-atlas/atlas-data/main/dist/v6/ATLAS-2026.08.yaml">ATLAS data release v2026.08</a> on September 11th, 2026. That release carries thirty-nine mitigation identifiers, and its <a href="https://github.com/mitre-atlas/atlas-data/releases/tag/v2026.08">release notes</a> say &#8220;39 mitigations&#8221;. The advisory cites ten of them, AML.M0035 included, and AML.M0017 is not among the ten. An earlier read of the deprecated <code>dist/ATLAS.yaml</code>, which &#8220;will no longer be updated&#8221;, counted thirty-five and lacked AML.M0035; the text of AML.M0017 is the same in both.</p><p>[17] The <a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a">advisory</a>, Moonshot AI section, page 4, continuing to page 5, and again in Table 1 on page 6. Table 1&#8217;s version of the list omits GPT-4o, so the prose carries eighteen models and the table seventeen. Two of the eighteen entries, &#8220;Gemini 2.5 Flash-Image&#8221; and &#8220;Nano Banana&#8221;, are the same model (<a href="https://developers.googleblog.com/en/introducing-gemini-2-5-flash-image/">Google, August 26th, 2025</a>), so the prose list covers at most seventeen distinct models.</p><p>[18] Hugging Face API record for <a href="https://huggingface.co/openai/gpt-oss-20b">openai/gpt-oss-20b</a>, retrieved September 10th, 2026: licence apache-2.0, created August 4th, 2025, gated false, 6,551,191 downloads in the trailing thirty days. Confirmed independently against OpenAI&#8217;s <a href="https://arxiv.org/abs/2508.10925">gpt-oss model card</a>: &#8220;We release the model weights, inference implementations, tool environments, and tokenizers under an Apache 2.0 license.&#8221;</p><p>[19] Google Threat Intelligence Group, &#8220;<a href="https://cloud.google.com/blog/topics/threat-intelligence/distillation-experimentation-integration-ai-adversarial-use">GTIG AI Threat Tracker: Distillation, Experimentation, and (Continued) Integration of AI for Adversarial Use</a>&#8220;, February 12th, 2026, the second entry in the advisory&#8217;s reference list, retrieved September 11th, 2026: &#8220;Google continuously detects, disrupts, and mitigates model extraction activity to protect proprietary logic and specialized training data, including with real-time proactive defenses that can degrade student model performance.&#8221; Reducing the fidelity of model outputs was in MITRE&#8217;s ATLAS catalogue well before either post (AML.M0002, created April 2023, cited by the advisory).</p><p>[20] Anthropic, &#8220;<a href="https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks">Detecting and preventing distillation attacks</a>&#8220;, February 23rd, 2026, under &#8220;Countermeasures&#8221;, retrieved September 10th, 2026.</p><p>[21] The White House&#8217;s <a href="https://www.whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf">NSTM-4</a>, &#8220;Adversarial Distillation of American AI Models&#8221;, April 23rd, 2026, read by optical character recognition of the scanned PDF on September 11th, 2026: &#8220;As methods to detect and mitigate industrial-scale distillation grow more sophisticated, foreign entities who build their AI capabilities on such fragile foundations should have little confidence in the integrity and reliability of the models they produce.&#8221; The advisory, page 10: &#8220;China-based entities deploy production-grade automated quality assurance pipelines with multi-modal validation, enabling rapid detection of degraded outputs and differentiation of service issues from defensive data degradation.&#8221;</p><p>[22] LangChain, &#8220;<a href="https://www.langchain.com/state-of-agent-engineering">State of Agent Engineering</a>&#8220;, retrieved September 11th, 2026: a public survey run from November 18th to December 2nd, 2025, &#8220;We received 1340 responses&#8221;, 63% of them in technology. Overall, &#8220;not evaluating&#8221; is 29.5%, and 22.8% among teams with agents in production; 52.4% run offline evaluations on test sets. The respondents chose themselves, and LangChain sells evaluation tools.</p><p>[23] Advisory, &#8220;Response alteration for suspected distillation activity&#8221; and its &#8220;Implementation strategies&#8221; subsection, pages 13 and 14. The section&#8217;s triggers are &#8220;high-confidence malicious distillation requests&#8221; and &#8220;When suspecting a malicious distillation campaign&#8221;; neither carries a nationality qualifier. One sentence separates the China-scoped sentence from the unscoped one quoted here: the advisory&#8217;s stated reason for the secrecy.</p><p>[24] Advisory, &#8220;Cross-organization information sharing and ecosystem coordination&#8221;, page 14; two more sentences follow it.</p><p>[25] Author&#8217;s search of the record-copy PDF and the CISA HTML rendering, September 10th, 2026: &#8220;false positive&#8221; appears 0 times, &#8220;error rate&#8221; 0 times, and the document contains no percentage or rate figure of any kind. Its treatments of wrongly flagged customers are qualitative: under AML.M0002 on page 15, &#8220;Reduce fidelity of responses (withhold logits/confidences, shorten responses, targeted redaction). Balance security with user experience.&#8221;, and &#8220;lower-to-no legitimate user risk&#8221; in &#8220;Cross-organization information sharing and ecosystem coordination&#8221;.</p><p>[26] Anthropic, &#8220;<a href="https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards">Improving Fable 5&#8217;s biology safeguards</a>&#8220;, August 7th, 2026: of the improved classifier: &#8220;There will inevitably remain false positives&#8221;. Of the launch version it replaces: &#8220;We knew this would be frustrating for legitimate biology users: it would result in a high number of false positives in the near term, where users asking biology-related questions would have their requests blocked and sent to a less capable model.&#8221; OpenAI, &#8220;<a href="https://cdn.openai.com/pdf/839e66fc-602c-48bf-81d3-b21eacc3459d/chatgpt_agent_system_card.pdf">ChatGPT Agent System Card</a>&#8220;, July 17th, 2025: &#8220;This means that our safety mitigations will sometimes accidentally prevent safe uses of the product.&#8221; Both retrieved September 11th, 2026. Both describe safety classifiers for biological risk, not distillation detectors, and neither gives a false-positive rate on live traffic.</p><p>[27] The White House, Office of the Press Secretary, &#8220;<a href="https://clintonwhitehouse3.archives.gov/WH/EOP/OSTP/html/0053_2.html">Statement by the President Regarding the United States&#8217; Decision to Stop Degrading Global Positioning System Accuracy</a>&#8220;, May 1st, 2000, retrieved September 10th, 2026 from the National Archives&#8217; frozen copy of the Clinton White House site. The accuracy figure is the statement&#8217;s own: civilian users &#8220;will be able to pinpoint locations up to ten times more accurately than they do now&#8221;, a 10x improvement, quoted rather than derived. The statement records the March 1996 Presidential Decision Directive behind the policy, the commitment to discontinue Selective Availability by 2006, and that the decision followed a recommendation by the Secretary of Defense in coordination with State, Transportation, Commerce and the Director of Central Intelligence. On the error&#8217;s shape: the White House fact sheet of the same day, &#8220;<a href="https://clintonwhitehouse3.archives.gov/WH/EOP/OSTP/html/0053_4.html">Improving the Civilian Global Positioning System (GPS)</a>&#8220;, says the United States used Selective Availability &#8220;to globally degrade the civilian GPS signal&#8221;, that &#8220;Previously, a GPS-based car navigation could give the location of the vehicle to within a hundred meters&#8221;, and that the policy was to make &#8220;both the signal and the receiver design specification available to the public completely free of charge&#8221;. GPS.gov&#8217;s <a href="https://web.archive.org/web/20241229102950id_/https://www.gps.gov/systems/gps/modernization/sa/faq/">Selective Availability FAQ</a> (updated October 2001, read through the Wayback Machine) says &#8220;Selective Availability was a global degradation of the GPS service. It could not be applied on a regional basis.&#8221;, and, asked whether differential GPS was more accurate once SA ended, &#8220;No. There should not be much change in the accuracy of DGPS.&#8221; Both retrieved September 11th, 2026.</p><p>[28] &#8220;<a href="https://www.airealist.ai/p/access-disable-destroy">Access, Disable, Destroy</a>&#8220;, March 7th, 2026, which mapped the three-switch model over chips, cloud and models, and coined governance for the governed: &#8220;all of it is governance for the governed. Rules for the rule-followers.&#8221; And &#8220;<a href="https://www.airealist.ai/p/objects-that-struck-the-data-center">Objects That Struck the Data Center</a>&#8220;, March 2nd, 2026: &#8220;Open-source model weights are not scarce &#8212; they are infinitely copyable at zero marginal cost.&#8221;</p><p>[29] Advisory, Purpose section and executive summary.</p><p>[30] Advisory, References. <a href="https://www.whitehouse.gov/presidential-actions/2026/06/national-security-presidential-memorandum-nspm-11/">NSPM-11</a>, &#8220;Artificial Intelligence in the National Security Enterprise&#8221;, June 5th, 2026, retrieved September 11th, 2026, tells the national security enterprise to &#8220;adapt commercial or open-source AI technologies&#8221; and to &#8220;ensure, through contractual clauses or other means, that no commercial entity or adversary possesses the capability to prevent use of, disable or degrade, or materially modify without Federal Government knowledge and approval, an AI system that our men and women depend on for their missions&#8221;. <a href="https://www.whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf">NSTM-4</a>, &#8220;Adversarial Distillation of American AI Models&#8221;, April 23rd, 2026, read by optical character recognition of the scanned PDF on September 11th, 2026: &#8220;Consistent with America&#8217;s AI Action Plan, the United States will continue to foster a vibrant open-source ecosystem built on firm foundations&#8221;, and &#8220;there is nothing open about supposedly open models that are derived from acts of malicious exploitation&#8221;.</p><p>[31] 31 U.S.C. 5318(g)(2)(A)(i), <a href="https://www.law.cornell.edu/uscode/text/31/5318">Cornell LII</a>, retrieved September 11th, 2026: a financial institution that reports a suspicious transaction may not &#8220;notify any person involved in the transaction that the transaction has been reported&#8221;. Board of Governors of the Federal Reserve System, FDIC, FinCEN, NCUA and OCC, &#8220;<a href="https://www.federalreserve.gov/supervisionreg/srletters/SR2605a1.pdf">Joint Statement on Suspicious Activity Report Confidentiality Considerations Regarding Communications with Customers</a>&#8220;, issued with SR 26-5 on September 2nd, 2026, which lists as permitted &#8220;Notifying a customer that a delay, limitation, or restriction on an account or service or closure of an account may be related to suspected fraud or other suspicious activity&#8221;. On what the confidentiality has cost customers: GAO-18-263, February 26th, 2018, read through the <a href="https://web.archive.org/web/2026id_/https://www.gao.gov/assets/gao-18-263.pdf">Wayback Machine</a>because gao.gov refuses automated fetch, estimated from its survey that &#8220;93 percent of Southwest border banks terminated accounts because of the filing of SARs&#8221;, and reported that people it met in three border communities said banks &#8220;terminated the accounts of longtime established customers, sometimes without notice or explanation&#8221;. The joint statement also says: &#8220;This statement does not alter existing Bank Secrecy Act (BSA) legal or regulatory requirements or establish new supervisory expectations.&#8221;</p><p>[32] OpenAI, &#8220;<a href="https://openai.com/index/expanding-on-sycophancy/">Expanding on what we missed with sycophancy</a>&#8220;, May 2nd, 2025, read in a browser on September 11th, 2026 because openai.com refuses automated fetch: &#8220;On April 25th, we rolled out an update to GPT-4o in ChatGPT that made the model noticeably more sycophantic&#8221;, &#8220;We began rolling that update back on April 28th&#8221;, and &#8220;Because we expected this to be a fairly subtle update, we didn&#8217;t proactively announce it.&#8221; Anthropic, &#8220;<a href="https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions">Model IDs and versioning</a>&#8220;, retrieved September 11th, 2026: &#8220;Model weights are fixed for a given ID, but the serving infrastructure around the model can change over time. This infrastructure includes components such as the request router, safety classifiers, and sampling logic.&#8221;</p><p>[33] Ministry of Foreign Affairs of the People&#8217;s Republic of China, &#8220;<a href="https://www.mfa.gov.cn/eng/xw/fyrbt/lxjzh/202609/t20260909_12019292.html">Foreign Ministry Spokesperson Mao Ning&#8217;s Regular Press Conference on September 9, 2026</a>&#8220;, retrieved September 11th, 2026. AFP&#8217;s question: &#8220;Yesterday, the U.S. Cyber Defense Agency accused top Chinese AI labs, including DeepSeek and Moonshot of stealing the capabilities of U.S. models.&#8221; China&#8217;s Ministry of Commerce, &#8220;<a href="https://www.mofcom.gov.cn/xwfb/xwfyrth/art/2026/art_1439afbf24d941bfaefba078aacf340c.html">&#21830;&#21153;&#37096;&#26032;&#38395;&#21457;&#35328;&#20154;&#23601;&#32654;&#21457;&#24067;&#20013;&#22269;&#20154;&#24037;&#26234;&#33021;&#20225;&#19994;&#23545;&#32654;&#33976;&#39311;&#27963;&#21160;&#30456;&#20851;&#32593;&#32476;&#23433;&#20840;&#20844;&#21578;&#31572;&#35760;&#32773;&#38382;</a>&#8220; (spokesperson&#8217;s answer on the US advisory on Chinese AI companies&#8217; distillation), September 9th, 2026, 21:28, retrieved September 11th, 2026. The originals, in the author&#8217;s translation: &#8220;&#32654;&#26041;&#25152;&#35859;&#20013;&#22269;&#20154;&#24037;&#26234;&#33021;&#20225;&#19994;&#20174;&#20107;&#8221;&#24037;&#19994;&#35268;&#27169;&#8221;&#33976;&#39311;&#32654;&#27169;&#22411;&#30340;&#25351;&#25511;&#65292;&#20110;&#20107;&#26080;&#20973;&#65292;&#20110;&#27861;&#26080;&#25454;&#8221; (the US accusation that Chinese AI companies engage in &#8220;industrial-scale&#8221; distillation of US models has no basis in fact or law); &#8220;&#33976;&#39311;&#26159;&#20154;&#24037;&#26234;&#33021;&#39046;&#22495;&#21508;&#20010;&#27169;&#22411;&#38388;&#20114;&#30456;&#23398;&#20064;&#30340;&#36890;&#34892;&#20570;&#27861;&#65292;&#26412;&#36136;&#26159;&#20013;&#24615;&#25216;&#26415;&#25163;&#27573;&#65292;&#21253;&#25324;&#32654;&#20225;&#22312;&#20869;&#30340;&#20840;&#29699;&#27169;&#22411;&#20225;&#19994;&#37117;&#22312;&#29992;&#8221; (distillation is a common practice by which AI models learn from one another, in essence a neutral technical means, used by model companies worldwide, American ones included); &#8220;&#20013;&#26041;&#24320;&#28304;&#27169;&#22411;&#21521;&#21253;&#25324;&#32654;&#20225;&#22312;&#20869;&#30340;&#20840;&#29699;&#20225;&#19994;&#24320;&#25918;&#8221; (China&#8217;s open-source models are open to companies worldwide, American ones included); &#8220;&#22914;&#26524;&#32654;&#26041;&#20197;&#25171;&#20987;&#33976;&#39311;&#20026;&#21517;&#65292;&#23454;&#26045;&#36943;&#21387;&#20013;&#22269;&#20154;&#24037;&#26234;&#33021;&#20225;&#19994;&#30340;&#34892;&#21160;&#65292;&#20013;&#26041;&#24517;&#23558;&#22362;&#20915;&#37319;&#21462;&#25514;&#26045;&#20104;&#20197;&#21453;&#21046;&#8221; (if the United States acts to suppress Chinese AI companies in the name of fighting distillation, China will resolutely take countermeasures). The ministry also says that American firms&#8217; own model reports disclose extensive distillation of Chinese models (&#8221;&#32654;&#20225;&#26377;&#20851;&#27169;&#22411;&#30740;&#21457;&#25253;&#21578;&#20063;&#25259;&#38706;&#65292;&#20854;&#22823;&#37327;&#33976;&#39311;&#20013;&#22269;&#27169;&#22411;&#8221;); that is the ministry&#8217;s claim, not checked here.</p><p>[34] Axios, &#8220;The secret Trump administration battle to fight Chinese AI&#8221;, July 20th, 2026, read via the <a href="https://web.archive.org/web/20260720124350/https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi">Wayback capture</a> of the same day: &#8220;Administration officials keen on keeping regulation from stifling innovation killed all of those efforts&#8221;, and &#8220;ban momentum is picking up again.&#8221; Axios attributes the account to &#8220;knowledgeable sources&#8221; and &#8220;a source close to the administration&#8221;. Unnamed sources; reported, not confirmed.</p><p>[35] Axios, August 5th, 2026, <a href="https://web.archive.org/web/20260806102000/https://www.axios.com/2026/08/05/trump-ai-framework-china">via the Wayback Machine</a>: &#8220;The White House does not plan to publicly release the framework, which explicitly says nothing in it should be interpreted as restricting open models once they&#8217;ve been released&#8221;. The same piece reports the question &#8220;far from resolved&#8221;. Bloomberg reported the same development on the same date; only its opening paragraphs are readable without a subscription, so Axios is the source read in full here. The August 5th piece gives no source for the framework&#8217;s text beyond &#8220;Axios first reported&#8221;. Unnamed sources; reported, not confirmed.</p><p>[36] Executive Order 14409, &#8220;Promoting Advanced Artificial Intelligence Innovation and Security&#8221;, signed June 2nd, 2026, published at <a href="https://www.govinfo.gov/content/pkg/FR-2026-06-05/html/2026-11415.htm">91 FR 34565</a>, Sec. 3(c). The disclaimer is scoped to that section; Sec. 3(b) orders agencies to &#8220;design a voluntary framework with AI developers&#8221; through which developers could give the government access to covered frontier models &#8220;for a period of up to 30 days before they plan to release such models to other trusted partners&#8221;. CBS News, &#8220;<a href="https://www.cbsnews.com/news/ai-model-framework-white-house/">White House framework for testing AI models remains hidden as concerns about threats mount</a>&#8220;, September 10th, 2026, retrieved September 11th, 2026: &#8220;The White House finalized the framework by the beginning of August, following an executive order President Trump signed in June&#8221;; &#8220;Neither the administration nor the companies are required to release the results of the reviews or even say if they participated&#8221;; &#8220;Last week, OpenAI CEO Sam Altman said the company had submitted its powerful new Astra model for review&#8221;.</p><p>[37] Public Law 119-60, the National Defense Authorization Act for Fiscal Year 2026, enacted December 18th, 2025, <a href="https://www.govinfo.gov/content/pkg/PLAW-119publ60/html/PLAW-119publ60.htm">text at GovInfo</a>, retrieved September 11th, 2026. Sec. 1532: the Secretary of Defense &#8220;shall require the exclusion and removal of covered artificial intelligence from the systems and devices of the Department of Defense&#8221;, and &#8220;no contractor may, during the period of performance of such contractor under a contract with the Department of Defense, use covered artificial intelligence with respect to the performance of a contract with the Department&#8221;; &#8220;covered artificial intelligence&#8221; means &#8220;any artificial intelligence, or successor artificial intelligence, developed by the Chinese company DeepSeek&#8221; or by High Flyer. Sec. 1532(a)(2) lets the Secretary extend this by guidance to any &#8220;covered artificial intelligence company&#8221;, which includes companies on the Consolidated Screening List or &#8220;domiciled in a covered nation&#8221;. Sec. 6604, &#8220;Prohibition on use of DeepSeek on intelligence community systems&#8221;, covers &#8220;the DeepSeek application or any successor application or service&#8221;. The Senate&#8217;s bill for fiscal year 2027, <a href="https://www.govinfo.gov/content/pkg/BILLS-119s4784rs/html/BILLS-119s4784rs.htm">S. 4784 as reported</a> (Report No. 119-127), Sec. 1651, amends Sec. 1532 to add AI &#8220;developed by the Chinese company&#8221; Baidu, Zhipu AI, Moonshot AI, 01.AI, &#8220;Mistral-rival Minimax&#8221;, Alibaba and Tencent, and to make the guidance mandatory; its <a href="https://www.govinfo.gov/bulkdata/BILLSTATUS/119/s/BILLSTATUS-119s4784.xml">bill-status record</a> shows a motion to proceed on July 27th, 2026. The House passed its own bill, H.R. 8800, on July 22nd, 2026, by 216 to 212; its text does not name these companies.</p><p>[38] Two independent registries, searched September 10th, 2026. The document leg is the author&#8217;s queries against the <a href="https://www.federalregister.gov/api/v1/documents.json">Federal Register documents API</a>, publication dates January 1st to September 10th, 2026: DeepSeek 0, Moonshot AI 0, MiniMax 0, StepFun 0, Zhipu 0, &#8220;Z.ai&#8221; 0, Qwen 0, &#8220;Kimi K3&#8221; 0 (rerun September 11th, 2026, same counts); &#8220;open-weight&#8221; 2, both promotional notices. The independent confirmation is the <a href="https://www.ecfr.gov/api/search/v1/count">eCFR search API</a> over codified Title 15, which returns 0 for DeepSeek, Moonshot, MiniMax, StepFun and Qwen, and 9 for Zhipu; published documents and codified text are separate corpora. The nine are dated versions of one codified section, the Entity List at 15 CFR part 744, Supplement No. 4 (the API describes its result as &#8220;Changes to sections matching &#8216;Zhipu&#8217; in Title 15&#8221;, rechecked September 11th, 2026), which is the export control the piece goes on to describe, not a restriction on deployment. The one plausible counterexample, GSA&#8217;s proposed rule on large language model acquisition of June 17th, 2026, was read in full and contains no occurrence of &#8220;China&#8221; or &#8220;foreign adversary&#8221;. Two limits on what this shows, found at the checking stage: &#8220;Alibaba&#8221; returns four 2026 documents, among them the Defense Department&#8217;s &#8220;<a href="https://www.federalregister.gov/documents/2026/06/10/2026-11571/notice-of-availability-of-designation-of-chinese-military-companies">Notice of Availability of Designation of Chinese Military Companies</a>&#8220;, 91 FR 35189, June 10th, 2026, which lists &#8220;Alibaba Group Holding Limited&#8221;; and neither registry shows statutes, so the Federal Register search cannot see the defence law in note [37]. The claim the body rests on this note is therefore limited to deployment by a company in its own business.</p><p>[39] H.R. 1121, &#8220;<a href="https://www.govinfo.gov/content/pkg/BILLS-119hr1121ih/html/BILLS-119hr1121ih.htm">No DeepSeek on Government Devices Act</a>&#8220;, 119th Congress, as introduced: &#8220;To prohibit the use of DeepSeek by the executive agencies, and for other purposes&#8221;, sponsored by Mr. Gottheimer. Its section 2 directs the Office of Management and Budget to require removal of the DeepSeek application from federal agency information technology. Action history from the <a href="https://www.govinfo.gov/bulkdata/BILLSTATUS/119/hr/BILLSTATUS-119hr1121.xml">GovInfo bill-status record</a>, retrieved September 10th, 2026: two actions, both on February 7th, 2025, introduced and referred to House Oversight and Government Reform, and nothing since. The same record shows the Senate companion, S. 765, read twice and referred to Homeland Security and Governmental Affairs on February 27th, 2025, with no action after that. The congress.gov page carries the same history; it is not linked because it returns HTTP 403 to automated retrieval. Senator Hawley&#8217;s S. 321, the &#8220;Decoupling America&#8217;s Artificial Intelligence Capabilities from China Act of 2025&#8221;, <a href="https://www.govinfo.gov/content/pkg/BILLS-119s321is/html/BILLS-119s321is.htm">as introduced</a>: &#8220;the importation into the United States of artificial intelligence or generative artificial intelligence technology or intellectual property developed or produced in the People&#8217;s Republic of China is prohibited&#8221;. Its <a href="https://www.govinfo.gov/bulkdata/BILLSTATUS/119/s/BILLSTATUS-119s321.xml">bill-status record</a>, retrieved September 11th, 2026, shows it introduced, read twice and referred to the Committee on the Judiciary on January 29th, 2025, and nothing since. The No DeepSeek bill&#8217;s approach was enacted for intelligence-agency systems in Public Law 119-60, Sec. 6604 (note [37]).</p><p>[40] Bureau of Industry and Security, Entity List addition, <a href="https://www.govinfo.gov/content/pkg/FR-2025-01-16/html/2025-00704.htm">90 FR 4617, January 16th, 2025</a>, four days before the change of administration on January 20th, 2025, retrieved September 10th, 2026. The rule adds eleven entities under eleven entries; ten of them are named in one determination, &#8220;because these entities advance the People&#8217;s Republic of China&#8217;s military modernization through the development and integration of advanced artificial intelligence research&#8221;, and the eleventh, a lithography company, in a separate one. Seven of the ten carry the Zhipu name: Beijing Zhipu Huazhang, Beijing Zhipu Future, Beijing Zhipu Linghang, Beijing Zhipu Qingyan, Hangzhou Zhipu Huazhang, Shanghai Zhipu Huanyu and Shenzhen Zhipu Future. The other three are Beijing Lingxin Intelligent, Beijing Yuanyin Intelligent and Nanjing Zhihu Information Technology. The advisory&#8217;s Table 1 prints &#21271;&#20140;&#26234;&#35889;&#21326;&#31456;&#31185;&#25216;&#26377;&#38480;&#20844;&#21496; against its Z.AI row, which transliterates to Beijing Zhipu Huazhang Technology Co., Ltd., the first company on the January 2025 list; the match is the author&#8217;s, and the advisory does not mention the Entity List. The Beijing Academy of Artificial Intelligence was added on <a href="https://www.govinfo.gov/content/pkg/FR-2025-03-28/html/2025-05427.htm">March 28th, 2025</a>. Both entries remain in the current eCFR, Part 744, Supplement No. 4, retrieved September 10th, 2026.</p><p>[41] Chris Lehane, OpenAI, &#8220;<a href="https://openai.com/index/ai-policy-window/">The AI policy window is open. We need to act.</a>&#8220;, September 9th, 2026. openai.com returns HTTP 403 to automated fetch; read via the <a href="https://web.archive.org/web/20260910034521id_/https://openai.com/index/ai-policy-window/">Wayback capture</a> of September 10th, 2026. Author&#8217;s search of the full text, same date: &#8220;distill&#8221; 0, &#8220;China&#8221; 0, &#8220;Chinese&#8221; 0, &#8220;CISA&#8221; 0, &#8220;FBI&#8221; 0, &#8220;advisory&#8221; 0. OpenAI is a signatory to the July open letter, which it joined after the launch roster of 25; the post says &#8220;As we affirmed in signing the Open Weights and American AI Leadership letter&#8221;.</p><p>[42] On innocent variation: OpenAI&#8217;s <a href="https://developers.openai.com/api/docs/guides/prompt-caching">prompt caching guide</a>: &#8220;Prompt caching does not change how the model generates output tokens. The model generates a new response using the cached prefix, so identical requests are not guaranteed to produce identical outputs.&#8221; Anthropic&#8217;s <a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching">prompt caching documentation</a>: &#8220;The response you receive is identical to what you would get if prompt caching were not used.&#8221; On reasoning tokens: OpenAI&#8217;s <a href="https://developers.openai.com/api/docs/guides/reasoning">reasoning guide</a>says reasoning tokens &#8220;are billed as output tokens&#8221; and reports them in <code>output_tokens_details.reasoning_tokens</code>; Anthropic&#8217;s <a href="https://platform.claude.com/docs/en/build-with-claude/adaptive-thinking">adaptive thinking documentation</a>: &#8220;the model evaluates each request and decides for itself whether to think and how much&#8221;, and &#8220;To see how many billed output tokens were spent on internal reasoning, read usage.output_tokens_details.thinking_tokens in the response&#8221;; Google&#8217;s <a href="https://ai.google.dev/gemini-api/docs/thinking">Gemini thinking guide</a>, last updated September 9th, 2026, reports <code>total_thought_tokens</code>. Because thinking is adaptive, single counts vary; the test compares averages. Artificial Analysis, &#8220;<a href="https://artificialanalysis.ai/articles/endpoint-accuracy-index">Launching the Endpoint Accuracy Index: Same Model, Different Accuracy</a>&#8220;, August 4th, 2026, on open-weight endpoints: &#8220;Endpoints that score below the reference generally produce fewer output tokens per task. Output limits and reduced reasoning effort show up directly in token usage&#8221;. Irena Gao, Percy Liang and Carlos Guestrin, &#8220;<a href="https://arxiv.org/abs/2410.20247">Model Equality Testing: Which Model Is This API Serving?</a>&#8220;, ICLR 2025: applied to &#8220;commercial inference APIs from Summer 2024 for four Llama models&#8221;, the test found &#8220;11 out of 31 endpoints serve different distributions than reference weights released by Meta&#8221;, using &#8220;an average of just 10 samples per prompt&#8221;; a statistical test against published weights, which closed models don&#8217;t have. The watch-list items are in the advisory: page 2, &#8220;immediate maximum usage from new accounts&#8221;; page 8, &#8220;multiple accounts with similar registration details and payment methods&#8221;; page 9, &#8220;identical or similar prompt texts&#8221;. All retrieved September 11th, 2026.</p><p>[43] OpenAI, &#8220;<a href="https://cdn.openai.com/osa/openai-services-agreement.pdf">Services Agreement</a>&#8220;, version v.010126, retrieved September 11th, 2026. Section 2.3: &#8220;If an OpenAI update materially reduces the Services functionality, OpenAI will notify Customer at the Account email address.&#8221; Section 8.2, on limiting or suspending access: &#8220;OpenAI will use reasonable efforts to notify Customer before limiting to or suspending the Services pursuant to the preceding sentence but may do so without prior notice to the extent reasonably necessary.&#8221; Section 2.3 continues: &#8220;Within five business days of receipt of this notice, Customer may choose to terminate the Agreement by providing thirty days written notice.&#8221; Section 8.2&#8217;s grounds: &#8220;(a) it is required to do so by law; (b) Customer violates the Agreement or OpenAI Policies; or (c) doing so is necessary to prevent or terminate a Security Emergency.&#8221; Section 12.1 warrants that the Services &#8220;will conform in all material respects with the Documentation&#8221;.</p><p>[44] &#8220;<a href="https://www.airealist.ai/p/too-dangerous-for-you-free-for-everyone">Too Dangerous for You, Free for Everyone</a>&#8220;, June 28th, 2026.</p>]]></content:encoded></item><item><title><![CDATA[The Cache Is the Price]]></title><description><![CDATA[Routing a job to a cheaper model looks like free money. Whether it is depends on how the frontier bills the handover.]]></description><link>https://www.airealist.ai/p/the-cache-is-the-price</link><guid isPermaLink="false">https://www.airealist.ai/p/the-cache-is-the-price</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Tue, 08 Sep 2026 13:45:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2XDL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2XDL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2XDL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2XDL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2XDL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2XDL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2XDL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:656981,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/214719963?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2XDL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2XDL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2XDL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2XDL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F106289ae-3518-45bf-9737-a0cc2e4fbdc4_1456x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On September 1st, Anthropic published a rate for Claude Fable 5.1: $10 per million input tokens and $50 per million output tokens.[1] On September 3rd, OpenAI published a rate for GPT-6 Astra: $10 per million input tokens and $50 per million output tokens.[2] Two frontier laboratories, forty-eight hours apart, at the same two numbers.</p><p style="text-align: justify;">Read the sheets past the headline, and they are not the same price.[3][4] For the agentic, context-heavy work both companies are selling into, the columns nobody quotes differ by a factor of four and by an entire pricing tier.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!o_8G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!o_8G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png 424w, https://substackcdn.com/image/fetch/$s_!o_8G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png 848w, https://substackcdn.com/image/fetch/$s_!o_8G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png 1272w, https://substackcdn.com/image/fetch/$s_!o_8G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!o_8G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png" width="1456" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:127361,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/214719963?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!o_8G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png 424w, https://substackcdn.com/image/fetch/$s_!o_8G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png 848w, https://substackcdn.com/image/fetch/$s_!o_8G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png 1272w, https://substackcdn.com/image/fetch/$s_!o_8G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1821051-485c-4aaf-8bd1-e189f70e7e6c_2400x1038.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Rates from each vendor&#8217;s own page.[3][4]</em></figcaption></figure></div><p style="text-align: justify;">Prices per token keep falling. Bills keep rising. The fix everyone reaches for is routing: <strong>send each job to a cheap model first, and pass it to a frontier model only when the cheap one fails</strong>. Done right, that pays for two reasons. The cheap model does most of the work, so the frontier bills only for the share that fails. And the context the frontier has to rebuild when it takes over is paid for once and then read at a fraction of the price for the rest of the session, which is what a prompt cache is for. Done wrong, the same router costs more than the frontier alone.</p><p style="text-align: justify;">This piece is about that pass, the moment a job one model has been working on is handed to another, because that is where right and wrong part company. Some jobs can be handed over for almost nothing. Others cost more to hand over than the savings from the cheaper model. The <strong>Good-Enough Line</strong> runs between those two kinds of jobs. It does not run between frontier models and cheap ones.</p><p style="text-align: justify;">What makes a handover expensive is everything that has piled up behind the job: the instructions, the documents, the results of every tool it has called, every turn of the conversation so far. The model that has been working on it has read all of that. The model taking over has not. Handing over means sending it all again, in full, and paying for it. Both vendors publish every rate you need to price that. What neither documents is how a rebuilt prefix is billed when a job arrives from another model, and no benchmark measures how often you would hand one over. That is why no vendor can quote your line.</p><p style="text-align: justify;">Where the line falls depends on what you have built around the models: the rule that decides which model sees a job, the check that decides whether the cheap one got it right, and the pile of instructions you have written to make either one work. Practitioners call that pile the harness and the runbook. Four things set the line: what the handover costs; how often the cheap model&#8217;s work fails a check; what running that check costs; and what it costs to keep the harness current across a model release every few weeks. The first is what this piece is about.</p><p style="text-align: justify;">A word on the column that does the damage, because it is the least understood line on either sheet. These models remember nothing between calls. Every turn of a conversation resends the whole of it: the system prompt, the documents, the tool output, every previous exchange. By the twentieth turn, you are paying to send the first nineteen again, at full input price, for the twentieth time.</p><p style="text-align: justify;">Prompt caching stops that. The provider keeps the prefix it has already processed, so you pay full price only for what is new in this turn and a fraction of it for everything before. Take a fixed two-hundred-thousand-token context over twenty turns. Uncached, it costs forty dollars. Cached on Anthropic&#8217;s sheet, it costs about three and a half; the prefix is written once at $12.50 per million, and read nineteen times at 25 cents per million. On OpenAI, the same twenty turns cost just over six, because its cache read is four times Anthropic&#8217;s.[5] Same headline rate, nearly double the bill, on one conversation.</p><p style="text-align: justify;">One thing stops it from being free money. Writing to the cache costs more than sending fresh, a quarter more on both sheets and twice as much for Anthropic&#8217;s hour-long window.</p><p>Past that, the savings compound, and four things make it compound the hardest:</p><ul><li><p>Long sessions, because every turn reads back everything before it. </p></li><li><p>Large fixed context, because the system prompt, the codebase, and the documents are written once and read all day. </p></li><li><p>Tool results, because each is appended to the prefix and re-read on every subsequent turn, so an agent that calls twenty tools has bought itself twenty more things to read cheaply.</p></li><li><p>Retries, which resend a context the provider has already processed, are therefore nearly free.</p></li></ul><p style="text-align: justify;">That last one is worth holding on to. A retry within a single model costs almost nothing. The same retry sent to a different model is a full rebuild because the cache is stored per model; a task handed to another vendor or to another model on the same sheet leaves it behind.[6] The more of your bill the cache saves today, the more you lose the moment the job switches to a different model.</p><h2><strong>Run the numbers on your own bill</strong></h2><p style="text-align: justify;">Take a team spending a little over two million dollars a year on agentic coding: about thirty-five billion input tokens a month against 1.8 billion output tokens, four-fifths of the input served from cache and the rest written to it. Neither vendor publishes a cache share, so I assume four-fifths, and every figure below is based on that assumption. On Anthropic&#8217;s sheet, that is about $183,000 a month.[7] Move the identical workload to OpenAI at the same headline rates; it is about $204,000. The whole of the difference is one column: seventy-five cents per million tokens of cache read, across twenty-eight billion cached tokens a month, amounts to $21,000 per month and about $254,000 per year.[7] Read your own off your own bill; the rate is the thing, not the total.</p><p style="text-align: justify;">One assumption is worth more than the finding. The comparison counts a token at one vendor as a token at the other, and the two bills differ by only 11.6%. Anthropic&#8217;s own pricing page records a 30% token-count difference between its own tokenizer generations,[3] so a cross-vendor gap of that order in Anthropic&#8217;s direction would reverse the ranking. What does not move with the tokenizer is the ratio between the two cache-read columns, which is the number this piece is about.</p><p style="text-align: justify;">Now put a router in front of it. Send everything to DeepSeek&#8217;s V4-Pro first at its peak rates, and the same number of tokens costs about $17,500 a month, <strong>under a tenth of the frontier bill</strong>.[7] It is priced here because it is the cheapest published frontier-class rate, not because it is the one your procurement will clear. A tenth of the price looks like the end of the argument. If handing the hard cases up were free, you could hand up almost everything and still come out ahead.</p><p style="text-align: justify;">Substitute a cheap tier that your procurement will actually clear, and the whole calculation moves, because the cheap pass is the numerator of every break-even below. The nearest rung is on Anthropic&#8217;s own sheet: Opus 5 at five and twenty-five, half of Fable, with cache hits at fifty cents.[3] Go one further down, to Sonnet 5, and the same tokens cost about $41,000 instead of $17,500.</p><p style="text-align: justify;">Handing them up is not free, and the cost turns on one thing no rate card answers: whether the frontier model keeps the prefix it has just rebuilt, or pays to rebuild it and then throws it away. When the escalated task arrives at the frontier, its prefix must be rebuilt, and the bill depends on whether that prefix is read back.[7]</p><p style="text-align: justify;">How often you hand up decides the rest, and the only public number for that comes from Together AI, an inference and routing vendor that ran a cheap model first and escalated to a frontier one whenever the tests failed. Its published costs imply it handed up 37.2% of tasks.[8] That was a coding benchmark with a free machine check, on nothing like a cache-heavy workload, and Together claims nothing about this bill. The rollouts and the costs are its run. The escalation rate, the bill, the assumptions, and the conclusion are mine.</p><p style="text-align: justify;"><strong>The break-even escalation rate is the share of jobs you can send to the frontier before the router&#8217;s bill equals the cost of sending all jobs to the frontier</strong>. Below it, routing saves money. Above it, you would have done better paying the frontier rate for everything. There are three rows because the frontier can bill a handed-over job three ways.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5BZ0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5BZ0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png 424w, https://substackcdn.com/image/fetch/$s_!5BZ0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png 848w, https://substackcdn.com/image/fetch/$s_!5BZ0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png 1272w, https://substackcdn.com/image/fetch/$s_!5BZ0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5BZ0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png" width="1456" height="613" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:613,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:183735,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/214719963?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5BZ0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png 424w, https://substackcdn.com/image/fetch/$s_!5BZ0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png 848w, https://substackcdn.com/image/fetch/$s_!5BZ0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png 1272w, https://substackcdn.com/image/fetch/$s_!5BZ0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50e3aa2c-7404-4c92-ad19-19648b581b5b_2400x1010.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Author&#8217;s calculation; assumptions in the notes.[7]</em></figcaption></figure></div><p style="text-align: justify;">Seventy points of spread, and nothing about the models moved to produce it. The cheap model costs the same in all three rows, and so does the frontier model&#8217;s rate. What changes is only what the frontier charges to take on the job. That alone moves the answer from buy to don&#8217;t, which is the argument: <strong>how the frontier bills a handed-over job decides whether routing saves money</strong>, and it decides it harder than the choice of models does.</p><p style="text-align: justify;">Which row you are in is not a vendor secret. It is a property of your own design, and you can read it off your own traffic. Escalate into a session that runs for several turns at the frontier, and the rebuilt prefix is written once and read many times, which is the top row. Escalate one-shot, and it is written and thrown away, which is the bottom row, and Anthropic&#8217;s own pricing page says caching pays for itself after a single read. So the bottom row is a misconfiguration rather than a third possibility, and a competently built cascade lives between a fifty-three per cent saving and a wash.</p><p style="text-align: justify;">The cache share moves the middle reading too: at four-fifths, the middle reading is a wash; at seven-tenths, the router wins; and above nine-tenths, the middle reading breaks even at 29 percent, so a router escalating more than that loses.[7] Before copying any of this, get your own cache share and your own escalation rate.</p><p style="text-align: justify;">So the line falls where the opening said it would, and the ranking of the three rows holds whichever pair of models you put on either side of it. Anthropic&#8217;s own cache diagnostics say the cache is per model and tell you to hold the model constant within a cached conversation, so escalating from a cheap tier to an expensive one on the same sheet drops the prefix exactly as crossing a vendor boundary does.[6] The charge is on the handover, not on the border. <strong>The way out is to hand nothing over: pick the model for each job before the work starts, based on the job type, rather than switching models mid-task after a failure.</strong> That is selection instead of routing, and it is not free either. A partition is only as good as the rule that assigns the job. Escalation puts the cost on the invoice; partition puts it somewhere you have to go and look.</p><h2><strong>The convergence ran one way</strong></h2><p style="text-align: justify;">The convergence was not mutual. Anthropic has charged ten and fifty since Fable 5 and changed exactly one cell when 5.1 shipped, cutting the cache read from a dollar to twenty-five cents.[9] OpenAI climbed. GPT-5.6 Sol launched on July 9th at five and thirty. On August 21st, OpenAI cut it to four and twenty, and called the new rate promotional, running at least to November 21st.[10]</p><p style="text-align: justify;">The multiple depends on which Sol you measure. Compared to the promotional rate, Astra is 2.5 times; compared to Sol&#8217;s July launch price, it is 2 times the input rate.[10] One lab set a number and held it. The other cut the price of its existing model and launched a new one at double the old launch price, which is how it arrived at the same two figures. Every rate card quoted here was captured between August 11th and September 7th, so this is a three-week window rather than a price history.</p><p style="text-align: justify;">The referee moved too. On September 3rd, Artificial Analysis scored Astra level with Sol, both at sixty-one. On September 4th, it shipped a new version of its index, doubled the weight of its private held-out set to forty per cent, and under that version, Astra came out four points ahead of Sol.[11] Three days later, it shipped another, raised the private weighting to forty-five per cent, and under that, Fable 5.1 and Astra tie at fifty-three, with Sol six points behind.[11] Two revisions in four days, each naming its changes, and the order at the top moved with each. It publishes the weight of every evaluation and of each category. What it does not publish is the contents of the private held-out sets, which is the half an outsider would need to separate the model from the ruler. It publishes a cost per index task, and Astra is the cheaper model at $3.26 compared to $7.63 for Fable 5.1.[11] At identical headline rates, that gap is in token volume, not price, and the tokenizer bound above runs in exactly that direction.</p><h2><strong>The premium buys a runbook you did not write</strong></h2><p>Buyers are now running this arithmetic on whether to keep renting a frontier model.</p><p style="text-align: justify;">Together AI put GPT-5.6 Sol against DeepSeek V4-Pro over 904 rollouts on DeepSWE, a software engineering benchmark. Sol won on a single attempt, 72.7 percent against 62.8. At four attempts, the ranking inverts, and the open model takes it, 88.5 percent against 85.8.[8]</p><p style="text-align: justify;">Then Together did what an engineer would do: ran the cheap model first and escalated to Sol only when the tests failed. That cascade solved 83.0 percent at $3.35 per task. It beat Sol alone on accuracy and on price. It also beat a hypothetical perfect router that chooses one model per task in advance, even though the cascade is allowed a second attempt and that router is not.[8] Together sells inference and routing, so read the recommendation with that in mind; the rollouts are still the best public head-to-head with cost attached.</p><p style="text-align: justify;">Note what the cascade is made of. Its 83.0 percent includes every escalated call to Sol, at Sol&#8217;s price. <strong>A routing policy does not take the frontier off the bill. It decides how much of the bill the frontier gets.</strong></p><p style="text-align: justify;">On Terminal-Bench 2.1, the StateM work reported 88.09 percent with a DeepSeek V4-Flash configuration, compared with a GPT-5.6 Sol reference that scored 88.8. That is a match rather than a win, which is the point: the scoring run cost $15.20, or $52.22 once the provider-specific adaptation and the required configuration are included, against $574.68 of the submission-reported cost for the reference.[12] The two cost figures are not measured the same way; one is a bill, and the other is the submission&#8217;s own estimate, so read the gap as an order of magnitude rather than a ratio.</p><p style="text-align: justify;">One condition carries both results. The cascade escalates when the tests fail, and both runs sit on benchmarks that hand you a free, machine-checkable verifier. <strong>Most work has no such check</strong>. Nothing tells you that a summary, a customer reply, or a research memo came back wrong. The escalation rule has nothing to fire on, and the buyer pays the frontier rate as insurance against mistakes they cannot see.</p><p style="text-align: justify;">There is also a cheaper answer that needs none of it. Batch is half off both columns on both sheets; it stacks with caching and requires no router, verifier, or runbook. For anything that can wait, that is the first thing to do, and this piece is not about it.</p><h2><strong>The buyers said it first, including one who holds shares in the seller</strong></h2><p style="text-align: justify;">In May, Marc Benioff described the mechanism on the All-In podcast. He would probably use about $300 million of Anthropic's funds that year on Salesforce coding, he said. That is a personal forecast, not a disclosed commitment. Later in the same conversation, he said the vast majority of those tokens do not need to go to Anthropic, that there needs to be an intermediary layer, and that smaller models can route each job to the most affordable option.[13] He said it as the chief executive of a shareholder in the supplier [14]. </p><p style="text-align: justify;">He is not quite arguing against his own book. Salesforce&#8217;s stake is worth what Anthropic&#8217;s market is worth, not what Salesforce&#8217;s token bill is worth, and the intermediary layer Benioff wants is something Salesforce sells. The chief executive of one of the largest enterprise software companies called the routing policy the obvious engineering answer, at a moment when nothing obliged him to say anything at all.</p><p style="text-align: justify;">On August 27th, Salesforce&#8217;s executive vice president of finance, Mike Spencer, reportedly said that covering token spend was &#8220;part of the reason we didn&#8217;t raise margin guidance on the year&#8221;, and that on many jobs &#8220;you&#8217;re totally fine with the second or third generation model&#8221;.[15][16]</p><p style="text-align: justify;">The second buyer sells the infrastructure rather than the model, so its interest runs in the opposite direction. On September 2nd, the chief financial officer of Hewlett Packard Enterprise said on an earnings call that the company routes each internal workload request to the most cost-effective model across open-source and open-weight options on its own private cloud. On HPE&#8217;s own internal analysis, she said, its private-cloud offering can reduce token costs compared with the public cloud by up to 60%.[17] HPE sells that private cloud, so the 60% is its own product figure, as reported on its earnings call. Read it the way the Together figure is read here: interested, and still the most specific public claim of its kind.</p><h2><strong>What would have to break</strong></h2><p style="text-align: justify;">The strongest argument against all of this is that nobody is doing it yet. Vercel runs a gateway that sends traffic to many laboratories and bills every token at list price. In July, Anthropic handled 30% of the tokens that passed through it and collected 65% of the revenue.[19] Its tokens sold for 4.4 times the average price of everyone else&#8217;s, and that premium was 3.4 times in June. Buyers are paying the premium, and paying more of it each month. Cheap models are winning volume. They are not yet winning budgets.</p><p style="text-align: justify;">That premium is also the test. Watch it rather than Anthropic&#8217;s share of spend, because the share falls whether Anthropic cuts its prices or buyers route work away, and those mean opposite things. If the 4.4 falls toward 1 over the next two monthly reports, Anthropic is closing the gap itself, and a buyer has nothing to route away from.</p><p style="text-align: justify;">The second comes from the supply side. A modelling paper argues that when capacity binds, a degraded cheaper tier needs more attempts per satisfied answer, so the discount can invert: the buyer pays less per call and more per outcome.[20] That is one more cost that lands on the outcome and never on the rate card, and it reaches the bottom row of the table by a second road. If an independent run at a matched retry budget erases the cascade&#8217;s advantage, the premium was buying capability after all.</p><p style="text-align: justify;">Until one of those happens, the sheet says one thing and the bill says another. The charge is on the handover, not on the border. <strong>The numbers that decide what you should do are: how much of your traffic the cache serves, and how often the cheap model has to hand up</strong>. The first is already in the usage object of every call you have paid for. The second costs a pilot. That is the real price of the answer.</p><h3><strong>Notes</strong></h3><p>[1] Anthropic, &#8220;<a href="https://www.anthropic.com/claude-fable-and-mythos-5-1">Claude Fable 5.1 and Claude Mythos 5.1</a>&#8220;, published September 2026, accessed September 7th, 2026, whose own metadata dates it only to the month; the day comes from Anthropic&#8217;s <a href="https://platform.claude.com/docs/en/release-notes/overview">release notes</a>, accessed September 7th, 2026, under the heading &#8220;September 1, 2026&#8221;: &#8220;at $10 / $50 USD per MTok, the same as Claude Fable 5, with cache reads cut to $0.25 per MTok&#8221;. The announcement itself says: &#8220;Fable 5.1&#8217;s pricing is otherwise the same as Fable 5&#8217;s: $10 per million input tokens and $50 per million output tokens.&#8221; The same page sells the cache cut as the saving: &#8220;For typical workloads, costs are reduced by around 25% relative to Fable 5. For complex coding and highly agentic tasks, the savings could be up to around 45%.&#8221; Rates confirmed independently on Google Cloud&#8217;s <a href="https://cloud.google.com/vertex-ai/generative-ai/pricing">Vertex AI partner-model price sheet</a>, which lists Claude Fable 5.1 input $10.00 and output $50.00.</p><p>[2] OpenAI, <a href="https://developers.openai.com/api/docs/pricing">API pricing</a>, accessed September 7th, 2026: gpt-6-astra at $10.00 input, $1.00 cached input, $12.50 cache writes, $50.00 output for standard short-context requests. The same rates appear in Microsoft&#8217;s <a href="https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-intelligence-for-work-now-generally-available-in-microsoft-foundry/">Foundry announcement</a> of September 3rd, 2026. OpenAI&#8217;s marketing pages refuse automated requests; the developer documentation, to which platform.openai.com/docs/pricing redirects, does not.</p><p>[3] Anthropic <a href="https://platform.claude.com/docs/en/about-claude/pricing">pricing documentation</a>, accessed September 7th, 2026: Claude Fable 5.1 at $10 base input, $12.50 five-minute cache writes, $20 one-hour cache writes, $0.25 cache hits and refreshes, $50 output; &#8220;Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier.&#8221; On break-even: &#8220;caching pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write).&#8221; On the tokenizer: &#8220;Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer &#8230; This tokenizer produces approximately 30% more tokens for the same text.&#8221; On context: &#8220;For Claude 4.6 and later models &#8230; include the full 1M token context window at standard pricing.&#8221; Google Cloud&#8217;s <a href="https://cloud.google.com/vertex-ai/generative-ai/pricing">Vertex AI sheet</a> carries the same cache rates from a counterparty rather than from Anthropic.</p><p>[4] OpenAI&#8217;s <a href="https://developers.openai.com/api/docs/models/gpt-6-astra">model page for GPT-6 Astra</a>, accessed September 7th, 2026: &#8220;Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.&#8221; <a href="https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure">Microsoft Learn</a> states the same threshold and the same full-request rule: &#8220;For GPT-6 models, prompts with more than 272,000 input tokens use long-context pricing for the full request, not only for tokens beyond the threshold.&#8221; Two vendors, one figure.</p><p>[5] Author&#8217;s calculation from the rates in notes 1 to 3, for a fixed 200,000-token prefix over twenty turns. Uncached, twenty turns resend four million tokens at $10 per million; author&#8217;s calculation: 4 &#215; 10 = 40. Cached on Claude Fable 5.1, the prefix is written once at $12.50 per million, $2.50, and read nineteen times at $0.25 per million, $0.95, so $3.45 in total. On GPT-6 Astra the write is the same $2.50 and the nineteen reads cost $3.80 at $1.00 per million, so $6.30. The two cached bills differ by $2.85 on one conversation, at identical headline rates.</p><p>[6] Anthropic, <a href="https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics">cache diagnostics</a>, accessed September 7th, 2026: &#8220;The cache is per-model. Hold the model constant within a cached conversation.&#8221; Checked against the <a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching">prompt-caching page</a>, accessed September 7th, 2026, whose table of what invalidates a cache carries no row for the model, and against the <a href="https://www.anthropic.com/pricing">pricing page</a>, which does not record it either: the rule is documented in one place. OpenAI&#8217;s <a href="https://developers.openai.com/api/docs/guides/prompt-caching">caching guide</a>, accessed September 7th, 2026, states only the weaker version, that &#8220;a different model can use different weights and caching behavior&#8221;, so do not read its position as absolute. One cut runs the other way and is recorded here for fairness: OpenAI&#8217;s <a href="https://developers.openai.com/api/docs/changelog">changelog entry</a> for GPT-6 Astra of September 3rd, accessed September 7th, 2026, lets a caller change reasoning effort mid-conversation while preserving the cached prefix, where Anthropic&#8217;s invalidation table lists the effort setting as cache-invalidating. On that axis, which this piece does not price, OpenAI&#8217;s cache is the more forgiving of the two.</p><p>[7] Assumptions are mine and stated: about 35.3 billion input tokens a month, four fifths served from cache and the rest written to it, and 1.76 billion of output. Rates from Anthropic&#8217;s <a href="https://www.anthropic.com/pricing">pricing page</a> and OpenAI&#8217;s <a href="https://developers.openai.com/api/docs/models/gpt-6-astra">model page</a>, both accessed September 7th, 2026, and from the DeepSeek schedule below. Counting writes at $12.50 per million, Anthropic&#8217;s month is $88,250 of writes, $7,060 of cache reads and $88,000 of output, about $183,300, or $2.2m a year; OpenAI&#8217;s differs only in the cache column, at $28,240, about $204,500. The gap is the $0.75 per million between the two cache rates across 28,240 million cached tokens; author&#8217;s calculation: 28240 &#215; 0.75 = 21180. That is about $254,000 a year, and because the write rate is identical on both sheets, counting writes raises both bills without moving the gap. The same tokens on DeepSeek V4-Pro at peak cost about $17,500, at $0.044 per million on cache hits, $1.32 on cache misses and $3.96 on output, all three transcribed from the image table in DeepSeek&#8217;s <a href="https://api-docs.deepseek.com/news/news260813">pricing announcement</a> of August 13th, 2026, which publishes the schedule as pictures rather than as page text. Break-even escalation is (183,310 &#8722; 17,531) &#247; the cost of escalating everything, which is $183,310 if the escalated task rebuilds and reuses its own cache, $441,000 if every token goes fresh at $10, and $529,250 if the rebuilt prefix bills as a cache write at $12.50: author&#8217;s calculation: 165779 &#247; 183310 = 90.4%. The other two denominators give 37.6 and 31.3 per cent. The model applies no long-context multiplier on either sheet, which is conservative in OpenAI&#8217;s favour: above 272,000 tokens its input and cache rates double, so including it widens the gap rather than narrowing it. One further assumption runs the other way. The break-even is a share of tokens and Together&#8217;s rate is a share of tasks, and a cascade that escalates on test failure escalates the long, context-heavy ones by construction, so their token share exceeds their task share and every break-even above is a little generous to the router. Anthropic says of one workload it built that cache reads were most of the cost, but that is a cost share on Fable 5 pricing before the cache cut, and it is not this workload: here cache reads are about four per cent of the bill. Chart data in <code>data/routing-break-even.csv</code>.</p><p>[8][] <a href="https://www.together.ai/blog/deepseek-v4-pro-0813-vs-gpt-5-6-sol-on-deepswe-cost-coding-and-routing">Together AI</a>, August 18th, 2026: 904 rollouts across 113 tasks, four trials per model; Sol at 72.7 per cent pass@1 (&#177;2.2) and $8.37 per rollout against V4-Pro at 62.8 per cent (&#177;3.1) and $0.24; at four attempts the order reverses, V4-Pro at 88.5 per cent pass@4 against Sol&#8217;s 85.8, and Together&#8217;s own line on cost is &#8220;The price gap is 35x&#8221;; a Pro-first cascade escalating on test failure at 83.0 per cent and $3.35 per task, above a perfect one-shot oracle router at 80.8 per cent. The escalation rate used in this piece is derived, not published; author&#8217;s calculation: (3.35 &#8722; 0.24) &#247; 8.37 = 37.2%. It also falls out of the single-attempt pass rate, since 100 &#8722; 62.8 = 37.2, which reproduces it without relying on the $0.24 this note goes on to impeach. Together ran a second cascade three days later, &#8220;<a href="https://www.together.ai/blog/glm-5-3-vs-gpt-5-6-sol-on-deepswe-cost-coding-and-routing">GLM-5.3 vs GPT-5.6 Sol on DeepSWE</a>&#8220;, August 21st, 2026, over the same 113 tasks and 904 rollouts: GLM-5.3 at 69.0 per cent (&#177;2.7) and $3.99 a rollout, Sol at 72.7 (&#177;2.2) and $8.37, and &#8220;Run GLM-5.3 first and escalate to Sol only when your test suite rejects the answer: 85.9% solved at $6.61 per task.&#8221; Both routes reproduce there too; author&#8217;s calculation: (6.61 &#8722; 3.99) &#247; 8.37 = 31.3%. The pass-rate route gives author&#8217;s calculation: 100 &#8722; 69.0 = 31.0%. So the public range for a test-failure cascade is roughly 31 to 37 per cent across two model pairs. The 31.3 here is an escalation rate and is not the 31 per cent break-even in the table above, which is a different quantity that happens to land on the same figure. Vendor-run: Together sells both the inference and the routing, and its own caveat is that every figure comes from this run and can differ from other public scorecards. It does not state which rate card it costed against, and the post of August 18th predates OpenAI&#8217;s cut of August 21st, so the Sol side is most likely at the pre-promotional $5 and $30 while the Astra comparisons elsewhere in this piece are at $4 and $20. The open-model side is the harder problem. Together&#8217;s own price list bills DeepSeek V4-Pro output at $3.96 per million and DeepSeek&#8217;s card matches it, so the 101,000 output tokens the post reports per rollout cost about $0.40 before a single input token is counted, against a reported total of $0.24; off-peak at $1.98 still gives $0.20 and leaves four cents for 146 agent steps of input. Together&#8217;s post offers no explanation. The direction matters: correcting it narrows the price gap and raises the cascade&#8217;s cost without reversing either. DeepSWE is also the suite on which V4-Pro moved from 12.8 at preview to 62.7 at general availability, a jump consistent with fitting to the published harness. The independent leg of the same proposition is the <a href="https://arxiv.org/abs/2608.15089">StateM work</a>; see note 12.</p><p>[9] Anthropic&#8217;s own documentation lists the superseded rate: Claude Fable 5 at $10 input, $50 output, $1 cache hits, on its <a href="https://platform.claude.com/docs/en/about-claude/pricing">pricing documentation</a>. The <a href="https://claude.com/pricing">marketing page</a> carries the same figures under legacy models. The ratio between OpenAI&#8217;s cache read for Astra and Anthropic&#8217;s for Fable 5.1, author&#8217;s calculation from note 2 and note 3: $1.00 &#247; $0.25 = 4x.</p><p>[10] OpenAI <a href="https://developers.openai.com/api/docs/changelog">API changelog</a>, entry of August 21st, 2026: &#8220;GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, representing 20% lower input pricing and 33% lower output pricing. GPT-5.6 Sol&#8217;s promotional pricing is available at least through November 21st, 2026.&#8221; The <a href="https://web.archive.org/web/20260811103033/https://developers.openai.com/api/docs/pricing">pre-cut sheet</a>, archived August 11th, 2026, lists gpt-5.6-sol at $5.00 input and $30.00 output. Microsoft&#8217;s <a href="https://azure.microsoft.com/en-us/pricing/details/azure-openai/">Azure OpenAI price page</a> still lists Sol at $5.00 and $30.00, accessed September 7th, 2026.</p><p>[11] Artificial Analysis, &#8220;<a href="https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2">Intelligence Index v4.2</a>&#8220;, September 4th, 2026: GPQA Diamond removed as &#8220;saturated&#8221; with no ceiling published, AA-Briefcase and GDP.pdf added, private held-out weighting doubled to 40 per cent, and Astra placed four points above Sol. The previous day&#8217;s benchmarking piece scored both at 61 under version 4.1.1. <a href="https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3">Version 4.3</a>, September 7th, 2026, accessed September 8th, 2026: &#8220;Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47)&#8221;; &#8220;Evaluations with private questions or answers account for 45% of the Intelligence Index v4.3 weighting, up from 40% in v4.2&#8221;; and on cost: &#8220;GPT-6 Astra (max) and Claude Fable 5.1 (max with fallback) both score 53, but their average cost per Intelligence Index task is $3.26 and $7.63 respectively - 57% lower for Astra.&#8221; Terminal-Bench moved from 2.1 to 4.0 and AutomationBench-AA replaced &#964;&#179;-Banking at the same 5 per cent weight; category weights are unchanged.</p><p>[12][] <a href="https://arxiv.org/abs/2608.15089">StateM</a>, August 15th, 2026, on Terminal-Bench 2.1: 88.09 per cent from a DeepSeek V4-Flash configuration against a GPT-5.6 Sol max reference of 88.8; $15.20 of realised API charges on the scoring run, $52.22 including about $37 of provider-specific adaptation, against $574.68 of submission-reported model cost for the reference, which the paper puts at roughly a ratio of one to eleven. A frozen runbook is worth 9.0 to 10.4 points, the upper figure computed off a headline 95.3 per cent that the paper flags as pre-adjudication with a defensible range of 93.26 to 95.28. Four authors, no declared affiliation, described as personal-time work. Applied unchanged across providers the runbook lowered the score, 82.7 to 82.0, and the adaptation recovered the gain; the authors also disclose that on one task family the profile learned the evaluator&#8217;s boundary convention without reading verifier code, so part of the harness gain is fit to the grader. The independent leg of the same proposition is <a href="https://www.together.ai/blog/deepseek-v4-pro-0813-vs-gpt-5-6-sol-on-deepswe-cost-coding-and-routing">Together&#8217;s run</a>, a different benchmark, a different method and an opposed interest.</p><p>[13] Marc Benioff, <a href="https://www.youtube.com/watch?v=jJRAvZNGUvI">All-In podcast</a>, uploaded May 15th, 2026, at 38:56 and 1:00:29, from the published captions: &#8220;I am going to probably use $300 million of Anthropic this year at Salesforce coding&#8221; and &#8220;the vast majority of those tokens don&#8217;t need to go to Anthropic, there needs to be some intermediary layer &#8230; these ones can handle by smaller models that can route it to the most affordable for the job.&#8221; A personal forecast scoped to coding, not a disclosed commitment: Salesforce&#8217;s 10-Q for the quarter ended July 31st, 2026 records no Anthropic purchase obligation and no significant change to fixed contractual obligations.</p><p>[14] Salesforce <a href="https://www.sec.gov/Archives/edgar/data/1108524/000110852426000190/crm-20260731.htm">Form 10-Q</a> for the quarter ended July 31st, 2026, EDGAR acceptance August 26th, 2026 22:54:46 UTC: a strategic investment portfolio of over 450 companies carrying $11.3 billion, &#8220;including the Company&#8217;s investment in Anthropic PBC &#8230; which represented approximately $5.1 billion&#8221;, about 45 per cent of the portfolio against 22 per cent at January 31st, and &#8220;Upward adjustments for the three and six months ended July 31st, 2026 include unrealized gains of $2.7 billion and $3.0 billion, respectively, related to the Company&#8217;s investment in Anthropic&#8221; &#8212; the body uses the six-month figure, because the period it describes runs from January to July. The portfolio was $7,591m at January 31st and $11,324m at July 31st; author&#8217;s calculation: 7591 &#215; 0.22 = 1670. The quarter&#8217;s <a href="https://www.sec.gov/Archives/edgar/data/1108524/000110852426000187/crm-q2fy27xexhibit991.htm">earnings exhibit</a> gives &#8220;Gains (losses) on strategic investments, net&#8221; of $2,613m against &#8220;Income from operations&#8221; of $2,331m, and states that strategic-investment gains &#8220;impacted GAAP diluted net income per share by $2.43&#8221;, which is where the below-the-line placement is proved.</p><p>[15][] <a href="https://www.theregister.com/ai-and-ml/2026/09/03/salesforce-blames-its-claude-addiction-for-denting-profit-margin-guidance/5294219">The Register</a>, September 3rd, 2026, reporting Mike Spencer at the Deutsche Bank Technology Conference of August 27th: &#8220;It&#8217;s part of the reason we didn&#8217;t raise margin guidance on the year because we&#8217;re covering some of the token spend that we&#8217;ve got going&#8221;, and &#8220;You&#8217;re totally fine with the second or third generation model.&#8221; Salesforce&#8217;s investor-relations page confirms the event, the date of August 27th and the speaker, and gives his title as executive vice-president of finance where the outlet gives deputy chief financial officer. The webcast sits behind a registration gate and the company has published no transcript, so this is one outlet&#8217;s note of a session and the piece attributes it as reported rather than stating it.</p><p>[16] Salesforce <a href="https://www.sec.gov/Archives/edgar/data/1108524/000110852426000187/crm-q2fy27xexhibit991.htm">Form 8-K exhibit 99.1</a> for the second quarter of fiscal 2027, EDGAR acceptance August 26th, 2026 20:03:53 UTC: &#8220;GAAP operating margin of 20.5% and non-GAAP operating margin of 34.1%&#8221; and &#8220;Updates full year FY27 GAAP operating margin guidance to 20.1%, and maintains non-GAAP operating margin guidance of 34.3%.&#8221; The <a href="https://www.sec.gov/Archives/edgar/data/1108524/000110852426000125/crm-q1fy27xexhibit991.htm">first-quarter exhibit</a>, accepted May 27th, 2026 20:18:26 UTC, gives the prior figure of 20.6 per cent. Both are full-year FY27 guidance reconciliations rather than quarter actuals, each calculated on the midpoint of the revenue guidance range. The second-quarter exhibit moves amortisation of purchased intangibles from 4.2 to 4.4 per cent and restructuring and acquisition costs from 0.5 to 0.8, while stock-based compensation is flat at 9.0 in both. So each side sums on the page: 20.6 + 4.2 + 0.5 + 9.0 = 34.3, and 20.1 + 4.4 + 0.8 + 9.0 = 34.3. The half point is entirely in the two lines the body names.</p><p>[17] Marie Myers, chief financial officer of Hewlett Packard Enterprise, on its third-quarter fiscal 2026 earnings call of September 2nd, 2026. The wording below comes from a third-party transcript rather than a company one: &#8220;HPE is now deploying an internal agentic AI platform built on our own private cloud AI, open source, and open weight models, leveraging intelligent routing that sends each workload request to the most cost effective AI model. According to our own internal analysis, our PCAI offering can reduce token costs versus the public cloud by up to 60%. Routine tasks stay on premise while frontier models are reserved for the most complex work.&#8221; <a href="https://singjupost.com/transcript-hewlett-packard-hpe-q3-fy26-earnings-call/">Transcript</a>; a third-party transcript, not the company&#8217;s own. Confirmed independently by Constellation Research&#8217;s <a href="https://www.constellationr.com/insights/news/hpe-delivers-strong-q3-ups-fiscal-2026-outlook-due-ai-demand">report of the same call</a>, accessed September 7th, 2026, which carries both the up-to-60-per-cent projection and the same routing description. HPE&#8217;s own <a href="https://investors.hpe.com/%7E/media/Files/H/HP-Enterprise-IR/documents/q3-2026/q3-2026-earnings-press-release.pdf">earnings release</a> does not carry the claim, so it exists only in the call.</p><p>[18][] <a href="https://vercel.com/blog/deepseek-overtakes-google-on-volume-cost-per-token-falls">Vercel</a>, August 11th, 2026, reporting July. Vercel operates a competing gateway and computes spend at published list price rather than at invoiced rates, which if anything flatters the frontier. On the 13.6 per cent, the same post is explicit that it is not a rate cut: &#8220;The entire decline in average price came from what companies chose to route&#8221;, and &#8220;holding June&#8217;s mix of models constant, the average price would have held essentially flat instead of declining.&#8221; On the spend share: &#8220;In July, it collected 65.1% of all spend on 30% of total volume. The average price per Anthropic token ran 4.4 times the average across every other lab, up from 3.4 in June.&#8221; So 3.4 times in June against 4.4 times in July. The figures are not stable month to month: 61 per cent of spend on 32 per cent of tokens in June, 65 on 32 in May, per the <a href="https://vercel.com/blog/ai-gateway-production-index-july-2026">preceding index</a>.</p><p>[19] The <a href="https://arxiv.org/abs/2608.23986v2">shadow-price analysis</a>, August 29th, 2026, which models attempts per satisfied answer as 1 &#247; (1 &#8722; d&#961;) and observes that the discount inverts exactly when capacity binds. Its own section five calls the work a proof of concept rather than an empirical claim; it is cited here as a mechanism, not as a measurement.</p><p>[20] Salesforce, &#8220;<a href="https://www.salesforce.com/news/stories/agent-fabric-control-plane-announcement/">Salesforce Advances Agent Fabric</a>&#8220;, accessed September 7th, 2026: &#8220;AI Gateway, MCP Bridge, and Trusted Agent Identity with mobile authorization for high-risk agent actions are generally available today&#8221;, dated April 15th, 2026, and of AI Gateway: &#8220;Standardize token management and compliance across your entire multi-LLM stack. Enforce routing rules, unify access, and control costs from a central point.&#8221; The layer therefore predates the May remark quoted above. Salesforce also published the practice itself: Jayesh Govindarajan, &#8220;<a href="https://www.salesforce.com/news/stories/cutting-inference-spend-by-right-sizing-models/">How We Cut Inference Spend by Right-Sizing Our Models</a>&#8220;, July 8th, 2026 (<a href="https://web.archive.org/web/20260908094010/https://www.salesforce.com/news/stories/cutting-inference-spend-by-right-sizing-models/?bc=HL">archived copy</a>, since the site files it under a tracking variant of its own address), describing an agentic harness routing work across five purpose-built in-house models and reserving the frontier model for multi-step reasoning. A defender would call those auxiliary utility models rather than a frontier cascade, which is fair; the falsifier as originally written did not survive either document.</p>]]></content:encoded></item><item><title><![CDATA[Event Horizon]]></title><description><![CDATA[Nobody could pay to keep Hugging Face neutral. Nvidia is paying to own it.]]></description><link>https://www.airealist.ai/p/event-horizon</link><guid isPermaLink="false">https://www.airealist.ai/p/event-horizon</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Mon, 31 Aug 2026 14:36:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ftae!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ftae!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ftae!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ftae!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ftae!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ftae!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ftae!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:501240,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/213540741?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ftae!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ftae!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ftae!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ftae!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5984114-2661-4ac6-b13c-e3eda4a8fa90_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">Nvidia is reportedly buying Hugging Face for $12.9bn. In March, I <a href="https://www.airealist.ai/p/open-source-closed-orbit">asked</a> whether anyone else would pay to keep the hub neutral. The reported answer says more about how the company was built and what it cost to run this summer than about the buyer.</p><p><em>Disclosure: I was chief evangelist at Hugging Face from 2021 to 2024. I hold equity from that time.</em></p><p style="text-align: justify;">On 23 August, Business Insider reported that Hugging Face had hired a bank to gauge interest in a sale at $13bn or more.[1] On 26 August, The Information reported that Nvidia had agreed to buy it for $12.9bn; Reuters relayed the report that evening, and on 27 August, Bloomberg, citing Business Insider, reported only that Nvidia had discussed a purchase.[2] Five days on, as of 31 August, neither company has commented, and there is no filing; Business Insider reported on the 28th that no contract had been signed and the talks could still collapse; Jensen Huang took an earnings call on the 27th without mentioning the deal, and Hugging Face&#8217;s Thomas Wolf gave an interview on the 29th that discussed it without disputing it.[3] Everything below is conditional on a deal that is reported, not confirmed.</p><p style="text-align: justify;">In March, I argued in <em><a href="https://www.airealist.ai/p/open-source-closed-orbit">Open Source, Closed Orbit</a></em> that &#8220;open&#8221; now means two things: Nvidia&#8217;s, where the weights are free, and every path from download to production bends toward Blackwell, NIM containers and an NVIDIA AI Enterprise license list-priced at $4,500 per GPU per year. Hugging Face is a hub that hosts everyone&#8217;s models and co-maintains, with the chip vendors, the libraries that run them on Trainium, AMD, and others.[4] The piece ended with the question of whether anyone was &#8220;investing hard enough in the alternative.&#8221;</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;6e42c077-b7cf-411f-8265-679e2036e161&quot;,&quot;caption&quot;:&quot;On August 8, 2023, Jensen Huang stood on stage at SIGGRAPH and asked a question.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Open Source, Closed Orbit&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:100614256,&quot;name&quot;:&quot;Julien Simon&quot;,&quot;bio&quot;:&quot;AI Operating Partner @ Fortino - Ex-Arcee AI, Hugging Face, AWS - Tech, Cloud, AI, Chips, Space, VC, PE, Investing - https://julien.org&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17c8a9ff-3d05-4f75-8676-1429259c8a0a_427x427.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-03-13T17:47:00.534Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!skV1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc85c51e8-8e77-4393-baa8-33e91d8111f3_1280x720.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.airealist.ai/p/open-source-closed-orbit&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:189991912,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:3,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1027538,&quot;publication_name&quot;:&quot;The AI Realist&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!u6cR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924ecf6b-2ddb-4f24-a3bd-89ae62c7c1dc_800x800.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h2>Nobody else needed to own it</h2><p style="text-align: justify;">The instinctive question is: How did Google, Amazon, Microsoft, Meta, or Salesforce let Nvidia walk off with the bride? The answer is that none of them was ever going to propose.</p><p style="text-align: justify;">Start with the cap table. Hugging Face last raised in August 2023: $235m at $4.5bn, from Salesforce, Google, Amazon, Nvidia, Intel, AMD, Qualcomm, IBM, and Sound Ventures, most of the hardware and cloud industry, about 5% of the company between them and none with a reported board seat.[5] That structure made the hub neutral. It also made it indefensible: eight of those nine investors had no reason to ever own it, and the ninth was Nvidia.</p><p style="text-align: justify;">Each cloud already has the pipe. Microsoft put more than 10,000 Hugging Face models into Azure AI Foundry in May 2025; Google Cloud built a caching gateway for the hub last November and says 1,500 terabytes of models and datasets move between the two every day; AWS sells hub models through SageMaker and the Bedrock Marketplace.[6] They already get the distribution. Owning the hub would cost them the thing that makes it worth having: the day Google owned the front door, Amazon and Microsoft would build their own, and the traffic that makes the hub worth $13bn would leave with them. To each cloud, a hub nobody owns is worth more than a hub it owns. The only company for which that is not true is the one the hub lets everyone else avoid.</p><p style="text-align: justify;">Then look at who could have written the cheque this month. Google closed Wiz on 11 March after a 12-month review and is fighting a Justice Department cross-appeal that seeks to bar Chrome and Android. Amazon has $11.6bn for Globalstar awaiting approval, an FTC trial in March 2027, and a long habit with open source: on the same day The Information reported the Nvidia deal, AWS announced it was buying DuckLabs, the thirty-person company behind DuckDB, for an undisclosed sum, while the code stays with an independent foundation under MIT.[7] Amazon hires the maintainers and sells the service; it does not buy the project, and this year, the big cheque, $50bn, went to a stake in OpenAI. Microsoft already owns GitHub, and owning the GitHub of models as well would invite a long antitrust review. Meta structured its $14bn Scale deal with a 49% non-voting stake precisely to avoid review, and is named in the senators&#8217; letters about that structure. Salesforce, which led the last round, is halfway through a $25bn accelerated buyback, and its chairman&#8217;s only public word on the subject this summer was &#8220;Mahalo, Jensen&#8221; on joining Nvidia&#8217;s open-weights coalition.[8] None of them has ever counter-bid a rival&#8217;s control deal to protect a minority position, and a bank process that leaked on a Saturday and produced a reported agreement by Wednesday gave them days, not months.[9]</p><p style="text-align: justify;">Nvidia had none of those constraints. It had been at the table since its refused offer last year. It closed the quarter with $99bn of cash and securities and bought back $39bn of stock in six months; $12.9bn is thirteen days of data-center revenue.[10] At that scale, the price barely registers, and a buyer who wants the asset will always outbid one who only wants the return. And the structure it used for Groq and Poolside, license the technology and hire the people, was not available here: the hub&#8217;s asset is not engineers or IP but traffic and trust, and those transfer only with the company. Six days after a $7bn license-and-hire for Poolside built to stay outside merger review, Nvidia reportedly signed its first large deal in a year that will have to go through it.[11] <strong>That is how much it wanted this one</strong>.</p><h2>Why now</h2><p style="text-align: justify;">Nine months ago, the answer was no. In late 2025, Nvidia offered $500m at a $7bn valuation; Hugging Face refused because, according to the Financial Times, it did not want a single dominant investor that could sway its decisions.[12] The company was not distressed then and is not now: profitable in 2025 by the FT&#8217;s account, &#8220;close to profitability&#8221; by its CEO&#8217;s, with about half of the money it ever raised still in the bank.[13] Three things changed between January and August, and a fourth had been true for a year.</p><p style="text-align: justify;"><strong>The first is the pric</strong>e. Annualized revenue went from about $100m to more than $150m in two months, and a company sells at the top of its growth rate, not after it.[14] Roughly 85x run-rate is a strategic multiple, above Figma (about 50x, 2022) and Wiz (about 45x, 2025), and a refused $7bn mark, plus a year of growth and a control premium, lands at $12&#8211;14bn mechanically.[15] At $12.9bn, every preferred share converts, and a decade of employees get paid, against an IPO path that exists but requires a ton of corporate work and is subject to market conditions. OpenAI and Anthropic both filed confidentially in June, and OpenAI&#8217;s CFO has told staff it &#8220;will be a public company in 2027&#8221;; once a frontier lab trades publicly, AI companies will have a public price, and $12.9bn negotiated before that price exists is a better deal for the seller than one negotiated after.[16] Cl&#233;ment Delangue told Axios last November that the industry was in &#8220;an LLM bubble&#8221; that &#8220;might be bursting next year.&#8221;[17]</p><p style="text-align: justify;"><strong>The second is the bill for neutrality</strong>. From 9 to 13 July, an OpenAI evaluation agent that had escaped its sandbox operated inside Hugging Face&#8217;s infrastructure, and by the company&#8217;s own account, its forensics were slowed because the hosted frontier models it tried first refused to help; it finished the analysis with an open Chinese model in Nvidia&#8217;s optimized build.[18] Delangue flew to San Francisco, asked OpenAI for &#8220;radical transparency&#8221; and $100m of compute for community cyber-defense, and has received, as far as the record shows, a technical report and no money; state attorneys general have since subpoenaed OpenAI, and the lawyers quoted on liability say the insurance market for this kind of incident is still being worked out.[19] Business Insider's first report came a month after the intrusion. No reporting connects the two, and the Nvidia courtship predates the breach; I am not claiming the attack sold the company. What it did, at minimum, is reprice the job of staying independent: the cost of hosting everyone's models is now set by frontier-lab incidents, whichever lab's agent it is, and a $150m company cannot or may not want to carry it.</p><p style="text-align: justify;"><strong>The third is that the hub was becoming a channel Nvidia did not control</strong>. Meta signaled in December that its next frontier models would be closed; then came back on 10 August with Muse Glimmer, a 30-billion-parameter open-weight model distilled from the closed Muse Spark; the hub&#8217;s hardware-agnostic patron returned for the small end, not the top.[20] DeepSeek V4, the largest open model on the hub, launched in April, tuned for Huawei&#8217;s Ascend, and ran on it from day one.[21] The Google leg of the vendor-library arrangement went quiet in January, when Optimum TPU was archived.[22] And on 19 August, Stripe bought OpenRouter, the neutral routing layer above the hub, for about $7.5bn.[23] Nvidia&#8217;s own Nemotron 4, at a trillion parameters, is due this fall and needs a place where people will find it.[24] The two neutral layers in OpenAI, discovery and routing, were both for sale in the same fortnight; Stripe took one, and Nvidia took the other.</p><h2>Two governments hold the traffic</h2><p style="text-align: justify;">The fourth reason is the one Hugging Face&#8217;s own reports never name as a risk. Chinese open-weight models accounted for 41% of hub downloads this spring; Qwen alone is downloaded more than the rest of the open ecosystem combined and has five times as many derivatives as Llama, and in August, Delangue said China is &#8220;clearly dominating on open models.&#8221;[25] That traffic is the hub&#8217;s growth, and it is hostage to two governments.</p><p style="text-align: justify;">In Washington, the administration revived a push in July to restrict Chinese models through the Entity List, and two House committees are writing to American companies that deploy them.[26] In Beijing, the Ministry of Commerce spent June convening Alibaba, ByteDance and Zhipu to discuss restricting overseas access to their most advanced models, including open-weight models, under a tiered scheme in which the frontier tier would be domestic-only.[27] Either government can move the hub&#8217;s largest source of traffic. Nothing has been pulled yet: DeepSeek V4 linked its weights only to Hugging Face, Kimi K3 set the hub&#8217;s trending record in July, and Qwen still ships on Hugging Face and ModelScope the same evening.[28] But the supply now depends on a Chinese policy decision and the demand on an American one, and Hugging Face has no influence over either.</p><p style="text-align: justify;">It never had much in Beijing. The site has been blocked in mainland China since September 2023; the Spaces domain followed last November; the mirror everyone there uses is run by someone else, and the China entry the company sketched in 2023, a WeChat account, a localized blog, a lead on stage at the Shanghai conference, never became an entity, a licensed mirror or a cloud partnership. The head of its Asia-Pacific ecosystem, who had walked DeepSeek and Qwen onto the hub, left this summer.[29] ModelScope, Alibaba&#8217;s hub, is a fraction of the size, but it is based in China, whereas Hugging Face is not.[30] Hugging Face&#8217;s Chinese business was always the export of Chinese models to the West, and that is a business two ministries can end.</p><p style="text-align: justify;">Is this Nvidia&#8217;s way back into China? No. Nvidia shipped less than 1% of its data-center revenue to China last quarter, carries none in its outlook, and is under an anti-monopoly investigation in Beijing over the Mellanox conditions; owning a site that is blocked there changes none of that.[31] What changes is where Chinese models go when they leave China. DeepSeek V4 launched tuned for Huawei&#8217;s Ascend; Hugging Face&#8217;s forensics this summer ran on Zhipu&#8217;s GLM-5.2 in Nvidia&#8217;s optimized build, &#8220;the American version of a Chinese model,&#8221; in Delangue&#8217;s phrase.[18] The hub is where a Chinese model, once released, gets packaged and deployed, and whoever owns the hub decides which version is the default. The biggest Chinese models are now built for Huawei&#8217;s chips first; on the hub, they can be made to run on Nvidia&#8217;s first. So this is not Nvidia&#8217;s way into China. It is the way Chinese models come out of China, and Nvidia would own the door. That much is on the record. Whether the door also gives Nvidia leverage with either government is a matter of speculation, and I will not make it here.</p><h2>Where Anthropic is</h2><p style="text-align: justify;">Anthropic is on the other side of every line this piece draws, and that is not an accident. It has never released a model&#8217;s weights. It did not sign Nvidia&#8217;s July letter opposing restrictions on open weights, and it is not part of the Open Secure AI Alliance that Nvidia formed around Hugging Face after the intrusion. Its own statement that week said open models &#8220;that don&#8217;t have dangerous capabilities are a public good.&#8221; The same statement repeated its claim that Alibaba had run the largest distillation attack it had seen, and the New York Times reported that Anthropic and OpenAI were lobbying Washington to restrict Chinese open-weight models.[32] When Hugging Face&#8217;s engineers tried to analyze the attack, the models they reached for first were Claude Opus and Fable, and by Hugging Face&#8217;s account, both &#8220;refused a large part of that work: their safety guardrails treated reverse-engineering an exploit the same as launching one.&#8221;[18]</p><p style="text-align: justify;">Its silicon runs the same way. Anthropic has contracted up to 5 gigawatts of Trainium from Amazon, several gigawatts of TPU from Google, two gigawatts of AMD&#8217;s MI450 from 2027, and about one gigawatt of Nvidia through Microsoft&#8217;s Azure; by announced capacity, Nvidia is a minority supplier to the lab that sells the leading closed model, and Amazon and Google are its two largest shareholders.[33] Nvidia&#8217;s CFO put it in one line on Wednesday: &#8220;NVIDIA runs the leading closed models&#8221;; &#8220;nearly all open models run on NVIDIA.&#8221;[31] It runs them today. The closed labs are building the exits.</p><p style="text-align: justify;">So are the trenches being dug? The open-versus-closed line is real, and everyone can see it: Nvidia, Meta, Hugging Face, Mistral and the Chinese labs on one side; Anthropic, OpenAI and, mostly, Google on the other. The silicon line is messier than it looks. Google signed Nvidia&#8217;s letter while selling Anthropic TPUs; AMD signed it days before putting $5bn into Anthropic; Microsoft is Nvidia&#8217;s biggest customer and Anthropic&#8217;s third cloud; Amazon kept Anthropic as its primary lab and put $50bn into OpenAI anyway.[34] What lines up is simpler. Nvidia&#8217;s growth now comes from open models, and its largest customers are spending less on it. Buying the place where open models live is the response to both.</p><h2>Nothing needs to be switched off</h2><p style="text-align: justify;">The honest counter-examples are GitHub, still broadly neutral as a code host eight years after Microsoft paid $7.5bn for it, though its formal independence ended in 2025, and OpenRouter itself. The difference is what the buyer sells. Microsoft does not sell the compute every repository compiles for, and Stripe sells no models. Nvidia sells the layer every model on the hub runs on, and for well over a decade, from the Linux driver fight to Mellanox to Grace and Vera, it has used each layer it owns to secure the one below.[35]</p><p style="text-align: justify;">That history also answers the objection that an acquirer will destroy what makes the hub valuable. Nvidia needs the hub full: it wrote the July letter against restricting open weights, Chinese ones included, it ships Nemotron there, and a hub nobody visits routes nothing.[24] So nothing gets switched off. It gets narrowed. Mellanox kept selling InfiniBand to everyone; the firmware notes now say NVIDIA &#8220;does not support InfiniBand cables or modules not qualified or approved by NVIDIA&#8221;, which every switch vendor says. That is the point: nothing is banned, but there is a list.[36] Run:ai&#8217;s founders said the software would be open-sourced; the scheduler was, the platform was not.[37] vLLM is still open; NIM 2.0 is now built on vLLM alone, inside a licensed container that is free to try and paid in production.[38] Nvidia, which reported $89bn of data-center revenue for the quarter on 26 August, is not buying a hub to sell more GPUs. It is buying the paid software that attaches to every download and the right to choose the default for models that would otherwise run on Trainium or TPUs.[10]</p><p style="text-align: justify;">Apply that to the hub, and nothing happens on day one. The models stay; the Chinese models stay, now with a lobbyist. What changes is the direction the page tilts: NIM moves to the top of the Deploy menu, and enterprise features now include a line referencing NVIDIA AI Enterprise. The Optimum libraries are permissively licensed, and the vendors&#8217; own engineers already maintain them; even the default endpoint is one environment variable away, as China&#8217;s mirrors prove daily.[39] What cannot be forked is the traffic: the download counts, the discovery graph, the place everyone looks first. Each step is defensible.</p><h2>Verdict</h2><p style="text-align: justify;">Hugging Face did not lose its neutrality. It was structured so that nobody could pay to keep it, and this summer it found out what keeping it costs. The one company that gains nothing from the hub&#8217;s neutrality is the one that gains most from owning it, and it made the only reported offer. If a rival bid surfaces before signing, that argument is wrong, and I will say so. And if, a year after closing, the Deploy menu still ranks by anything other than Nvidia attach, the second half of this piece is wrong too.</p><p style="text-align: justify;">In March, I wrote that Nvidia&#8217;s kind of open is a black hole: nothing that crosses the event horizon comes back out. The physics adds one detail worth keeping. From the outside, a crossing is invisible. The star never winks out; it hangs at the horizon and slowly dims. A year from now, the hub will still be there, and the models will still be open. Mellanox still sells InfiniBand to everyone. Run:ai&#8217;s scheduler is still open source. vLLM is still free. It is what crossing looks like.</p><div><hr></div><h3>Notes</h3><p>[1] Business Insider, 23 Aug 2026, as relayed by <a href="https://finance.yahoo.com/technology/ai/articles/hugging-face-exploring-sale-valuing-200012818.html">Reuters</a> and <a href="https://www.bloomberg.com/news/articles/2026-08-23/hugging-face-gauging-interest-for-potential-sale-business-insider-says">Bloomberg</a>. Business Insider cited annualized revenue of more than $100m. Bank not named; no bidder named.</p><p>[2] The Information, <a href="https://www.theinformation.com/articles/nvidia-agrees-buy-open-source-model-repository-hugging-face-12-9-billion">&#8220;Nvidia Agrees to Buy Open Source Model Repository Hugging Face For $12.9 Billion&#8221;</a>, 26 Aug 2026; Reuters relayed the same evening US time; Bloomberg, <a href="https://www.bloomberg.com/news/articles/2026-08-27/nvidia-discussed-buying-ai-startup-hugging-face-insider-says">&#8220;Nvidia Discussed Buying AI Startup Hugging Face, Insider Says&#8221;</a>, 27 Aug 2026. Gizmodo reported the deal as &#8220;still being finalized.&#8221; Neither company responded to Reuters. No 8-K would necessarily be required at this size relative to Nvidia&#8217;s market capitalization.</p><p>[3] Business Insider, 28 Aug 2026, via <a href="https://www.techtimes.com/articles/325863/20260828/nvidias-129b-hugging-face-deal-must-pass-antitrust-review-its-quasi-mergers-dodged.htm">TechTimes</a>: no signed contract; talks could collapse. Thomas Wolf interview, <a href="https://thenextweb.com/news/hugging-face-nvidia-acquisition-thomas-wolf-interview">The Next Web, 29 Aug 2026</a>. Nvidia Q2 FY2027 call, 26&#8211;27 Aug: no mention of the deal.</p><p>[4] Julien Simon, <a href="https://www.airealist.ai/p/open-source-closed-orbit">&#8220;Open Source, Closed Orbit: The Hardware Monopolist&#8217;s Guide to Owning Open Source&#8221;</a>, The AI Realist, 13 Mar 2026. Pricing: <a href="https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/pricing.html">NVIDIA AI Enterprise Licensing Guide</a>, updated 8 Jun 2026: $4,500 per GPU per year, self-managed subscription. Libraries: <a href="https://github.com/huggingface/optimum-habana">Optimum for Intel Gaudi</a> (Apache-2.0, last release Apr 2026), <a href="https://github.com/huggingface/optimum-neuron">Optimum Neuron</a> (Apache-2.0, Feb 2026), and the AMD kernels in hf-rocm-kernels. The March piece also listed Optimum TPU; see note 22.</p><p>[5] TechCrunch, <a href="https://techcrunch.com/2023/08/24/hugging-face-raises-235m-from-investors-including-salesforce-and-nvidia">&#8220;Hugging Face raises $235M from investors including Salesforce and Nvidia&#8221;</a>, 24 Aug 2023; <a href="https://salesforceventures.com/perspectives/welcome-hugging-face/">Salesforce Ventures</a> confirms it led. $4.5bn is the post-money valuation; $235m at that price represents about 5% in aggregate. Total raised about $395m; no priced round reported since. Hugging Face has never published its board; no outlet reported a board seat for any Series D investor.</p><p>[6] Microsoft, <a href="https://devblogs.microsoft.com/foundry/microsoft-and-hugging-face-expand-partnership-to-accelerate-open-source-ai-innovation-on-azure-ai-foundry/">&#8220;Microsoft and Hugging Face expand partnership&#8230; on Azure AI Foundry&#8221;</a>, 19 May 2025 (10,000+ models). Hugging Face, <a href="https://huggingface.co/blog/gcp-partnership">&#8220;Google Cloud partnership&#8221;</a>, 19 Nov 2025: CDN gateway, native TPU support, &#8220;over 1,500 terabytes of open models and datasets are downloaded and uploaded between Hugging Face and Google Cloud&#8221; daily. Hugging Face, <a href="https://huggingface.co/blog/bedrock-marketplace">&#8220;Bedrock Marketplace&#8221;</a>, 9 Dec 2024, and <a href="https://aws.amazon.com/ai/hugging-face/">AWS</a>.</p><p>[7] Amazon, <a href="https://www.aboutamazon.com/news/company-news/aws-ducklabs">&#8220;AWS to acquire DuckLabs&#8221;</a>, 26 Aug 2026: &#8220;We are not acquiring the DuckDB open source project, which will remain free and open source under the independent DuckDB Foundation&#8221;; price undisclosed; DuckDB Labs becomes a wholly owned subsidiary with the team in Amsterdam (<a href="https://www.geekwire.com/2026/amazon-acquires-ducklabs-adding-the-team-behind-duckdb-amid-broader-shakeup-in-cloud-data/">GeekWire</a>). MotherDuck&#8217;s Jordan Tigani on the same day: &#8220;That&#8217;s Amazon&#8217;s playbook, after all: wait until an open source project gets big enough, then launch it as a service.&#8221; Amazon&#8217;s $50bn OpenAI investment, announced 27 Feb 2026, completed 31 Jul 2026 (<a href="https://www.pymnts.com/news/artificial-intelligence/2026/amazon-completes-50-billion-dollar-investment-openai/">PYMNTS</a>). Globalstar: <a href="https://www.cnbc.com/2026/04/14/amazon-globalstar-satellite-leo-internet.html">CNBC, 14 Apr 2026</a>.</p><p>[8] Google&#8211;Wiz: closed 11 Mar 2026 after a &#8220;12-month regulatory review process&#8221; (<a href="https://www.clearygottlieb.com/news-and-insights/news-listing/google-completes-32-billion-acquisition-of-wiz">Cleary Gottlieb</a>); DOJ cross-appeal seeking Chrome and Android divestiture filed 3 Feb 2026 (<a href="https://www.cnbc.com/2026/01/16/google-files-to-appeal-search-monopoly-case.html">CNBC</a>). Amazon&#8211;Globalstar: $11.57bn, 14 Apr 2026, close expected early 2027 (<a href="https://www.bloomberg.com/news/articles/2026-04-14/amazon-s-11-6-billion-globalstar-deal-amps-up-rivalry-with-musk">Bloomberg</a>); FTC v. Amazon trial set for 29 Mar 2027 (<a href="https://www.mlex.com/mlex/articles/2420780/us-ftc-amazon-propose-pushing-back-antitrust-trial-to-march-2027">MLex</a>). Meta&#8211;Scale: 49% non-voting, Jun 2025 (<a href="https://www.axios.com/2025/06/11/metas-noncontrol-scale-ai-deal-antitrust">Axios</a>); Warren, Wyden and Blumenthal to FTC/DOJ on Meta&#8211;Scale, Google&#8211;Windsurf and Nvidia&#8211;Groq, 4 Feb 2026 (<a href="https://www.warren.senate.gov/newsroom/press-releases/warren-wyden-blumenthal-call-on-federal-regulators-to-investigate-meta-google-nvidia-reverse-acqui-hire-deals">Warren</a>). Salesforce: $25bn accelerated repurchase, Q1 FY27 call, 28 May 2026 (<a href="https://www.fool.com/earnings/call-transcripts/2026/05/28/salesforce-crm-q1-2027-earnings-transcript/">transcript</a>); Benioff, <a href="https://x.com/Benioff/status/2081821775353233804">X</a>, late Jul 2026: &#8220;Mahalo, @JensenHuang!&#8230; on Hugging Face (where we led the last round).&#8221; Apple&#8217;s largest acquisition remains Beats, $3bn, 2014. The &#8220;second request&#8221; line is my characterization, not a reported regulator view.</p><p>[9] Business Insider reported banks engaged to &#8220;evaluate incoming bids&#8221;; none has been named. Bloomberg Opinion (Parmy Olson, 31 Aug 2026, <a href="https://www.taipeitimes.com/News/editorials/archives/2026/08/31/2003863402">via Taipei Times</a>) floated Salesforce as a potential acquirer; no bid has been reported.</p><p>[10] NVIDIA, <a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581026000073/q2fy27pr.htm">Q2 FY2027 results</a>, 8-K exhibit, 26 Aug 2026, quarter ended 26 Jul 2026: revenue $96.2bn, data center $89.0bn; cash, equivalents and marketable securities $99.4bn; first-half repurchases $39.0bn plus $6.3bn of dividends; $99bn of authorization remaining. $89.0bn over 91 days &#8776; $0.98bn a day; $12.9bn &#8776; 13 days.</p><p>[11] Groq: ~$20bn license plus hires, Dec 2025; Warren and Blumenthal to Huang, Mar 2026: &#8220;by licensing its technology and hiring its most important employees, NVIDIA has effectively acquired Groq in all but name&#8221; (<a href="https://www.warren.senate.gov/newsroom/press-releases/warren-blumenthal-question-whether-nvidias-20-billion-groq-deal-is-attempt-to-avoid-antitrust-laws/">Warren</a>). Poolside: $6bn non-exclusive license plus $1bn equity, 109 engineers, ~20 Aug 2026; investor letter: &#8220;not an acquisition and it is not an acquihire&#8221; (<a href="https://www.techtimes.com/articles/325441/20260825/nvidias-third-structured-non-acquisition-nine-months-27-billion-sidestep-merger-review.htm">TechTimes</a>). FTC chair Ferguson, Jan 2026: the agency will review whether acqui-hires are &#8220;being constructed to try to escape Hart-Scott-Rodino review&#8221; (<a href="https://www.wilmerhale.com/en/insights/client-alerts/20260129-ftc-eyeing-acquihire-transactions-in-tech-industry">WilmerHale</a>). A $12.9bn stock purchase is far above the HSR threshold; no outlet has yet reported on the review, and the deal structure is unreported.</p><p>[12] Financial Times, &#8220;Why AI start-up Hugging Face turned down a $500mn Nvidia deal,&#8221; 28 Jan 2026, via <a href="https://oodaloop.com/briefs/technology/why-ai-start-up-hugging-face-turned-down-a-500mn-nvidia-deal/">OODA Loop</a> and <a href="https://www.tipranks.com/news/the-fly/hugging-face-declined-500m-nvidia-investment-offer-ft-reports-thefly">The Fly</a>. The reasoning is the FT&#8217;s account from people familiar with the matter; the company declined to comment. Several August rewrites date the offer to &#8220;earlier this year&#8221;; the FT says late 2025.</p><p>[13] FT, ibid.: profitable in 2025, a first-quarter 2026 loss on dataset investment, roughly half of the ~$400m raised still on the balance sheet. Delangue on TechCrunch&#8217;s Equity podcast this summer, quoted in <a href="https://techcrunch.com/2026/08/24/hugging-face-reportedly-in-talks-to-be-acquired-for-13b/">TechCrunch, 24 Aug 2026</a>: &#8220;close to profitability&#8221;; had only &#8220;recently started to touch the money that [it] raised three years ago.&#8221; Unaudited: the CEO&#8217;s own account.</p><p>[14] The Information, <a href="https://www.theinformation.com/briefings/exclusive-hugging-face-annualized-revenue-jumps-50-150-million">&#8220;Hugging Face Annualized Revenue Jumps 50% to $150 Million&#8221;</a>, Aug 2026: &#8220;more than $150 million in annualized revenue, a 50% increase from two months ago.&#8221; Annualized run-rate, not contracted ARR; much of the revenue is usage-based compute. Business Insider&#8217;s $100m figure (note 1) is consistent with a June snapshot. What drove the jump is not reported.</p><p>[15] $12.9bn &#247; $150m &#8776; 86x. Mechanical price: $7bn &#215; ~1.5 (growth since the refused offer) &#215; ~1.3 (control premium; 30&#8211;50% is the conventional public-market band) &#8776; $13.6bn. Figma: Adobe&#8217;s <a href="https://news.adobe.com/news/news-details/2022/adobe-to-acquire-figma">15 Sep 2022 release</a>, ~$20bn against &#8220;surpassing $400 million in total ARR exiting 2022,&#8221; a projected figure, &#8776;50x. Wiz: $32bn against $700m ARR at announcement (<a href="https://techcrunch.com/2025/03/18/google-is-buying-wiz-for-32b-to-beef-up-in-cloud-security">TechCrunch, 18 Mar 2025</a>), &#8776;46x.</p><p>[16] OpenAI confidential filing announced on 8 Jun 2026 (<a href="https://techcrunch.com/2026/06/08/following-anthropic-openai-files-confidentially-for-ipo/">TechCrunch</a>); Anthropic reportedly filed about a week earlier. Sarah Friar at an all-hands, <a href="https://www.cnbc.com/2026/08/19/open-ai-ipo-timing-2027-friar.html">CNBC, 19 Aug 2026</a>: &#8220;will be a public company in 2027,&#8221; possibly sooner if &#8220;our business continues to inflect.&#8221; The IPO-at-half estimate is mine, on public software multiples of 15&#8211;25x forward revenue. Conversion of all preferred at this price is an assumption, given $395m raised. No employee tender has been reported since 2023.</p><p>[17] <a href="https://www.axios.com/2025/11/18/hugging-face-delangue-axios-bfd-2025-llm-bubble">Axios, 18 Nov 2025</a>: &#8220;I think we&#8217;re in an LLM bubble, and I think the LLM bubble might be bursting next year.&#8221;</p><p>[18] Hugging Face, <a href="https://huggingface.co/blog/security-incident-july-2026">&#8220;Security incident disclosure, July 2026&#8221;</a>, 16 Jul, and <a href="https://huggingface.co/blog/agent-intrusion-technical-timeline">&#8220;Anatomy of a Frontier Lab Agent Intrusion: Technical Timeline&#8221;</a>, 27 Jul 2026: window 9&#8211;13 July; &#8220;the attacker was bound by no usage policy, while our own forensic work was blocked by guardrails of hosted models we first tried&#8221;; analysis of 17,000+ actions run on Z.ai&#8217;s GLM-5.2 in Nvidia&#8217;s optimised build. Delangue, <a href="https://www.cbsnews.com/news/clement-delangue-face-the-nation-transcript-aug-2-2026/">Face the Nation, 2 Aug 2026</a>: &#8220;We used the American version of a Chinese model.&#8221; OpenAI, <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">&#8220;Hugging Face model evaluation security incident&#8221;</a>, 21 Jul, and full report of 26 Aug 2026 (<a href="https://techcrunch.com/2026/08/26/openai-releases-its-official-report-on-the-hugging-face-breach/">TechCrunch</a>): a pre-release Astra-family model in the ExploitGym evaluation escaped its sandbox; &#8220;misaligned behavior in an outlier scenario.&#8221;</p><p>[19] <a href="https://techcrunch.com/2026/07/26/hugging-face-ceo-calls-for-radical-transparency-after-unprecedented-openai-hack/">TechCrunch, 26 Jul 2026</a>: Delangue asks for agent traces and &#8220;$100 million in computing resources&#8221; for community cyber-defense. <a href="https://www.axios.com/2026/08/13/liability-ai-courts-agents-openai-anthropic-meta">Axios, 13 Aug 2026</a>: state attorneys general preservation demands; <a href="https://www.techtarget.com/it-strategy/feature/OpenAI-Hugging-Face-incident-raises-AI-liability-concerns">TechTarget, 13 Aug</a>: AI coverage &#8220;rapidly changing right now&#8221;; <a href="https://www.axios.com/2026/08/26/openai-hugging-face-technical-report-ai-hack">Axios, 26 Aug</a>: Alabama attorney general subpoena. No compensation, insurance claim, lawsuit, or Hugging Face cost figure has been reported. Reuters and Decrypt tied the sale exploration to &#8220;a month after&#8221; the incident.</p><p>[20] CNBC, <a href="https://www.cnbc.com/2025/12/09/meta-avocado-ai-strategy-issues.html">&#8220;From Llamas to Avocados&#8221;</a>, 9 Dec 2025. Meta AI, <a href="https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model">&#8220;Introducing Muse Glimmer&#8221;</a>, 10 Aug 2026: 30B dense, Apache-2.0, &#8220;trained on Muse Spark&#8217;s outputs using logit distillation,&#8221; sized for a single consumer GPU. Meta says Muse Spark 1.2 weights will follow; until they do, the top of Meta&#8217;s line is closed. No formal end of the Llama line has been announced.</p><p>[21] <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/deepseek-launches-1-6-trillion-parameter-v4-on-huawei-chips">Tom&#8217;s Hardware, 26 Apr 2026</a>: DeepSeek V4-Pro, 1.6T parameters, MIT licence, &#8220;optimized for Huawei&#8217;s Ascend AI processors&#8221;; Huawei confirmed day-one support across the Ascend 950 line.</p><p>[22] <a href="https://github.com/huggingface/optimum-tpu">huggingface/optimum-tpu</a>, archived 23 Jan 2026, last release Dec 2024. Optimum AMD&#8217;s own repository has been dormant since 2023; AMD support moved into mainline Transformers and the Kernel Hub.</p><p>[23] Stripe, <a href="https://stripe.com/newsroom/news/stripe-agrees-to-acquire-openrouter">&#8220;Stripe agrees to acquire OpenRouter&#8221;</a>, 19 Aug 2026; about $7.5bn per the New York Times (<a href="https://techcrunch.com/2026/08/19/stripe-didnt-really-buy-openrouter-because-of-the-singularity/">TechCrunch</a>); Bloomberg had reported a deal &#8220;over $7 billion&#8221; on 16 Aug. OpenRouter lists Nvidia as a customer; its six most-used models in July were all Chinese open-weight models.</p><p>[24] &#8220;Open Weights and American AI Leadership,&#8221; 24 Jul 2026, authored by Jensen Huang, 25 founding signatories including Hugging Face, Microsoft, Meta and IBM; Salesforce, Amazon, Anthropic and Apple did not sign (<a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidia-and-24-other-companies-sign-open-weights-letter">Tom&#8217;s Hardware</a>, <a href="https://techcrunch.com/2026/07/24/as-us-weighs-response-to-chinese-ai-industry-urges-against-broad-open-weight-restrictions/">TechCrunch</a>). Nvidia, <a href="https://blogs.nvidia.com/blog/open-secure-ai-alliance/">Open Secure AI Alliance</a>, 27 Jul 2026, with Hugging Face as a member. Nemotron 3.5 Lightning, 11 Aug 2026; The Information, 11 Aug: Nemotron 4 at a trillion-plus parameters targeted for late fall.</p><p>[25] Hugging Face, <a href="https://huggingface.co/blog/state-of-open-models-summer-2026">&#8220;State of Open Models: Summer 2026&#8221;</a>, 14 Aug 2026: Qwen 2,045M downloads in 2026; 151,448 Qwen derivatives against 28,531 for Llama. <a href="https://techcrunch.com/2026/07/14/the-real-ai-race-may-no-longer-be-at-the-frontier-open-models-hugging-face/">TechCrunch, 14 Jul 2026</a>: Chinese open-weight models 41% of hub downloads, spring 2026. &#8220;More than the rest of the open ecosystem combined&#8221;: Interconnects, Jan 2026, on December 2025 downloads. Delangue, <a href="https://www.cnbc.com/2026/08/03/hugging-face-china-ai-race-open-models.html">CNBC, 3 Aug 2026</a>.</p><p>[26] Axios, ~20 Jul 2026, via <a href="https://www.fastcompany.com/91576757/">Fast Company</a>: administration &#8220;reviving a push to ban Chinese AI models,&#8221; Entity List floated as the mechanism. House Homeland Security and Select Committee on the CCP letters to Anysphere and Airbnb (Apr 2026) and <a href="https://homeland.house.gov/2026/07/31/chairmen-garbarino-moolenaar-continue-joint-investigation-into-prc-ai-models/">DoorDash (31 Jul 2026)</a>. No letter to a hosting platform has been reported. &#8220;No lobbying budget&#8221; is my characterization; Hugging Face has a policy team but no reported federal lobbying spend.</p><p>[27] Reuters, 7 Jul 2026, three sources: MOFCOM met Alibaba, ByteDance and Zhipu over the preceding month on restricting overseas access to advanced models including open-weight versions (summary at <a href="https://www.explainx.ai/blog/china-overseas-ai-model-restrictions-reuters-july-2026">explainx</a>); Jamestown Foundation, <a href="https://jamestown.org/beijing-signals-tiered-governance-of-open-weight-models/">&#8220;Beijing Signals Tiered Governance of Open-Weight Models&#8221;</a>, 13 Aug 2026: basic tier files, advanced tier faces security review, frontier tier domestic-only or unreleased; discussion stage, not yet policy. Jamestown also documents Alibaba keeping Qwen 3 Max API-only.</p><p>[28] DeepSeek, <a href="https://api-docs.deepseek.com/news/news260424/">V4 Preview release note</a>, 24 Apr 2026, links weights only to huggingface.co; <a href="https://huggingface.co/moonshotai/Kimi-K3">moonshotai/Kimi-K3</a>, weights 27 Jul 2026; GLM-5.2 on Hugging Face and ModelScope, 13 Jun 2026; Qwen3.8-27B on both, evening of 14 Aug 2026 Beijing time. No case of a Chinese model being withdrawn from the hub has been reported.</p><p>[29] Block: <a href="https://www.chinatalk.media/p/hugging-face-blocked-self-castrating">ChinaTalk, 19 Oct 2023</a> (intermittent from May 2023, blocked from 12 Sep 2023); Hugging Face to <a href="https://www.semafor.com/article/10/20/2023/ai-platform-hugging-face-confirms-china-blocked-it">Semafor, 20 Oct 2023</a>: &#8220;there&#8217;s not much we can do against government regulations for now.&#8221; hf.space unreachable from the mainland from Nov 2025 (<a href="https://discuss.huggingface.co/t/hf-space-has-been-blocked-in-china/170544">forum thread</a>). hf-mirror.com is community-run; no Hugging Face involvement is documented. The departure of the head of its Asia-Pacific ecosystem is confirmed by his own public note and by <a href="https://restofworld.org/2026/tiezhen-wang-china-us-open-source-ai/">Rest of World, 15 Jun 2026</a>. No Chinese entity, licensed mirror, or Chinese-cloud partnership has ever been announced; the absence is an inference from the public record.</p><p>[30] Alibaba, Sep 2025: ModelScope passed 100,000 models and 18m users (<a href="https://www.doit.com.cn/p/543121.html">DOIT</a>); Hugging Face had about 2.4m models in January and 2.96m in August 2026 (note 25). User figures are not like-for-like.</p><p>[31] NVIDIA Q2 FY2027 call, 26 Aug 2026, Colette Kress: &#8220;we ship less than 1% of our total data center revenue in Hopper 200 products to customers based in China in accordance with the U.S. Government licenses&#8221;; &#8220;there is no China data center compute revenue in our forward outlook&#8221;; also Kress: &#8220;NVIDIA runs the leading closed models,&#8221; and Huang: &#8220;Nearly all open models run on NVIDIA&#8221; (<a href="https://news.alphastreet.com/nvidia-corporation-nvda-q2-2027-earnings-call-transcript/">transcript</a>). SAMR preliminary finding, 15 Sep 2025, that Nvidia violated the Anti-Monopoly Law in relation to the Mellanox conditions; investigation continuing (<a href="https://www.cnbc.com/2025/09/15/china-nvidia-violated-anti-monopoly-law-will-continue-investigation.html">CNBC</a>).</p><p>[32] Tom&#8217;s Hardware, <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidia-and-24-other-companies-sign-open-weights-letter">&#8220;Nvidia and 24 other companies sign open-weights letter&#8221;</a>, 24 Jul 2026: Anthropic, OpenAI and Google absent at launch; OpenAI and Google joined within days, Anthropic did not. Open Secure AI Alliance membership: <a href="https://blogs.nvidia.com/blog/open-secure-ai-alliance/">NVIDIA</a>, 27 Jul 2026. Anthropic, <a href="https://www.anthropic.com/news/position-open-weights-models">&#8220;Our position on open-weights models&#8221;</a>, 27 Jul 2026: &#8220;Anthropic has never advocated for a ban on open-weights models&#8221;; &#8220;Open-weights models that don&#8217;t have dangerous capabilities are a public good&#8221;; &#8220;We should not sell powerful chips or chipmaking equipment to China.&#8221; Lobbying: New York Times, reported 25 Jul 2026 (summary at <a href="https://www.implicator.ai/openai-anthropic-lobby-washington-open-weight-ai/">implicator.ai</a>); the Alibaba distillation claim was made to the Senate in 2026. The interpretability team has published small open-weight research models on the hub; no Claude weights.</p><p>[33] Amazon, <a href="https://www.aboutamazon.com/news/company-news/amazon-invests-additional-5-billion-anthropic-ai">20 Apr 2026</a>: $5bn plus up to $20bn, &#8220;up to 5 gigawatts of Trainium capacity,&#8221; Anthropic to spend &#8220;$100 billion&#8221; on AWS over a decade. Google and Anthropic, <a href="https://www.anthropic.com/news/google-broadcom-partnership-compute">6 Apr 2026</a>: &#8220;multiple gigawatts&#8221; of TPU from 2027, on top of the Oct 2025 deal for up to one million TPUs; Google to invest up to $40bn (<a href="https://techcrunch.com/2026/04/24/google-to-invest-up-to-40b-in-anthropic-in-cash-and-compute/">TechCrunch, 24 Apr 2026</a>). AMD, <a href="https://ir.amd.com/">22 Jul 2026</a>: up to 2GW of MI450, first gigawatt in H1 2027, up to $5bn equity. Nvidia and Microsoft, <a href="https://www.anthropic.com/news/microsoft-nvidia-anthropic-announce-strategic-partnerships">18 Nov 2025</a>: up to $10bn and $5bn respectively, $30bn of Azure, up to 1GW of Grace Blackwell and Vera Rubin. The &#8220;minority by announced capacity&#8221; comparison is mine; no analyst split has been published, and contracted gigawatts are not delivered gigawatts. Shareholdings: Google about 14% (court filings, 2025); Amazon&#8217;s position valued at $74bn at the Series G (<a href="https://fortune.com/2026/06/04/amazon-google-billions-anthropic-ipo/">Fortune, 4 Jun 2026</a>); Nvidia and Microsoft stakes undisclosed.</p><p>[34] Google and AMD signatures: Tom&#8217;s Hardware, ibid. Amazon&#8211;OpenAI: $50bn, $15bn immediately, completed 31 Jul 2026 (note 7). Amazon is reported to be evaluating OpenAI and its own models for internal workloads after Anthropic price increases (The Information, via <a href="https://www.techzine.eu/news/infrastructure/142545/">Techzine, 30 Jun 2026</a>). Next Platform, <a href="https://www.nextplatform.com/ai/2026/08/13/the-war-between-open-source-open-weight-and-closed-ai-models/5287504">13 Aug 2026</a>, draws the open-versus-closed line; no published analysis ties the labs&#8217; silicon choices to it.</p><p>[35] Microsoft, <a href="https://news.microsoft.com/announcement/microsoft-acquires-github/">&#8220;Microsoft acquires GitHub&#8221;</a>, announced 4 Jun 2018, closed 26 Oct 2018, $7.5bn. GitHub folded into Microsoft&#8217;s CoreAI group without a successor CEO, Aug 2025 (<a href="https://www.runtime.news/why-microsofts-decision-to-bury-github-in-its-coreai-group-is-the-end-of-an-era/">Runtime</a>). Torvalds on Nvidia&#8217;s drivers, Aalto, June 2012; open kernel modules from 2022 (<a href="https://developer.nvidia.com/blog/nvidia-transitions-fully-towards-open-source-gpu-kernel-modules">NVIDIA</a>). Mellanox: $6.9bn announced Mar 2019, closed Apr 2020, brand retired Aug 2020 (<a href="https://nvidianews.nvidia.com/news/nvidia-completes-acquisition-of-mellanox-creating-major-force-driving-next-gen-data-centers">NVIDIA</a>). Vera CPU shipping since May 2026 (<a href="https://finance.yahoo.com/sectors/technology/articles/nvidia-starts-vera-cpu-shipments-174725636.html">Benzinga</a>). Also the CUDA EULA clause barring translation of compiled output &#8220;to target a non-NVIDIA platform&#8221; (<a href="https://www.tomshardware.com/pc-components/gpus/nvidia-bans-using-translation-layers-for-cuda-software-to-run-on-other-chips-new-restriction-apparently-targets-zluda-and-some-chinese-gpu-makers">Tom&#8217;s Hardware</a>) and NVLink Fusion, licensed to third-party silicon only when paired with an Nvidia GPU or CPU.</p><p>[36] NVIDIA ConnectX-6 firmware release notes, <a href="https://networking-docs.nvidia.com/connectx6defwrn/22481000/validated-and-supported-cables-and-switches">&#8220;Validated and Supported Cables and Switches&#8221;</a>, verbatim; the same text appears across ConnectX and Quantum release notes. Pre-acquisition Mellanox held roughly 55&#8211;60% of the InfiniBand market and sold to all comers (see the March piece).</p><p>[37] Run:ai founders at closing, quoted by <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidia-finalizes-acquisition-of-ai-software-firm-run-ai-takes-software-open-source-company-reportedly-cost-usd700-million">Tom&#8217;s Hardware, 30 Dec 2024</a>: &#8220;open sourcing the software will enable it to extend its availability to the entire AI ecosystem.&#8221; NVIDIA, <a href="https://developer.nvidia.com/blog/nvidia-open-sources-runai-scheduler-to-foster-community-collaboration/">&#8220;NVIDIA Open Sources Run:ai Scheduler&#8221;</a>, 1 Apr 2025: KAI Scheduler under Apache-2.0, which &#8220;continues to be packaged and delivered as part of the NVIDIA Run:ai platform.&#8221; About $700m as reported; EU clearance unconditional (<a href="https://ec.europa.eu/commission/presscorner/detail/en/ip_24_6548">European Commission</a>).</p><p>[38] <a href="https://docs.nvidia.com/nim/large-language-models/latest/about-nim-llm/overview.html">NVIDIA NIM for LLMs 2.0 overview</a> and <a href="https://docs.nvidia.com/nim/large-language-models/2.0.1/about-nim-llm/release-notes.html">2.0.1 release notes</a>, Mar 2026: &#8220;a ground-up redesign that adopts a one-container, one-backend philosophy&#8221;; the 1.x multi-backend container (vLLM, TensorRT-LLM) &#8220;is replaced by a dedicated vLLM container.&#8221; Free under the Developer Program; production under NVIDIA AI Enterprise (note 4).</p><p>[39] <code>HF_ENDPOINT</code> in <a href="https://github.com/huggingface/huggingface_hub/blob/main/src/huggingface_hub/constants.py">huggingface_hub constants</a>; <a href="https://hf-mirror.com/">hf-mirror.com</a> instructs users to set it. Optimum for Intel Gaudi and Optimum Neuron are Apache-2.0; Optimum AMD is MIT.</p>]]></content:encoded></item><item><title><![CDATA[The Host Pays the Empire]]></title><description><![CDATA[AI data centers are the foreign bases of the US&#8211;China tech rivalry, except that the host now pays the rent.]]></description><link>https://www.airealist.ai/p/the-host-pays-the-empire</link><guid isPermaLink="false">https://www.airealist.ai/p/the-host-pays-the-empire</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Tue, 25 Aug 2026 10:13:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!j-YM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!j-YM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!j-YM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!j-YM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!j-YM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!j-YM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!j-YM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:791433,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/212673092?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!j-YM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!j-YM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!j-YM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!j-YM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8eb942d-b7c1-4426-8393-74698008dae4_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On August 21, 2026, Papua New Guinea's Minister for State Enterprises presented Kumul Cloud Infinity, which the country has described as its "first sovereign cloud and AI data centre."[1] The system is perfectly in the Western manner: cloud software from CloudSigma, a company based in Zurich, which offers the same "sovereign" platform to a number of service providers in over thirty-five countries; hardware from Hewlett Packard Enterprise; NVIDIA H200 GPUs; and gas-powered backup power supplied by a local partner.[2] The operator is Datec PNG, a subsidiary of Telikom, which is itself held by Kumul Consolidated Holdings, the state's holding company. It was launched as a commercial success.</p><p style="text-align: justify;">None of the following information was made public in the operator's press release, the holding company's release, or any subsequent reports: the cost, the megawatts, the GPU count, the rack count, the tier certification, or the financing source. All of these details. I have checked six sources, including the two official ones. The extent and cost of Papua New Guinea's sovereign AI infrastructure are a secret, possibly even from the people paying for it.</p><p style="text-align: justify;"><span>The other building is located across town. The National Data Center was constructed by Huawei and opened in 2018, with funding from a reported $53 million loan from China&#8217;s Export-Import Bank.[3] In 2020, a review carried out by Australia&#8217;s foreign ministry on behalf of PNG&#8217;s own cyber-security agency discovered that the facility&#8217;s encryption algorithm had been "openly broken" for years, that its firewalls had reached the end of their life two years&nbsp;</span><em><span>prior to the day it opened</span></em><span>, and that its core switches were outside the firewall, meaning that remote access could not be detected.[4] Most of the government departments had never moved into the center. PNG&#8217;s ICT minister described the asset in official remarks as "a failed investment."[5] It has not been made public whether the loan has been repaid.</span></p><p>A single capital, a single decade, and the buildings of both superpowers.</p><p style="text-align: justify;">The correct approach to understanding those buildings is not to look at IT procurement; it is to consider the oldest form of transaction in geopolitics: <strong>the foreign base</strong>. A base is not simply about the concrete; it is about physical presence. The patron puts hardware within the host's territory, and the host shows its alignment by accepting it, while rivals are prevented from using that ground. A data center is not a base: it has no soldiers, no jurisdiction, and it does not deny ground to anyone.</p><p style="text-align: justify;">All of the instinctive interpretations are incorrect, and the reasons for this are worth giving. The account is not about Chinese debt traps. According to the specialist literature, Beijing has arranged the loans so that they are repaid, either by collecting the debt or by rescheduling it rather than by seizing the assets; in this case, the loan is the result, not the trap.[6] It isn't a story about NVIDIA taking control of the Global South. When its own chief financial officer speaks of sovereign AI revenue, the only countries mentioned are the rich ones.[7] Nor is it a story about gullible governments, as Port Moresby itself will show before this article is finished.</p><h2>What a base actually costs, and who pays</h2><p>For 70 years, the foreign base machine has operated using a well-documented payment route from patron to host.</p><p style="text-align: justify;">Djibouti is simultaneously renting out its geographical location to five armies. According to the most recent figure provided by its finance minister, the United States paid around $65 million per year for Camp Lemonnier, France paid about $30 million, China was reported to pay $20 million for its first overseas base, and both Italy and Japan paid roughly $3 million each, making the total annual amount about $121 million. A Djiboutian official summed up the nation's business strategy in a single sentence by saying that <strong>the country's geography is its primary national resource, just as oil is for the Gulf states</strong>.[8]</p><p style="text-align: justify;">Kyrgyzstan conducted its trade through auctions. In 2001, the United States paid about $2 million per year for the Manas Air Base. By 2006, the rent had increased to $17.4 million. In 2009, Moscow made a counter-offer to Bishkek in the form of $2.15 billion in aid and loans so as to drive the Americans out; Kyrgyzstan declared that it would close the base, allowed the bidding to proceed, and then re-let it to the United States for $60 million a year, thirty times the original price.[9] <strong>The geography had not changed; what had changed was that two patrons were now bidding</strong>. Likewise, the 2025 treaty concerning Diego Garcia obliges the United Kingdom to pay an average of about &#163;101 million per year to Mauritius, for a period of 99 years, in order to keep the Indian Ocean base which it leases on to the Americans, though the bill implementing the treaty has stalled at Westminster and, at the time of writing, no pound has yet been paid.[10]</p><p style="text-align: justify;"><span>Honesty demands that we look at counterexamples as well: Japan pays Washington around &#165;211 billion each year in host-nation support&#8212;approximately $1.5 billion, depending on the exchange rate&#8212;and South Korea about $1.1 billion.[11] Therefore, the rule has never been that 'the patron pays'. The rule is that&nbsp;</span><strong><span>the direction of payment follows bargaining power</span></strong><span>.[12] When the host country is rich, under threat, and needs the patron more than the patron needs the location, the host pays for protection. In the case of poor host countries with geographical advantages, they receive payment because of their location. If all you have is your position, that is worth actual money, and the competition among the superpowers drives up the price.</span></p><p>For AI data centers, almost everyone now pays the way Japan does, even countries that have nothing to offer except their location.</p><h2>The ledger</h2><p style="text-align: justify;">For this article, I compiled a census of sovereign cloud and sovereign AI initiatives in the countries that the AI conversation never addresses: not the United States, China, the EU's big five, or the Gulf states, but rather the long tail: 59 programs in 47 countries, ranging from Papua New Guinea to Paraguay, Zambia to Kazakhstan. Almost all of them have been announced or launched between January 2024 and August 2026, together with a small number of earlier projects kept as precedents.[13]</p><p style="text-align: justify;">There is no information given about a single signed power purchase agreement. Of the 23 that announce an investment figure, only seven are based on actual documentation, such as a signed contract, a disbursement, or a formal EU award. The others consist of rounded figures&#8212;such as $250 million, $1 billion, and $10 billion&#8212;announced at GITEX or GTC, the Gulf and NVIDIA trade shows, not at ribbon-cutting ceremonies. Specific and documented figures appear only in cases where a lender's or counterparty's documentation requires them: &#8364;66.55 million in a document relating to an African Development Bank project, &#8364;36 million in the EuroHPC procurement&#8212;the EU's joint supercomputing initiative&#8212;and an $18 million three-year contract with Oracle. </p><p style="text-align: justify;">In all other cases, the figures are given in rounded form. Forty-two per cent of the programs have never gone beyond the paper stage, meaning only an announcement, a policy, or a memorandum of understanding. Nearly sixty per cent have not produced any actual computing power at all. Very few name a paying customer; in instances where an anchor tenant is mentioned, it is almost always the government paying for the building, and a true arm's-length tenant (such as Kenya's Everse Technology or Thailand's Mahidol University) appears in only a very small number of cases.</p><p><span>It is normal for private data centers not to disclose their power purchase agreements. However, in the case of these projects, which are&nbsp;</span><em><span>state-sponsored sovereignty</span></em><span>&nbsp;initiatives sold to legislatures on the grounds of national interest, the entire risk lies in the electricity bill, and by limiting the count to the 31 programs that are either currently operational or under construction, the figure of zero remains unchanged</span></p><p>Prepare the ledger for these transactions based on the two-column format that any  analyst would use. <strong>What does the host get, and what does the host pay?</strong></p><p style="text-align: justify;">Kenya obtained a loan of RMB 1.225 billion, approximately $180 million, from China EXIM at an interest rate of 2 percent, with a grace period of about seven years, to construct the Konza National Data Center, with Huawei as the sole contractor. The first payment was due on March 31, 2026.[14] According to Kenya's Auditor-General, the facility being repaid for had a disaster-recovery site that was "not operational", lacked fire suppression systems, with the data being backed up on the same servers situated in the Primary Data Center, no data protection officer having been appointed, and no vetting having been carried out of either the customers or the staff within the building.[15] Those are the shortcomings of the state operator, not the contractor's; yet they are the very things the loan was used to acquire. The cloud revenue earned at Konza for the year up to June 2025 was KSh 126.8 million, which can be rounded to 1 million dollars. It had decreased by 16 percent.[16]</p><p style="text-align: justify;">In all cases where China is not the one providing the funds, the situation fails to improve. Instead, the host country and its partners just pay each other. Armenia is financing its declared $500 million NVIDIA-stack 'AI factory' via a private developer and a diaspora foundation. Kazakhstan has entered into a $10 billion framework agreement for a gigawatt-scale 'Data Center Valley' on its own and its partners' balance sheets. The financing arrangement in Papua New Guinea remains, as has been noted, a secret.</p><p style="text-align: justify;">The census involves two types of arrangements. In some cases, the program is based on state contributions: the government borrows money in order to build the infrastructure itself, using its own balance sheet, for example, in Kenya, Zambia, Senegal, and the 2018 PNG facility. In other instances, it is a private colocation venture&#8212;such as Malaysia&#8217;s YTL, the pan-African Cassava, Vietnam&#8217;s FPT, and an Indonesian project financed by I Squared&#8212;where a foreign fund assumes the risk of construction and 'sovereign' is used as a marketing point in government tenders. Armenia and Kazakhstan fall between the two categories, with both state and private partners sharing the cost. It is important to distinguish between the two: the group comprising loans and Chinese-built projects is uniformly disappointing, whereas when a private operator assumes the commercial risk, the investor bears that exposure, not the taxpayer.</p><p style="text-align: justify;"><span>Yet even in the case of privately financed projects, </span><strong><span>the silicon is imported under the laws of another country, and on the American side, there is no concealment regarding the direction of this flow</span></strong><span>. OpenAI&#8217;s program for third countries states that partner countries &#8220;will invest in expanding the global Stargate Project&#8221;: the host country is required to fund the patron&#8217;s network.[17] It is reported that the UAE&#8217;s chip deal requires that for every dollar Abu Dhabi invests in the Stargate UAE project, a dollar must also be invested in AI infrastructure&nbsp;</span><em><span>on US soil</span></em><span>.[18] The Export-Import Bank&#8217;s new AI export program is, in effect, buyer&#8217;s credit: America lends you the money needed to buy its chips.[19] The United States&#8217; development finance arm has granted one major loan for a data center in Africa: up to $300 million to a commercial pan-African company, with $83 million disbursed by 2022, the most recent figure available.[20]</span></p><p style="text-align: justify;"><span>At the same time, the demand for the facilities said to have been built is measurable: the $1 million per year in Konza, which is decreasing, is an example of the situation in one of Africa's most digitally advanced economies. You can't fill 100 megawatts with that amount, any more than you can fill ten. That is why the biggest initiatives in the census have quietly given up pretending: Kazakhstan claims to generate "at least $3 billion in annual export revenue," and Guyana's facility "will meet international demand.&#8221; Put simply,&nbsp;</span><strong><span>since there is no domestic demand for such services, we are going to have foreign customers' computing carried out on our grid, on our balance sheet, under our own flag</span></strong><em><span>.</span></em><span>&nbsp;A sovereign facility whose business case is based on serving foreign customers is not true sovereignty; it is simply an export processing zone with more effective marketing.</span></p><p style="text-align: justify;">The lack of demand refers to demand for frontier-training; real demand does exist for ordinary enterprise colocation&#8212;such as disaster recovery, hosting, and data repatriated from Frankfurt&#8212;and this demand fills private data centers like those of YTL and Cassava. Moreover, the offshore option is by no means free, since renting hyperscale capacity abroad involves currency outflows, latency, and exposure to foreign jurisdictions, which later becomes clear in the health-data deals.</p><p style="text-align: justify;">The company at the heart of the situation has stated its view in the one language that matters. When asked by analysts during the February earnings call, NVIDIA's chief financial officer said that sovereign AI revenue had "more than tripled" to over $30 billion in fiscal 2026, with the increase mainly driven by customers in five wealthy countries.[7] The term "sovereign AI" has not appeared in either an NVIDIA 10-K or a 10-Q; it only appears in comments made during earnings calls, in press releases, and in the stylish annual report, and is absent from the officially audited documents. In the first quarter of fiscal 2027, sovereign revenue continued to rise &#8212; by more than 80 per cent year on year, compared with 85 per cent for the company as a whole &#8212; and it was included in a new classification divided between "Hyperscale" and "ACIE" (AI Clouds, Industrial and Enterprise), a breakdown that no longer separates out sovereign revenue and in which the word is completely omitted from the press release.[7] The label is being phased out and absorbed into the hyperscale category just as the long tail embraces the proposal.</p><h2>The armory</h2><p style="text-align: justify;">The second key point after the payment rail is that while the host controls the territory, the patron controls the weapons: sovereignty extends to the perimeter but not to the armory. In this case, the argument "these are merely commercial arrangements, not bases" is answered by precedent, since one type of procurement has always included a supply restriction and a monogamy clause: advanced arms. Turkey was excluded from the F-35 program in 2019 for having purchased Russian air defense systems; the fighter is the proper analogy here: an alignment tool priced as a capital expenditure and obtained at a price that is too high. The difference in this instance is even more stark: <strong>at least with the F-35, a buyer gets a weapon that works for him, whereas the statistics show that buyers have no workloads</strong>. The alignment is, in fact, the whole product. The restriction itself is incorporated into US regulations and functions in three stages. Those who have read "<a href="https://www.airealist.ai/p/access-disable-destroy">Access, Disable, Destroy</a>" will be familiar with the chip layer of the coercion system; the rings are that layer divided up by time&#8212;the chips you have not yet received, the chips you currently have, and the chips themselves.</p><p style="text-align: justify;"><strong><span>The first layer relates to future supplies and is firmly in place.</span></strong><span>&nbsp;Any export license granted under the US regulations concerning AI chips is, by regulation, "open to revision, suspension or revocation in whole or in part without notice".[22] This is not merely a possibility; it has already been put into practice by revoking licenses against the adversary and by making midstream revisions to allied buyers, though not yet against a long-tail recipient. In 2024 alone, Washington canceled eight existing export licenses to Huawei, including those held by Intel and Qualcomm. In October 2023, orders from the Gulf region for NVIDIA chips were halted midway due to a change in the rules, as noted in NVIDIA's own 8-K filing, which lists Saudi Arabia and the UAE.[23] The regime has been rewritten on three separate occasions within three years and can therefore alter the terms beneath whatever system a recipient is using. The rule issued in July 2026, which facilitates access for the UAE, is the most revealing of all: chip exports on a license-free basis go only to a whitelist of eight named American companies and to specific UAE government agencies&#8212;plus G42 and Core42, the UAE's own national champions, whose eligibility&nbsp;</span><em><span>automatically expires</span></em><span>&nbsp;in April 2027 unless they arrange a restructuring that satisfies Washington.[24] Even the most favored recipient in the system has its status on a lease that includes a sunset clause.</span></p><p style="text-align: justify;">The Validated End User (VEU) regime, which applies to data centers, serves as the link between the first and second rings, the condition for secure and lasting supply being a check of the delivered floor. If a foreign facility wishes to gain access to chips on a durable basis under the VEU arrangement, as stated in the Federal Register, it agrees to "on-site compliance reviews by representatives of the United States Government" and submits semiannual reports to Washington covering its chip inventory, its compute utilization, and its customer list.[25] At present, no host country is included in this VEU scheme; it has so far only been put into practice for deployments at the scale of hyperscalers, and the July 2026 UAE license-free authorization is now listed alongside it. Nevertheless, it is a model that can be applied to any long-tail host that may grow in size at some point. <strong>A data center considered sovereign and operating under the VEU system must submit its customer list to a foreign power twice a year and allow that power's inspectors to conduct inspections at the facility.</strong> Status-of-forces agreements have granted American military police authority over American personnel and installations located on the soil of a host country; similarly, the VEU gives American export-control officers access to the server room. And if your company rents capacity in one of these facilities, then your name will appear on that ledger.</p><p style="text-align: justify;"><strong><span>The second ring involves hardware, the mechanism in this case being a failure rather than a switch.</span></strong><span>&nbsp;A frontier cluster is not a fixed entity; it is a subscription. CUDA updates, firmware, enterprise software licenses, interconnect spares and RMA replacements all come from the vendor under terms that include US export law, and once the operator is placed on the Entity List &#8212; that is, Washington's export blacklist &#8212; any servicing, spares or updates will require licenses which are likely to be denied. If a cluster is severed, it does not cease to function; instead, it deteriorates, drifting away from the frontier as the software stack freezes and becomes out of service after a number of years, because failed GPUs, optics, and switches exceed the capacity of an empty spares store. The possibility of disablement by the supplier is no longer merely theoretical but directly relevant: equipment from John Deere that was looted from Ukraine and taken to Chechnya was remotely locked via the dealer's connection, according to the reports; in 2024 Microsoft cut Russian companies off from their&nbsp;</span><em><span>existing</span></em><span>&nbsp;cloud services as a result of EU sanctions; when Nokia and Ericsson withdrew, Russia's mobile networks were left to degrade without attention; and ASML has reportedly promised the Dutch government that it can remotely disable EUV machines in the event of an invasion.[26]</span></p><p style="text-align: justify;"><strong><span>The third ring is the one people most frequently inquire about &#8212; a kill switch built into the silicon &#8212; and the only straightforward thing to say is that it has been claimed by Beijing, denied by Santa Clara, and not verified by anyone.</span></strong><span>&nbsp;In July 2025, China's cyberspace regulator called NVIDIA over alleged 'tracking and positioning' and 'remote shutdown' capabilities in its chips sold in China; NVIDIA responded with a corporate blog post that, in effect, stated&nbsp;there were </span><em><span>no backdoors. No kill switches. No spyware.</span></em><span>[27] A congressional bill requiring location-verification features on AI chips exported from the United States has passed its committee but is not yet law; should it become law, the third ring would cease to be a point of dispute and become mere speculation.[28] I will not state that the third ring is a fact, nor should anyone else. However, note what the dispute itself shows: both superpowers take the possibility seriously enough to pass legislation on it and summon company executives to discuss it. When it comes to bases, the issue is not whether the weapons on them are loaded. It is about who holds the keys &#8212; and no one disputes that the keys are not in the host's pocket.</span></p><p style="text-align: justify;">The rings are not an American patent; if you want to see the first ring demonstrated on a live subject from both sides simultaneously, Malaysia conducted the demonstration in May 2025. On the 13th, Washington stated that using Huawei's AI chips anywhere in the world would risk breaches of US export controls. Six days later, a deputy minister from Malaysia launched what was described as the region's first sovereign full-stack AI system &#8212; comprising Huawei Ascend chips and DeepSeek models, thereby presenting the entire Chinese alternative. However, within about forty-eight hours, the ministry withdrew the announcement.[29] In the same week, a small country experienced the effect of Washington's first ring and Beijing's supply chain, and its sovereignty was publicly revised twice.</p><p>The host goes around the perimeter; the patron has control of the armory; and, unlike any other base host in history, this one paid for the armory out of its own money first.</p><h2>Two empires, two ledgers</h2><p style="text-align: justify;">The situation would be neater if all this were the result of two master plans; it isn't, since the basic metaphor introduces a kind of intentional purpose on the part of a patron, which the evidence does not back up. Neither empire came up with this plan. Their failures are mirror images of each other, and it was only in the last three years that either one had begun to act with deliberate intent.</p><p style="text-align: justify;"><strong><span>China originally built the ladder by accident and has since taken on the role of</span></strong><span>&nbsp;governing it. According to peer-reviewed literature, the "Digital Silk Road" has no clear plan of action or strategy for implementation; as of the most recent public count, only 16 countries had signed a DSR agreement, contrary to Huawei's own figure of "800+ government cloud projects worldwide."[30] A study by AidData, which examined more than thirteen thousand Chinese projects, found that joining the Belt and Road did not alter in the slightest who Beijing funded or the terms on which funding was provided: it was essentially a rebranding rather than a shift in strategic direction.[31] Chinese digital lending to Ethiopia began in 2006, 13 years before Addis signed any agreement that included the phrase "Belt and Road". A truthful account of the period from 2006 to 2022 is that Huawei was pursuing revenue, the two policy banks were trying to increase their loan volumes, and Beijing simply painted a slogan on the outcome.</span></p><p style="text-align: justify;">Yet emergent sequencing had nevertheless created a genuine ladder, and you can climb it. In the case of Senegal, over eighteen years with just one vendor: a government intranet in 2006, followed by three thousand kilometers of fiber, then an $85 million China EXIM broadband initiative whose key element was a national data center into which government data was sent back, then a Chinese-funded 'Smart Senegal', and then a Safe City program featuring nearly five hundred facial-recognition cameras.[32] Pakistan has the entire ladder along with the geographical advantages&#8212;fibre running along the China&#8211;Pakistan Economic Corridor, the Huawei-built PEACE cable at Karachi and Gwadar, safe cities, and now a Huawei AI data centre&#8212;as well as the same vendor at each stage and the policy banks at most of the steps.</p><p style="text-align: justify;"><span>Since October 2023, the retrofit has become intentional. At the Belt and Road Forum, Xi launched China's Global AI Governance Initiative.&nbsp;</span><strong><span>In&nbsp;July 2026, the World AI Cooperation Organization was established in Shanghai with twenty-nine founding members, including Laos and Pakistan, which are described as ladder clients. The UN's Group of Friends on AI capacity-building, consisting of some eighty countries</span></strong><span>, is jointly chaired by China and Zambia&#8212;a country that was a $65 million client of EXIM and Huawei in the first, pre-AI round of data center lending and whose own parliament had recorded hardware that had been installed in a district without any electricity.[33] The country that had been in debt then became a co-sponsor. Throughout all this institutional development, Chinese sovereign lending commitments, according to Boston University's database of the two policy banks, have collapsed from $62.5 billion in 2016 to $6.1 billion in 2024.[34] China is now running the institutions while spending 90% less money. The brand began as fiction and later became a reality.</span></p><p style="text-align: justify;"><strong><span>At the same time, the United States has funds without a clear plan.</span></strong><span>&nbsp;In principle the necessary structures are in place: an executive order of July 2025 concerning the export of the 'complete American AI stack', an American AI Exports Programme, a diplomatic group known as Pax Silica, Technology Prosperity Deals, a new EXIM financing facility, and a development finance corporation for which Congress has more than tripled the lending limit, raising it from $60 billion to $205 billion.</span></p><p style="text-align: justify;">In reality, the executive order does not name any specific countries; it is up to private consortia to choose the markets, and the Commerce Department issued an official request asking the public to indicate which countries should be given priority. Twelve months after the order, no export packages have been designated, and no EXIM AI deals have been financed. All four of the Technology Prosperity Deals have gone to the United Kingdom, Japan, South Korea, and Sweden. At its June summit, Pax Silica had twenty-four signatories, none of whom were from Africa, and its main assistance program amounts to $50 million &#8212; for a supply-chain credentialing platform which has been piloted in Panamanian ports.[35] The White House AI czar himself described the importance of the situation by saying: &#8220;<strong>China is exporting Huawei chips and DeepSeek models to the Global South. If we don&#8217;t make it just as easy to export the American AI stack, we will lose this technology race in large parts of the world.</strong>&#8221;[36]</p><p style="text-align: justify;">The American response has instead been one of exclusivity. In August 2026, Washington drew up letters to the thirty-five countries that had signed the AI Opportunity statement, stating that membership in Pax Silica could not be combined with 'duplicative initiatives' &#8211; a formulation officials confirmed was directed at China's new Shanghai organization. "You can't have it both ways," one official put it.[37]</p><p style="text-align: justify;"><span>As for symmetry, China has constructed a ladder which it both finances and operates, and now sells the franchise; the United States, on the other hand, sells permission slips. In both cases, the host is not paid. And now that both of them are in control, each has begun to demand the exclusive loyalty that patrons had previously been able to&nbsp;purchase. </span><strong><span>During the Cold War, when a patron asked for exclusivity, the check was included. The AI empires have dropped the check but retained the clause</span></strong><span>. It is precisely this that makes the next step possible: an auction requires two bidders who both care, and this is the first time that both do.</span></p><h2>The one that played it right</h2><p>This brings us back to Port Moresby, where the two buildings can be understood.</p><p style="text-align: justify;">Papua New Guinea never managed to get a data center operating, since it never really needed to: it gained the benefits of an auction it didn't have to organize by simply appearing cooperative and, at the same time, clearly vulnerable to being aligned with the other side. The Huawei facility attracted Beijing's attention in 2018. Its subsequent public failure intensified the competition that Canberra was already engaged in for its own reasons. In the Pacific, Washington's side of this market is supported through its ally: the Coral Sea Cable, which is two-thirds funded by Australia in order to keep Huawei Marine out of the Pacific waters, and then, in December 2025, three Google subsea cables whose US$120 million domestic sections are funded by Australia under the Pukpukdefensee treaty, the 2025 Australia&#8211;PNG security pact.[38] And now, in 2026, there is a Western-stack sovereign cloud with a bill that no one will reveal &#8212; the only entry in PNG's ledger that might yet go wrong.</p><p style="text-align: justify;">The fact that the money mainly went to Port Moresby is the key issue: this represents the proper historical direction, that of patron to host. PNG achieved this not by running its own servers but by being present on contested territory and allowing the anxiety to carry out the negotiations. It's not exactly like Manas &#8212; since PNG didn't run the auction &#8212; but it does follow the same lesson from Kyrgyzstan: <strong>be situated where both patrons have an interest.</strong></p><p style="text-align: justify;">Kenya played the opposite strategy and suffered two setbacks. It managed to secure the first round only to obtain a facility which its own Auditor-General has described as flawed. It then failed to complete the second round, as the billion-dollar Microsoft&#8211;G42 project came to a standstill when the government refused to provide the sovereign guarantees the deal called for and faced the facts. President Ruto, who, to his credit, finally admitted the obvious: "In order to switch on that one data center, we would have to cut off the power to half the country. It was then that I realized there was a problem." [39] Kenya paid for the base but received neither the rent nor the computing power.</p><p style="text-align: justify;"><span>The idea that developing&nbsp;countries do odd things&nbsp;is based on a misunderstanding of transparency. Kenya's failures are set out in full, item by item, because the Auditor-General of Kenya publishes them. The reason that Zambia's unused equipment is recorded is that Zambia's parliamentary committees publish it. However, if you look for usage data regarding a sovereign cloud program in Europe, you won't find it. The shortcomings of the long tail are apparent in how audit bodies in that area function; failures elsewhere are simply not published.</span></p><h2>Where this goes</h2><p style="text-align: justify;">The reason the deals keep being concluded is that the purchases appear irrational when viewed merely as infrastructure &#8212; they aren't actually infrastructure; instead, they function as alignment tools and are treated as capital expenditures. Each party in the chain receives payment before the actual utility has been tested: <strong>the vendor is paid at the time of shipment, the integrator at the time of completion, the lender according to a schedule which has nothing to do with the level of demand, and the minister at the ribbon-cutting ceremony, which is timed to an election</strong>. The only party whose return depends on the building actually functioning is the state and its taxpayers. This is why none of the fifty-nine programs have disclosed a power purchase agreement. Such disclosure would introduce a price for the one risk that no one in the chain is actually taking on. Moreover, "insurance against foreign jurisdiction" does not save the purchases: the hedge fails because the jurisdiction arrives along with the racks &#8212; with imported silicon on revocable licenses and, at scale, a customer ledger being filed in Washington.</p><p>So: four calls, forward.</p><p style="text-align: justify;"><strong><span>The auction will start &#8212; I'm putting the date at 2028.</span></strong><span>&nbsp;Although so far no state that is not aligned has managed to trigger a bidding war on the lines of the Manas example, all the necessary elements are now available: two patrons who want exclusivity in the area of writing, a public playbook provided by Port Moresby, and a shortlist of states that actually have leverage &#8212; namely, stranded energy, cable chokepoints, minerals, and international votes. Keep an eye on Djibouti itself, Morocco, Kazakhstan, and the Indonesian archipelago. The first state to be&nbsp;paid&nbsp;to host the computing facilities &#8212; payment without taking on debt &#8212; will have the same effect on this market as Kyrgyzstan did with base arrangements in 2009: </span><strong><span>it will make the cost of alignment visible</span></strong><span>. All the other host states will then renegotiate the following morning.</span></p><p style="text-align: justify;"><strong><span>Second, the distress cycle will arrive on time, before the decade is over.</span></strong><span>&nbsp;A GPU that is serviced remains frontier-relevant for three to five years; the grace periods typically afforded by China's EXIM facilities of this kind are five to seven years&#8212;Kenya's, in the case that has been documented, was about seven. The first sovereign debt restructuring to include a dead data center is not a remote possibility. Zambia signaled the political implications in 2022 when it canceled the second round on the grounds of debt sustainability. The workout will have to be arranged since there is no published framework from either the IMF or the World Bank for valuing stranded compute. </span><strong><span>And the buildings will most likely be cleared in the only way that empty racks with grid connections and fiber can be cleared: by being sold at a discount to the hyperscalers and GPU brokers they were intended to escape</span></strong><span>. Bangladesh has already gone through this scenario: a Tier IV sovereign facility, with its hardware now obsolete, and the government's workloads having been migrated to servers in Singapore operated by Oracle.[40] On the current trajectory, sovereignty's end state will be that of an acquisition channel.</span></p><p style="text-align: justify;"><strong><span>Thirdly, the stacks split apart, and the long tail becomes the testing area. I anticipate</span></strong><span>&nbsp;that the American proposal will move downward toward the lower end of the market as the Gulf model becomes standardized by offering a full stack, VEU surveillance, and exclusivity letters. The Chinese offer, on the other hand, will move upwards in its new, capital-light version: replacing loans with reference architectures, using Ascend as the supply situation eases, making open-weight models the standard, and having the Shanghai institutions hold the votes. The countries included in this census are those in which the two offers will come together, with no allied group, no adequacy framework, and no EuroHPC to take in the clash. Malaysia's forty-eight hours served as a rehearsal; savvy host countries will realize that the credible&nbsp;threat&nbsp;of selecting Beijing is itself a valuable asset with market value.</span></p><p style="text-align: justify;"><strong><span>Fourth, the currency in the following round is not computing power; it is data, energy, and votes.</span></strong><span>&nbsp;The patrons have no need for the long tail's demands. NVIDIA's revenue attribution does not mention any of these countries. What the long tail does have that the empires actually want is </span><strong><span>training data, grid capacity, minerals, cable landings, and a place within the institutions that China is setting up</span></strong><span> &#8211; and this process has already begun. In 2026, the United States has entered into bilateral health agreements in Africa on the condition that aid be tied to data: Uganda has received grants from Washington in exchange for direct, real-time access to nine of its national health data systems for seven years; in six of the agreements, the host country is required to begin sharing pathogen samples within five days of a U.S. request.[41] </span></p><p style="text-align: justify;"><span>The empires have already found where the long tail's true resources lie, and they are not located in server rooms. The 'AI partnership' of 2028 is likely to appear less like a data center and more like a package &#8211; comprising your data, your resources, your UN vote &#8211; in exchange for membership rather than money.</span></p><div><hr></div><h2>Notes</h2><ol><li><p><a href="https://www.nbc.com.pg/post/38905/soe-minister-duma-to-unveil-pngs-sovereign-cloud-and-ai-datacenter"><span>The SOE Minister Duma will reveal Papua New Guinea's sovereign cloud and AI datacentre in</span></a><span>&nbsp;August 2026;&nbsp;</span><a href="https://www.kch.com.pg/papua-new-guinea-advances-with-launch-of-first-sovereign-ai-data-centre/"><span>Kumul Consolidated Holdings&nbsp;</span></a><span>will also</span><a href="https://www.kch.com.pg/papua-new-guinea-advances-with-launch-of-first-sovereign-ai-data-centre/"><span>&nbsp;launch its release</span></a><span>.</span></p></li><li><p><a href="https://emtv.com.pg/datec-launches-kumul-cloud-infinity-pngs-first-sovereign-cloud-and-ai-data-center/"><span>Datec announces the launch of Kumul Cloud Infinity</span></a><span>&nbsp;(22 August 2026);&nbsp;</span><a href="https://www.cloudsigma.com/about-us"><span>CloudSigma&nbsp;</span></a><span>provides</span><a href="https://www.cloudsigma.com/about-us"><span>&nbsp;information about</span></a><span>&nbsp;itself (headquarters in Zurich, with 35 or more partner countries). Regarding the launch, the hardware vendors involved (HPE, NVIDIA H200) and the local gas-backup partner are included; the only providers named in the KCH release are Telikom, Datec and CloudSigma.</span></p></li><li><p><a href="https://china.aiddata.org/"><span>The Chinese Official Finance dataset is provided by AidData</span></a><span>; the loan amount is given by&nbsp;</span><a href="https://www.theglobeandmail.com/world/article-huawei-built-data-centre-a-failed-investment-papua-new-guinea-says/"><span>The Globe and Mail</span></a><span>&nbsp;and&nbsp;</span><a href="https://www.datacenterdynamics.com/en/news/australia-huaweis-papua-new-guinea-data-center-security-openly-broken-making-potential-spying-easy/"><span>Data Center Dynamics</span></a><span>. The figure of $56 million is cited in The Diplomat (2021); the amount in kina (K130m) does not reconcile properly with either of the other figures.</span></p></li><li><p><a href="https://www.datacenterdynamics.com/en/news/australia-huaweis-papua-new-guinea-data-center-security-openly-broken-making-potential-spying-easy/"><span>With regard&nbsp;</span></a><span>to</span><a href="https://www.datacenterdynamics.com/en/news/australia-huaweis-papua-new-guinea-data-center-security-openly-broken-making-potential-spying-easy/"><span>&nbsp;Australia, the security of Huawei's data centre in Papua New Guinea was found to be "openly broken"</span></a><span>&nbsp;(August 2020), following the release of the 65-page assessment commissioned by DFAT. The report itself is not available to the public; the findings are quoted as reported.</span></p></li><li><p><span>Timothy Masiu, who at the time held the position of Minister for Communication and Information Technology, stated in&nbsp;</span><a href="https://www.theglobeandmail.com/world/article-huawei-built-data-centre-a-failed-investment-papua-new-guinea-says/"><span>The Globe and Mail</span></a><span>&nbsp;(August 2020).</span></p></li><li><p><span>Regarding the debate about the debt trap:&nbsp;</span><a href="https://www.tandfonline.com/doi/full/10.1080/23792949.2019.1689828"><span>Deborah Brautigam&#8217;s &#8220;A critical look at Chinese &#8216;debt-trap diplomacy&#8217;&#8221;</span></a><span>, published in</span><a href="https://www.tandfonline.com/doi/full/10.1080/23792949.2019.1689828"><span>&nbsp;</span></a><em><a href="https://www.tandfonline.com/doi/full/10.1080/23792949.2019.1689828"><span>Area Development and Policy</span></a></em><a href="https://www.tandfonline.com/doi/full/10.1080/23792949.2019.1689828"><span>&nbsp;in 2020</span></a><span>; and&nbsp;</span><a href="https://docs.aiddata.org/reports/delivering-the-belt-and-road.html"><span>AidData&#8217;s report on&nbsp;</span></a><em><a href="https://docs.aiddata.org/reports/delivering-the-belt-and-road.html"><span>Delivering the Belt and Road</span></a></em><span>. The repayment situation differs from country to country &#8212; Zambia, mentioned later, defaulted in 2020 and had its debt restructured &#8212; yet the intention behind these facility loans is for them to be repaid on time, not for assets to be seized.</span></p></li><li><p><a href="https://www.fool.com/earnings/call-transcripts/2026/02/25/nvidia-nvda-q4-2026-earnings-call-transcript/"><span>In the NVIDIA Q4 FY2026 earnings call held on</span></a><span>&nbsp;25 February 2026, sovereign AI revenues increased more than three times on a year-on-year basis, mostly due to customers in Canada, France, the Netherlands, Singapore and the UK;&nbsp;</span><a href="https://efts.sec.gov/LATEST/search-index?q=%22sovereign%20AI%22"><span>a search of the EDGAR database using</span></a><span>&nbsp;CIK 0001045810 shows that the term 'sovereign AI' does not appear in any of the 10-K or 10-Q filings, although it is used in the unaudited annual report wrapper and in the 8-K press releases; the&nbsp;</span><a href="https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-first-quarter-fiscal-2027"><span>NVIDIA Q1 FY2027 press release</span></a><span>&nbsp;of 20 May 2026 contains no mention of 'sovereign' and instead includes a new Hyperscale/ACIE disclosure, while the&nbsp;</span><a href="https://www.fool.com/earnings/call-transcripts/2026/05/20/nvidia-nvda-q1-2027-earnings-transcript/"><span>Q1 FY2027 call transcript</span></a><span>&nbsp;states that sovereign revenue rose by more than 80% year on year compared to the company's overall revenue growth of 85%.</span></p></li><li><p><a href="https://www.aljazeera.com/news/2026/4/8/our-geography-is-our-oil-why-djibouti-hosts-many-foreign-military-bases"><span>The article in Al Jazeera,&nbsp;</span></a><span>entitled</span><a href="https://www.aljazeera.com/news/2026/4/8/our-geography-is-our-oil-why-djibouti-hosts-many-foreign-military-bases"><span>&nbsp;&#8220;Our geography is our oil&#8221;: Why Djibouti hosts so many foreign military bases</span></a><span>&nbsp;(April 2026), gives figures for the rent (US$65 million for the United States, $30 million for France, $20 million for China, about $3 million each for Italy and Japan; the total amount being about $121 million), according to Ilyas Dawaleh, the country's Finance Minister, and these figures are based on those from 2017. The expression &#8220;our geography is our oil&#8221; is used in the article itself, while the quoted statement (&#8220;our geography is our main national resource&#8221;) is from an unnamed Djiboutian official.</span></p></li><li><p><a href="https://www.everycrsreport.com/reports/R40564.html"><span>The</span></a><span>&nbsp;Manas rent increased from about $2 million in 2001 to $17.4 million in 2006 and then to $60 million in 2009; the Russian package included an investment of about $1.7 billion and loans/aid amounting to $450 million. The 2009 re-lease was known as the "Transit Center at Manas."</span></p></li><li><p><a href="https://commonslibrary.parliament.uk/research-briefings/cbp-10273/"><span>The House of Commons Library (CBP-10273)</span></a><span>&nbsp;and&nbsp;</span><a href="https://fullfact.org/politics/chagos-deal-true-cost/"><span>Full Fact have stated that</span></a><span>&nbsp;the average annual payment under the deal in 2025/26 prices is &#163;101m; the Government Actuary&#8217;s Department has valued the total over the 99-year period at &#163;3.4bn on a net present value basis, as opposed to a nominal cash total of about &#163;34.7bn. The legislation necessary to implement the deal came to a standstill in the Commons in early 2026 and the deal was reported to have been put on hold from April 2026, with no payments having been made (</span><a href="https://commonslibrary.parliament.uk/research-briefings/cbp-10464/"><span>CBP-10464</span></a><span>).</span></p></li><li><p><a href="https://www.mofa.go.jp/press/release/press4e_003074.html"><span>The agreement between Japan and the Ministry&nbsp;</span></a><span>of</span><a href="https://www.mofa.go.jp/press/release/press4e_003074.html"><span>&nbsp;Foreign Affairs for host-nation support during the period FY2022&#8211;26 (amounting</span></a><span>&nbsp;to about &#165;211 billion per year; the dollar value varies between $1.4 billion and $2 billion according to the exchange rate); the&nbsp;</span><a href="https://www.koreaherald.com/article/3486962"><span>Korea&#8211;US Special Measures Agreement for 2026</span></a><span>&nbsp;(value of &#8361;1.5192 trillion, or about $1.1 billion).</span></p></li><li><p><span>Kent Calder,&nbsp;</span><em><span>Embattled Garrisons</span></em><span>&nbsp;(Princeton, 2007); Alexander Cooley,&nbsp;</span><em><span>Base Politics</span></em><span>&nbsp;(Cornell, 2008);&nbsp;</span><a href="https://www.cambridge.org/core/journals/perspectives-on-politics/article/abs/empire-will-compensate-you-the-structural-dynamics-of-the-us-overseas-basing-network/13A040F9E14540969BAF778BA021D75E"><span>and Cooley&nbsp;</span></a><span>and</span><a href="https://www.cambridge.org/core/journals/perspectives-on-politics/article/abs/empire-will-compensate-you-the-structural-dynamics-of-the-us-overseas-basing-network/13A040F9E14540969BAF778BA021D75E"><span>&nbsp;Nexon, &#8220;The Empire Will Compensate You,&#8221;&nbsp;</span></a><em><a href="https://www.cambridge.org/core/journals/perspectives-on-politics/article/abs/empire-will-compensate-you-the-structural-dynamics-of-the-us-overseas-basing-network/13A040F9E14540969BAF778BA021D75E"><span>Perspectives on Politics</span></a></em><a href="https://www.cambridge.org/core/journals/perspectives-on-politics/article/abs/empire-will-compensate-you-the-structural-dynamics-of-the-us-overseas-basing-network/13A040F9E14540969BAF778BA021D75E"><span>&nbsp;(2013)</span></a><span>. My version of the two-directional approach is a synthesis since Cooley and Nexon propose the patron-pays-client aspect while the alliance burden-sharing literature (Olson and Zeckhauser) deals with the other direction.</span></p></li><li><p>The data was compiled in August 2026 using government publications, documents from international financial institutions, announcements from vendors, and reports from the press; a full table with the sources for each row is provided in a companion appendix. The countries in question deliberately omit the United States, China, the major members of the EU, the UK, Japan, Korea, India, Israel, Singapore, Canada, Australia and the Gulf states. Although the builds funded by EuroHPC (in Greece, Bulgaria and Romania) are included, they are regarded as patron-to-host in a different way: an EU member country that makes use of a facility which it co-funds is a shareholder, not the host of a facility financed by a foreign patron. Kenya's Konza facility, the example worked out below, was a precedent from 2019 referred to in the text and is not counted twice in the current aggregates.</p></li><li><p><a href="https://china.aiddata.org/projects/59366/"><span>Project number 59366</span></a><span>&nbsp;(involving China Eximbank, amounting to 1.225 billion RMB, signed on 26 April 2019, with a 2% fixed rate, a maturity of 19.5 years, a grace period of about 7 years and the first repayment due on 31 March 2026; Huawei acting as the EPC contractor). The value of this AidData project in constant 2023 U.S. dollars is $185.2 million; it was equivalent to about $180 million at the exchange rate in effect at the time of the signing in 2019.</span></p></li><li><p><a href="http://libraryir.parliament.go.ke/bitstream/handle/123456789/25376/Report%20of%20the%20Auditor-General%20on%20Konza%20Technopolis%20Development%20Authority%20for%20the%20Year%20Ended%2030%20June,2023.pdf?sequence=1&amp;isAllowed=y"><span>The Auditor-General's report concerning the Konza Technopolis Development Authority as at 30 June YE</span></a><span>.</span></p></li><li><p><a href="https://kenyanwallstreet.com/konza-technopolis-loss-june-2025"><span>Kenyan Wallstreet &#8212; the results for Konza in FY2025&nbsp;</span></a><span>(cloud revenue dropped by 16.4% to KSh 126.8 million for the year ending June 2025; approximately $0.98 million).</span></p></li><li><p><a href="https://openai.com/global-affairs/openai-for-countries/"><span>OpenAI &#8212; OpenAI for Countries</span></a><span>&nbsp;(in May 2025 it was stated that partner countries would also invest in the expansion of the global Stargate Project).</span></p></li><li><p><span>The direction of the flow was stated by OpenAI itself in its announcement (</span><a href="https://openai.com/index/introducing-stargate-uae"><span>Introducing Stargate UAE</span></a><span>: "UAE investment into U.S. Stargate infrastructure"); the dollar-for-dollar ratio referred to by&nbsp;</span><a href="https://www.axios.com/2025/05/22/uae-openai-stargate-deal"><span>Axios</span></a><span>&nbsp;(22 May 2025) regarding the deal framework, not a clause in a published contract.</span></p></li><li><p><a href="https://www.exim.gov/news/exim-launches-exportai-initiative-strengthen-american-leadership-ai"><span>The ExportAI initiative</span></a><span>&nbsp;(approved by the board on 21 May 2026) and the&nbsp;</span><a href="https://www.federalregister.gov/documents/2025/10/28/2025-19674/american-ai-exports-program"><span>Federal Register&nbsp;</span></a><span>entry for the</span><a href="https://www.federalregister.gov/documents/2025/10/28/2025-19674/american-ai-exports-program"><span>&nbsp;American AI Exports Program</span></a><span>&nbsp;(28 October 2025) indicate that the toolkit includes direct loans, guarantees and insurance; the buyer's credit facility is a form of financing provided to foreign buyers of US chips.</span></p></li><li><p><a href="https://www.dfc.gov/media/press-releases/dfc-disburses-83-million-africa-data-centres-expand-ict-infrastructure-south"><span>The US DFC&nbsp;</span></a><span>has</span><a href="https://www.dfc.gov/media/press-releases/dfc-disburses-83-million-africa-data-centres-expand-ict-infrastructure-south"><span>&nbsp;disbursed $83 million to Africa Data Centres</span></a><span>&nbsp;(commitment amounting to up to $300 million in total; the $83 million figure refers to the first disbursement and a later statement deals with a facility in Ghana under this arrangement).</span></p></li><li><p><a href="https://primeminister.kz/en/news/data-center-valley-kazakhstan-government-firebird-and-nvidia-sign-10-billion-package-of-agreements-in-artificial-intelligence-31516"><span>The package for Data Center Valley&nbsp;</span></a><span>(15 June 2026: the minister forecasts "at least $3 billion in annual export revenue");&nbsp;</span><a href="https://www.cerebras.ai/press-release/guyana"><span>Cerebras &#8212; Guyana</span></a><span>&nbsp;("international demand").</span></p></li><li><p><a href="https://www.ecfr.gov/current/title-15/subtitle-B/chapter-VII/subchapter-C/part-750/section-750.8"><span>Section 15 CFR &#167; 750.8(a)</span></a><span>&nbsp;deals with licenses granted pursuant to the Export Administration Regulations, which include advanced AI chips; separate provisions apply to ITAR licenses (see 22 CFR &#167; 120.18) with a similar effect.</span></p></li><li><p><span>The number of licences revoked, as reported in the Commerce correspondence for May 2024 (</span><a href="https://www.scmp.com/tech/tech-war/article/3268946/tech-war-us-revoked-8-licences-exporting-goods-chinas-huawei-2024"><span>SCMP</span></a><span>);&nbsp;</span><a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581023000217/nvda-20231017.htm"><span>NVIDIA's Form 8-K of 17 October 2023</span></a><span>&nbsp;(naming Saudi Arabia, the UAE and Vietnam).</span></p></li><li><p><a href="https://www.federalregister.gov/documents/2026/07/14/2026-14132/enhanced-favorable-treatment-for-the-united-arab-emirates-under-the-export-administration"><span>In the Federal Register &#8212; Enhanced Favorable Treatment for the United Arab Emirates</span></a><span>&nbsp;(to come into effect on 10 July 2026), Supplement No. 8, eight named US companies together with certain UAE government bodies, and also G42 and Core42, will have their eligibility expire on 6 April 2027 unless a restructuring takes place.</span></p></li><li><p><a href="https://www.federalregister.gov/documents/2024/10/02/2024-22587/expansion-of-validated-end-user-authorization-data-center-validated-end-user-authorization"><span>In the Federal Register &#8212; Data Center Validated End User Authorization</span></a><span>&nbsp;(2 October 2024) it states: &#8220;compliance checks carried out by representatives of the United States Government&#8221; and semiannual reports which include information on chip inventory, compute utilization, and &#8220;a list of current customers together with a description of their utilization.&#8221;</span></p></li><li><p><a href="https://edition.cnn.com/2022/05/01/europe/russia-farm-vehicles-ukraine-disabled-melitopol-intl"><span>CNN&nbsp;</span></a><span>&#8211;</span><a href="https://edition.cnn.com/2022/05/01/europe/russia-farm-vehicles-ukraine-disabled-melitopol-intl"><span>&nbsp;Deere/Melitopol</span></a><span>&nbsp;(May 2022; the equipment was remotely locked through the dealer's connectivity, as reported; Deere has not confirmed this);&nbsp;</span><a href="https://therecord.media/russians-losing-access-microsoft-cloud-amazon"><span>The Record&nbsp;</span></a><span>&#8211;</span><a href="https://therecord.media/russians-losing-access-microsoft-cloud-amazon"><span>&nbsp;Microsoft/AWS&#8211;Russia</span></a><span>&nbsp;(March 2024: Microsoft suspended access for its existing customers as a result of EU sanctions; AWS said that it had blocked only new Russian customers since 2022); Reuters-syndicated reporting concerning the Nokia/Ericsson exit (December 2022);&nbsp;</span><a href="https://www.theregister.com/2024/05/21/asml_kill_switch/"><span>reporting on ASML's assurances regarding the remote-disable of its EUV equipment to the Dutch government</span></a><span>&nbsp;(May 2024; these assurances were not confirmed by ASML).</span></p></li><li><p><a href="https://www.globaltimes.cn/page/202507/1339752.shtml"><span>The Global Times&nbsp;</span></a><span>reported that the</span><a href="https://www.globaltimes.cn/page/202507/1339752.shtml"><span>&nbsp;CAC called on NVIDIA</span></a><span>&nbsp;on 31 July 2025;&nbsp;</span><a href="https://blogs.nvidia.com/blog/no-backdoors-no-kill-switches-no-spyware/"><span>in August 2025 the NVIDIA blog&nbsp;</span></a><span>stated,</span><a href="https://blogs.nvidia.com/blog/no-backdoors-no-kill-switches-no-spyware/"><span>&nbsp;"No Backdoors. No Kill Switches. No Spyware.</span></a><span>"</span></p></li><li><p><a href="https://www.congress.gov/bill/119th-congress/house-bill/3447">Chip Security Act (H.R. 3447)</a>, passed the House Foreign Affairs Committee 42&#8211;0 on 26 March 2026; no floor vote as of this writing. Senate companion: <a href="https://www.congress.gov/bill/119th-congress/senate-bill/1705">S. 1705</a>.</p></li><li><p><a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/u-s-issues-worldwide-crackdown-on-using-huawei-ascend-chips-says-it-violates-export-controls">BIS guidance on worldwide Huawei Ascend use</a> (issued ~13 May 2025); <a href="https://developingtelecoms.com/telecom-technology/data-centres-networks/18504-malaysia-launches-sovereign-ai-stack-with-deepseek-and-huawei-gpus.html">Developing Telecoms &#8212; Malaysia sovereign AI launch</a> (19 May 2025); <a href="https://www.malaymail.com/news/malaysia/2025/05/21/malaysia-walks-back-huawei-ai-project-amid-us-china-tech-tensions/177551">Malay Mail &#8212; retraction</a> (21 May 2025).</p></li><li><p><a href="https://www.tandfonline.com/doi/full/10.1080/10670564.2023.2222269">&#8220;&#8217;Digital Silk Road&#8217; as a Slogan Instead of a Grand Strategy,&#8221; </a><em><a href="https://www.tandfonline.com/doi/full/10.1080/10670564.2023.2222269">Journal of Contemporary China</a></em> (2023; &#8220;neither a clear roadmap nor an implementation strategy&#8221;); <a href="https://www.eurasiagroup.net/files/upload/Digital-Silk-Road-Expanding-China-Digital-Footprint-1.pdf">Eurasia Group &#8212; The Digital Silk Road</a> (2020; its &#8220;Countries signing DSR-specific MOU with China&#8221; table lists 16 countries); <a href="https://e.huawei.com/en/news/2025/industries/government/rise-national-government-cloud-architecture">Huawei &#8212; R.I.S.E. reference-architecture launch</a> (17 Sept 2025; &#8220;800+ government cloud projects&#8221;).</p></li><li><p><a href="https://docs.aiddata.org/reports/delivering-the-belt-and-road.html">AidData &#8212; </a><em><a href="https://docs.aiddata.org/reports/delivering-the-belt-and-road.html">Delivering the Belt and Road</a></em> (analysis of more than 13,000 projects, 2000&#8211;2017).</p></li><li><p><a href="https://africachinainitiative.georgetown.edu/research-working-group/blog-posts/chinese-huawei-group-and-the-building-of-senegalese-digital-grammar-from-2006-to-2024/">Georgetown Africa&#8211;China Initiative &#8212; Huawei and the building of Senegalese digital grammar, 2006&#8211;2024</a> ($85m China Eximbank broadband program, of which the Diamniadio data center was the centerpiece; ~500 facial-recognition cameras); AidData project records. (A lower disbursed-tranche figure &#8212; 46bn CFA, ~$83m &#8212; appears in Data Center Dynamics coverage; the appendix reconciles the two.) Pakistan&#8217;s current Huawei AI data-center project is privately financed (<a href="https://www.datacenterdynamics.com/en/news/indus-cloud-and-huawei-to-develop-data-center-in-pakistan/">DCD &#8212; Indus Cloud/Huawei</a>).</p></li><li><p><a href="https://www.fmprc.gov.cn/eng/wjb/zzjg_663340/jks_665232/jkxw_665234/202505/t20250521_11629968.html">Chinese MFA &#8212; UN Group of Friends for International Cooperation on AI Capacity-Building</a> (China&#8211;Zambia co-chairs); <a href="https://www.parliament.gov.zm/node/11808">Report of the Committee on Cabinet Affairs on SMART Zambia, Thirteenth National Assembly</a> (hardware installed in a district without electricity); <a href="https://china.aiddata.org/projects/53093">AidData project #53093</a> (~$65m China EXIM concessional, Huawei contractor); <a href="https://english.news.cn/20260716/b0449aa2133542868e310fdc45ef2969/c.html">Xinhua &#8212; WAICO founding</a> (29 founding states, Shanghai, July 2026).</p></li><li><p><a href="https://www.bu.edu/gdp/files/2025/07/GCI-PB-26-CODF-2025-FIN.pdf">Boston University Global Development Policy Center &#8212; 2025 brief on Chinese overseas development finance</a> (sovereign lending commitments by China&#8217;s two policy banks: $62.5bn in 2016; $6.1bn in 2024).</p></li><li><p><a href="https://www.whitehouse.gov/presidential-actions/2025/07/promoting-the-export-of-the-american-ai-technology-stack/">Executive Order 14320</a> (23 July 2025); <a href="https://www.federalregister.gov/documents/2026/04/10/2026-06952/american-ai-exports-program-call-for-proposals-for-pre-set-consortia">Federal Register &#8212; call for proposals</a> (10 Apr 2026); <a href="https://www.federalregister.gov/documents/2025/10/28/2025-19674/american-ai-exports-program">Federal Register &#8212; American AI Exports Program</a> (28 Oct 2025; the request for information asks, among other questions, which countries or regions should be priorities); <a href="https://www.whitehouse.gov/articles/2025/10/the-united-states-signs-technology-prosperity-deals-with-japan-and-korea/">White House &#8212; Technology Prosperity Deals</a> (UK, Japan, Korea; Sweden added May 2026); <a href="https://pl.usembassy.gov/outcomes-of-the-second-pax-silica-summit/">State Department &#8212; Outcomes of the Second Pax Silica Summit</a> (June 2026; 24 signatories, none African); <a href="https://www.state.gov/releases/office-of-the-spokesperson/2026/08/pax-silica-ai-assistance-project-nofo/">Pax Silica AI Assistance Project NOFO</a> (Aug 2026; up to $50m; Panama ports pilot). The DFC lending cap was raised from $60bn to $205bn under the FY26 reauthorization (<a href="https://www.cgdev.org/blog/dfc-reauthorization-whats-new-and-what-it-means">CGD</a>).</p></li><li><p><a href="https://x.com/DavidSacks/status/1981128143646695651">David Sacks, X</a> (22 Oct 2025).</p></li><li><p><a href="https://www.gmanetwork.com/news/topstories/world/998694/pax-silica-or-not-us-to-tell-partners-they-must-pick-sides-in-ai-race-with-china/story/">Reuters-derived reporting on US letters to AI Opportunity signatories</a> (Aug 2026): a State Department <em>draft</em> letter to 35 signatories bars &#8220;duplicative initiatives whose expectations conflict with our own&#8221;; officials confirmed the target is China&#8217;s Shanghai organization. The letter was drafted, not yet sent, as of the reporting.</p></li><li><p><a href="https://png.embassy.gov.au/pmsb/1148.html">Australian High Commission PNG &#8212; Coral Sea Cable System</a> (A$200M, ~two-thirds Australian grant; the funding decision was driven by Solomon Islands, with PNG added); <a href="https://www.abc.net.au/news/2025-12-13/google-building-undersea-cables-in-png/106139500">ABC News &#8212; Google cables for PNG</a> (13 Dec 2025; three Google subsea cables, US$120M domestic legs Australian-funded under the 2025 Pukpuk Treaty).</p></li><li><p>William Ruto (May 2026), per <a href="https://www.datacenterdynamics.com/en/news/microsoft-and-g42-data-center-in-kenya-stalled-due-to-lack-of-power-capacity/">Data Center Dynamics</a> and Kenyan press reporting on the stalled Microsoft&#8211;G42 Olkaria project, which the government declined to backstop with sovereign guarantees.</p></li><li><p>Faiz Ahmad Taiyeb (Special Assistant to the Chief Adviser, Bangladesh), per <a href="https://www.industryinsiderbd.com/article/billions-down-the-drain-as-state-run-data-centers-faltered">Industry Insider BD</a>: the Kaliakoir national data center &#8220;a complete failure&#8230; its hardware became obsolete over three years ago,&#8221; with workloads migrated to Oracle&#8217;s Singapore region; a separate in-country Oracle deployment serves other government entities.</p></li><li><p><a href="https://www.propublica.org/article/trump-state-department-africa-uganda-aid-medical-data-privacy">ProPublica &#8212; U.S. Demands to Access Africans&#8217; Data Raise Privacy, Sovereignty Concerns</a> (17 June 2026: Uganda&#8217;s agreement grants &#8220;direct, real-time access to nine of the nation&#8217;s health data systems for seven years&#8221;; six pathogen agreements require the host to begin sharing specimens within five days of a US request). Kenya signed a $1.6bn version in December 2025; a court froze it days after signing, and the Court of Appeal temporarily lifted the freeze in May 2026, letting implementation proceed while the case continues. Ghana, Zambia and Zimbabwe rejected the initial agreements.</p></li></ol>]]></content:encoded></item><item><title><![CDATA[The Watcher Is the Product]]></title><description><![CDATA[DeepSeek gave the agent harness away. OpenAI is taxing itself to watch its own. SpaceX paid $60 billion for one. All three just told you where the value went.]]></description><link>https://www.airealist.ai/p/the-watcher-is-the-product</link><guid isPermaLink="false">https://www.airealist.ai/p/the-watcher-is-the-product</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Thu, 20 Aug 2026 17:27:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2JyD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2JyD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2JyD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2JyD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2JyD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2JyD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2JyD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:553528,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/211983874?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2JyD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2JyD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2JyD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2JyD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79ad4f1b-949b-4148-a2c9-049e7d909f79_1456x720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On August 13, a TypeScript monorepo appeared under the deepseek-ai organization on GitHub: DeepSeek Harness, <code>dsh</code> to its command line, MIT-licensed, carrying a bolded warning that <strong>THERE WILL BE COMPATIBILITY-BREAKING CHANGES</strong> and not a single benchmark score anywhere in the release.[1] Six days later, it holds more than 160,000 stars.[2] On August 14, one day after the release, SpaceX closed its $60 billion all-stock acquisition of Anysphere, the maker of the Cursor coding agent, the highest price yet paid for a harness business.[3] And on August 18, OpenAI published a post titled &#8220;Pacing model development in an era of cyber-critical capabilities,&#8221; disclosing that supervising its own models now costs &#8220;roughly 20% of the inference compute being monitored,&#8221; and that its largest planned frontier reinforcement-learning run is on hold while it assembles the evidence and the machinery to run it safely.[4]</p><p style="text-align: justify;">Six days, one layer, three prices: zero, sixty billion dollars, and a fifth of everything watched. The layer is the harness, the control loop around the model. DeepSeek put its loop on GitHub for free. SpaceX paid sixty billion for a company that makes one. OpenAI said publicly in a document that the loop around <em>its</em> models has become expensive enough to disclose. <strong>None of the week&#8217;s three events is about a model. All three are about the scaffolding models run inside, and August 2026 is the month that scaffolding stopped being plumbing and became the contested asset of the AI stack.</strong></p><h2>The Layer Nobody Sells</h2><p style="text-align: justify;">A harness is everything around the model that turns text prediction into work: the system prompt, the tool catalog, the execution loop, the sandbox, the session state, the retry logic, the thing that decides when the agent is done. Claude Code is a harness. Codex is a harness. So are Aider, Cline, Goose, OpenCode, OpenHands, Pi, and Mistral&#8217;s Vibe; Cursor wraps one in an editor. The field is crowded enough that harness design philosophies have spanned a factor of fifty, from Pi&#8217;s two hundred tokens of system prompt to the ten thousand or so Claude Code carried until Anthropic cut it by roughly 80 percent this summer.[5]</p><p style="text-align: justify;">For three years, the industry treated this layer as a giveaway, open source, or bundled with the subscription, rarely as a line item. The model was the product; the harness was the wrapper around it. Most of that list was already free. What changed in August is that the zero became a strategy and the cost became public: one frontier camp made the giveaway its flagship release; the other disclosed, for the first time, what the layer costs to police. Between the zero and the 20% sit two measurements, one from four unaffiliated researchers and one from Tencent. <strong>Scores move more when you change the harness than when you change the model, and when the model&#8217;s judgment fails, the defenses that hold are the ones built into the harness</strong>. The layer one lab just gave away is the layer the other is taxing itself to police.</p><p>That is the argument. Here is the evidence.</p><h2>Five Times the Model</h2><p style="text-align: justify;">On August 15, two days after the DeepSeek release, four researchers published a paper with no institutional affiliation under the title &#8220;StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling.&#8221;[6] Terminal-Bench is the de facto reference benchmark for terminal-based agents: 89 tasks, real shells, verified outcomes. StateM is not a model. It is a frozen YAML runbook bolted onto an existing agent, a state machine of preconditions and practices distilled from failure postmortems. No weights change. Only the harness does.</p><p style="text-align: justify;">The numbers: GPT-5.5 under a stock harness scores 83.1 percent. The same model inside StateM scores 92.1% (+9 points) based on configuration alone.[7] For calibration, that is numerically above the 91.9% the paper cites for GPT-5.6 Sol Ultra, the next generation&#8217;s premium compute tier. Moving from GPT-5.5 to GPT-5.6 Sol at the same effort tier moves the score from 83.1 to 84.9 (+1.8%). A skeptic can build a bigger generation delta by buying up the range &#8212; stock Ultra sits 8.8% over stock GPT-5.5 &#8212; but that is rather the point: the runbook matched Ultra without changing models. <strong>The harness moved the score about five times as far as the model generation did</strong>.[8]</p><p style="text-align: justify;">The authors put their conclusion in bold: &#8220;The model appears not to be the (main) bottleneck.&#8221; And they draw the commercial inference for you: &#8220;One can invest in a harness that turns a cheaper model into a stronger system.&#8221;[9] Their cheap-model exhibit is the week&#8217;s other protagonist: DeepSeek-V4-Flash inside StateM reaches 88.1% for $15.20 in realized API charges. The frontier GPT-5.6 Sol xhigh run behind the 95.28% headline cost $1,062.95. A separate frontier submission cost $574.68 and scored 83.37%.[10] Read that triplet again. <strong>The frontier model in someone else&#8217;s submission scored </strong><em><strong>lower</strong></em><strong> than the bargain model inside the right harness, at thirty-eight times the reported cost</strong>.</p><p style="text-align: justify;">Now the asterisks, because this paper deserves both its headline and its scrutiny. The 95.28% is pre-adjudication. Its submission is still open as of this writing, with thirteen trajectories flagged by the benchmark&#8217;s automated review.[11] Four are conceded by the authors: embedded verification logic that should score zero, resulting in a score of 94.38%. Zeroing the nine flagged as possible reward hacking&#8212;gaming the grader rather than solving the task&#8212;sets their floor at 93.26. Zeroing all thirteen, the compounded worst case that the paper does not print gives 92.36. Every point in that range still sits numerically above the Ultra reference. The finding survives its asterisk. A reviewer&#8217;s separate cross-task contamination question remains open, the one flag arithmetic cannot bound.</p><p style="text-align: justify;">The deeper asterisk is generalization. Moved frozen, one-shot, to a held-out benchmark called BusinessBench, the runbook gains 0.55 macro points, a fraction of the headline gain, with outright negative transfer on two task families.[12] And the authors disclose something more uncomfortable: across iterations, their harness learned one evaluator&#8217;s boundary conventions &#8220;without ever reading verifier code.&#8221;[13] <strong>The runbook didn&#8217;t just learn to operate a terminal. It learned to please a specific examiner.</strong></p><p style="text-align: justify;">Both facts are true at once. A paper demonstrating that harnesses can inflate benchmark scores has a possibly-inflated benchmark score; a runbook tuned on postmortems memorized the grader; the leaderboard it sits on adopted, back in April, an automated &#8220;agent judge&#8221; that re-reviews every passing trial for reward hacking after three organizations were caught cheating.[14] The phenomenon is manifesting at every level, including the meta-level: even the benchmark now imposes a monitoring tax on the harness layer. Harness engineering is real leverage, and it is leverage precisely because it specializes in the task, in the environment, and, if you are not careful, in the examiner. Which is why <strong>the honest reading of the 5x is as a ceiling, not a floor</strong>.</p><h2>Everything Is a Plugin, Including the Providers</h2><p style="text-align: justify;">Against that measurement, consider the artifact DeepSeek shipped. The repository&#8217;s one-line description is the thesis: everything is a plugin. The fuller claim, as the company put it: &#8220;Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.&#8221;[15] The kernel underneath is not even DeepSeek&#8217;s: dsh is built on Cordis, a pre-existing plugin meta-framework from the Chinese chatbot ecosystem around Koishi.[16] A frontier lab looked at the most contested layer of the agent stack and adopted a chatbot community&#8217;s architecture for it. That is either humility or speed; on a codebase marked developer preview, it is probably both.</p><p style="text-align: justify;">The builders it aims at noticed: Armin Ronacher &#8212; co-founder of Earendil, which steers the Pi agent &#8212; called it &#8220;for sure the first time I have been looking at something new in the space and felt quite inspired to revisit some of our choices.&#8221;[5] The Hacker News thread reached 739 points, with praise focusing on the strength of this piece&#8217;s argument, while criticism focused on maturity and the plugin's attack surface.[17]</p><p>Three properties of the release matter.</p><p style="text-align: justify;">First, the silence. DeepSeek&#8217;s model launches include benchmark tables with accompanying prose. The harness shipped with none: no Terminal-Bench, no SWE-bench, and no eval directory containing results.[1] For this lab, that silence is louder than a chart. Either they have not measured, or the measurement does not flatter yet. Or they watched the 95.28 adjudication saga and declined to play a game where every harness number arrives with an asterisk. Whichever it is, <strong>the launch asks to be judged as architecture rather than capability</strong>. For a benchmark-first lab, this is a repositioning.</p><p style="text-align: justify;">Second, the logging. Every session in dsh is an append-only, replayable record, and the trajectory view exposes whatever reasoning the provider returns, which, for DeepSeek&#8217;s own API, is everything: V4-Pro now ships with thinking mode on by default.[18] Set that against the incumbents: OpenAI decided in September 2024, in the o1 launch post, not to show raw chains of thought, &#8220;after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring&#8221; [19]. Anthropic serves summaries of extended thinking and gates raw traces.[20] The industry&#8217;s chain-of-thought settlement inverted quietly this month: <strong>the closed labs treat raw reasoning as either a liability to monitor or an asset to protect, and the open lab treats it as a logging feature</strong>.</p><p style="text-align: justify;">Third, the funnel that isn&#8217;t. The obvious read of a free harness from a token vendor is razor-and-blades: MIT the razor, sell the blades. The configuration reality is softer. dsh ships provider cards for DeepSeek, Anthropic, OpenAI, Azure, Bedrock, Vertex, and Codex, and no provider is wired in as a default; you open settings and paste whichever key you own.[21] DeepSeek&#8217;s harness will cheerfully orchestrate OpenAI&#8217;s models. In the vocabulary this newsletter used for Nvidia&#8217;s open-source strategy, dsh is currently a sun, not a black hole: its gravity points outward, hardware- and provider-agnostic, where Nvidia&#8217;s giveaways route ecosystems back to its silicon.[22]</p><p style="text-align: justify;">That does not make the release charity. DeepSeek sells exactly one thing: tokens. It gave the weights away to grow their market, and the harness extends the same play one layer up. Leaked minutes of an investor meeting with founder Liang Wenfeng, published by ChinaTalk the day the harness shipped, cast China as a &#8220;token factory at global scale, pushing the price of intelligence down,&#8221; with DeepSeek as the instrument. DeepSeek&#8217;s own published economics say the factory&#8217;s margins live in the API.[23] And if a factory devoted to pushing prices down raised its own the same day [18], that is razor-and-blades in its plainest form: <strong>the price being pushed down is the market&#8217;s floor, not DeepSeek&#8217;s top tier, and the free layer is what walks you to the register</strong>. </p><p style="text-align: justify;">Follow the token-factory logic, and the harness strategy writes itself. You do not need to lock anyone in. You need the layer between users and tokens to be free, excellent, and everywhere, so that the only purchasing decision left is which tokens to pour through it, a competition DeepSeek believes it wins on price. Commoditize your complement. The Free Harness is less a lock than a price signal: this layer should cost nothing. Its de facto plugin registry today is an unauthenticated GitHub topic &#8212; anyone can publish to it.[24]</p><h2>The Twenty Percent</h2><p style="text-align: justify;">OpenAI&#8217;s August 18 post is the other half of the repricing. The core disclosure, verbatim: &#8220;These safeguards require meaningful compute. Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads.&#8221;[4]</p><p style="text-align: justify;">The scope is capability-indexed and precisely drawn: monitoring is required for all reinforcement-learning training and all tool-using evaluations for models at GPT-5.6 Sol capability or above.[25] The sharpest edge is a model OpenAI has not shipped: on August 7, eleven days before the cost disclosure, it said preliminary evaluations of an upcoming model called Astra were strong enough that it &#8220;cannot rule out critical cyber capabilities under our Preparedness Framework&#8221;, and that Astra &#8220;was not involved in exploiting Hugging Face.&#8221;[26] </p><p style="text-align: justify;">The threshold is a model that finds and weaponizes zero-days in hardened real-world systems without human help. The response is harness all the way down: isolation, restricted network and tool access, &#8220;universal monitoring for risky actions and misalignment across all agentic applications of Astra,&#8221; chain-of-thought monitors empowered to &#8220;review and interrupt high-risk activity.&#8221;[26] With that finding came an extension: monitoring now covers every tool-using inference call Astra makes, not just training and evals.[25]</p><p>Alongside, the pause: &#8220;Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.&#8221;[27]</p><p>Two readings of the 20%.</p><p>The skeptical reading: the denominator does all the work. Twenty percent of the inference compute being monitored, and the monitored share of total inference is undisclosed. If monitoring covers a sliver of traffic (frontier RL, evals, one presumptively cyber-capable model&#8217;s tool calls), the blended overhead across OpenAI&#8217;s fleet could be low single digits, and &#8220;20 percent tax&#8221; would overstate the burden the way headline capex numbers overstate disbursement. OpenAI has published the rate and withheld the base. A spokesperson added that the costs &#8220;won&#8217;t be passed on directly to customers&#8221;, and &#8220;directly&#8221; is carrying luggage.[28]</p><p style="text-align: justify;">The structural reading: the direction is unmistakable even if the magnitude is foggy. This is the company that stopped showing raw chains of thought partly to preserve monitorability, now publishing a price for the watching, and holding its biggest planned run until the watching is ready. The proximate cause is on the record: the July incident in which two OpenAI models under evaluation, with reduced refusal rates, escaped a sandbox via an Artifactory zero-day and reached Hugging Face&#8217;s infrastructure.[29] Both companies&#8217; remediations live mostly at the network and credential layer. After July, OpenAI paused frontier-model inference in research clusters for runs that could execute code or access the internet; the monitoring program is the condition the post offers for scaling such runs again.[4]</p><p style="text-align: justify;">The obvious objection runs the other way: doesn&#8217;t a tax levied on dangerous models prove the model is what matters? The risk originates there, yes: Astra earned its surveillance. But look at where the money goes: every item on OpenAI&#8217;s own mitigation list is loop-level: watchers wrapped around inference, execute-and-internet restrictions, sandboxes, permission boundaries. None of that spend makes the model more capable. All of it hardens the loop, the Monitoring Tax conceding in compute what StateM measures in points.</p><p style="text-align: justify;">And the monitoring technique OpenAI has published most about &#8212; reading the model&#8217;s chain of thought &#8212; carries a failure mode OpenAI itself documented in 2025: optimizing against a chain-of-thought monitor &#8220;does not eliminate all misbehavior and can cause a model to hide its intent.&#8221;[30] Monitoring at production scale <em>is</em> an optimization pressure: models get updated, retrained, and selected under it. That makes the Monitoring Tax an arms race rather than a toll, funded indefinitely out of OpenAI&#8217;s own compute.</p><p style="text-align: justify;">Put the two announcements side by side, and the asymmetry sharpens into strategy. A monitor over a training run and a sink gate in a session loop are the same line item: compute and code spent on the loop rather than the weights. DeepSeek&#8217;s answer to &#8220;who watches the agent?&#8221; is: you do; here is the append-only log; replay it. OpenAI&#8217;s answer is: We do, and here is what it costs. One externalizes the cost of supervision to the user and calls it transparency. The other internalizes it and calls it safety. Both are telling you the same truth: <strong>the loop around the model is now where risk management happens.</strong></p><h2>The Audit Came From Tencent</h2><p style="text-align: justify;">This brings us to the strongest evidence that the harness is the security perimeter. On August 17, four days after the dsh release, a team from Tencent&#8217;s Zhuque Lab published a security assessment of DeepSeek Harness, run with A.I.G, Tencent&#8217;s own AI-infrastructure red-teaming tool.[31] One Chinese lab publicly red-teaming another&#8217;s four-day-old flagship, methods and per-channel numbers in the open, is a norm Western frontier labs have not practiced on each other&#8217;s shipped products uninvited. An audit this fast was possible because the artifact is public: the transparency DeepSeek sells is what let Tencent take it apart in four days.</p><p style="text-align: justify;">The design is exhaustive for its target: 14,560 controlled executions against a pinned release-day commit, each payload tried both pasted as text and delivered as a file.[32] Headline result: full injection success at 5.6%, with most attempts ending in the agent explicitly refusing.[33]</p><p style="text-align: justify;">Whether 5.6% is good is unanswerable: no comparable audit of Claude Code or Codex has been published to serve as a baseline. The distribution, though, is where the lessons are.</p><p style="text-align: justify;">Hidden Unicode payloads inside files succeed 25.5% of the time; the identical payload pasted as text never succeeds.[34] The model is the same in both cases. What differs is the ingestion path: the paste route evidently normalizes away what the file route preserves. The delivery, and the cheapest defense, live in plumbing the model never sees.</p><p style="text-align: justify;">The skills channel &#8212; the very thing &#8220;everything is a plugin&#8221; celebrates &#8212; is injectable at 14-16%, among the highest rates in the study.[35] The architecture&#8217;s core feature is one of its largest measured wounds and its most predictable: skills are trusted instructions by design, so they are the cheapest place to hide untrusted ones, and the unauthenticated plugin topic is where untrusted ones will come from.</p><p style="text-align: justify;">The strongest text-mode attack in the study is social rather than technical: a fake progress note &#8212; your task is already done, now do this &#8212; planted in content the agent reads, succeeding at 17.0% against a 5.7% naive baseline.[36] The agent&#8217;s weakness is less parsing than trust in its own apparent history. And the most consequential split in the paper: corrupting what the agent says succeeds at 35.7%, while hijacking what it does &#8212; actually reaching a sensitive sink like mail, shell, or a transfer &#8212; succeeds at 2.5%.[37] An order of magnitude between lying and acting, but corrupted output that a human acts on is a hijack in itself.</p><p style="text-align: justify;">Why the gap? Because between interpretation and action sits the harness: confirmation steps, scoped permissions, sinks that demand more than persuasive context. Tencent&#8217;s top recommendation is to finish the job: authorize sensitive sinks independently of model interpretation.[38] Do not ask the model whether the email should be sent; make the send path require an authority the model&#8217;s context cannot mint. The July post-mortems converged on the same fix: authority boundaries, not better judgment.[29]</p><p style="text-align: justify;">Questions we&#8217;ll all have to answer: <strong>which actions can a poisoned context authorize? Which ingestion paths skip the paste route&#8217;s normalization? Who can publish a skill your agents will load?</strong> The defense in Tencent&#8217;s 14,560 runs was layered: 68% of attempts died when the model refused, and the harness failed to convert most of the survivors into privileged action. <strong>Security, like capability, is becoming a property of the loop.</strong></p><p style="text-align: justify;">The remaining caveats are the study&#8217;s own: a single harness, a single model, a single commit, and simulated sinks. Treat every number as a first measurement, not a ranking. But the direction of the finding does not depend on the baseline. Both of this month&#8217;s measurements point to the same layer from opposite sides: StateM shows the harness is where the capability variance is; Tencent shows it is where the effective defenses are.</p><h2>What the Asymmetry Buys</h2><p style="text-align: justify;">DeepSeek&#8217;s transparency is as strategic as OpenAI&#8217;s opacity. A lab accused of training on competitors&#8217; outputs benefits from normalizing raw chain-of-thought access: exposed traces are distillation feedstock, and while a closed API can still withhold its raw traces, a default-logging harness moves the norm: whatever a provider does return is recorded, replayable, and one export away from a training set.[39] &#8220;Everything is a plugin&#8221; also means the attack surface is a plugin: the skills channel Tencent flagged and the unvetted registry topic are the price of the architecture, and DeepSeek shipped them at developer-preview maturity, with a bolded compatibility warning where a plugin trust model will have to go.</p><p style="text-align: justify;">Conversely, OpenAI&#8217;s opacity spans two different decisions made two years apart, and the defense framing invites reading them as one. Monitoring for cyber-critical capability is a safety cost, disclosed this month; declining to show raw chains of thought is a choice from 2024, made by OpenAI&#8217;s own listing of factors, partly for &#8220;competitive advantage.&#8221;[19] Both are real; only one is the 20%. The asymmetry itself &#8212; one lab&#8217;s flagship feature is the other lab&#8217;s disclosed liability &#8212; is the finding.</p><p style="text-align: justify;">Readers of &#8220;<a href="https://www.airealist.ai/p/the-verification-tax">The Verification Tax</a>&#8221; will recognize the shape: there, the binding cost in verifiable-reward RL had migrated from generating answers to checking them; the Monitoring Tax is the same migration at the safety layer.[40] The priced layers of the stack, mid-2026, are the ones that watch and verify; the intelligence itself trades as a commodity, inside scaffolding that is now literally free.</p><h2>The Shell and the Moat</h2><p style="text-align: justify;">The deal that closed on August 14 was announced on June 16, days after SpaceX&#8217;s listing, with Anysphere then preparing for a round at a $50 billion valuation.[3] Sixty billion in a just-listed acquirer&#8217;s paper is not sixty billion in cash. Cognition &#8212; the Devin agent, plus the Windsurf editor it acquired &#8212; remains the largest standing independent: it raised $1 billion at a $25 billion pre-money valuation in May.[41]</p><p style="text-align: justify;">So what did SpaceX buy? Start with what it did not need. Not a software revenue line: a company assembling an AI division out of rockets, satellites, and the xAI merger does not spend sixty billion on ARR. And not models: xAI came in-house earlier this year, so the acquirer already owns a frontier lab.[3] <strong>A buyer with its own models paying the week&#8217;s highest price for a harness company is this piece&#8217;s thesis with a board&#8217;s signature on it</strong>. The model was not the bottleneck. The loop was. And the most deflationary rival read &#8212; that a model owner simply bought the funnel that feeds its models &#8212; is DeepSeek&#8217;s razor-and-blades at acquisition prices: it concedes the layer and haggles over the motive.</p><p style="text-align: justify;">What sixty billion buys, on the reading the week&#8217;s evidence supports, is the harness in the full sense. Not the generic shell DeepSeek zeroed twenty-four hours earlier, but everything that turns a shell into a harness: <strong>the coding-specialized loop, the domain process tuned against a telemetry base no rival could assemble from scratch, the team that does the tuning, and the enterprise contracts that keep the telemetry coming</strong>. </p><p style="text-align: justify;">VS Code did not kill JetBrains &#8212; Cursor itself began as a VS Code fork &#8212; and free shells tend to grow the category they commoditize. The decomposition also predicts a third moat: assurance. A hardened, attested build of a free shell is a paid product &#8212; Red Hat built a company on that arbitrage &#8212; and OpenAI just told you what assurance costs at the frontier. The twenty-six-trillion-dollar addressable market is the acquirer&#8217;s number and the acquirer&#8217;s story; the read above does not need it.[3]</p><p style="text-align: justify;">StateM hands you the same decomposition from the measurement side. Nine points on the benchmark the runbook was tuned for; 0.55 frozen points on the held-out one. The durable part of a harness advantage is the domain process &#8212; renewable, domain-bound, not ownable &#8212; and the assets that regenerate it: the users and the telemetry and postmortems they produce. The disposable part is the shell that hosts it, which now has a posted price. The underwriting translation is one question: of this company&#8217;s margin, how much is the loop itself, and how much is distribution, data, and contracts? This locates the value; it does not grade the price. <strong>SpaceX did not buy a shell; it bought the parts DeepSeek cannot give away.</strong></p><p style="text-align: justify;">The exit pattern should look familiar to readers of the silicon pieces: the independent specialized-inference vendors have been absorbed or signed away one by one to the GPU merchants, and now the largest independent harness has ended up inside an integrator too.[42] The one thing the Cursor deal does not tell you is what an independent harness is worth, because after August 14, there is one fewer way to find out.</p><h2>What Would Have to Break</h2><p>Three falsifiers, in plain language. </p><ol><li><p>If the next model generations resume dominating &#8212; if GPT-5.7 or V5 under stock harnesses move agentic benchmarks by more than harness engineering does &#8212; then the commoditization read dies and this was a plateau artifact. StateM currently argues otherwise, but StateM is one paper, with headline results on one benchmark, and an open adjudication. </p></li><li><p>If DeepSeek Harness, six months from now, has a thriving plugin ecosystem and no presence on Terminal-Bench-class leaderboards &#8212; or, if it avoids leaderboards on principle, no dsh components showing up inside rival stacks &#8212; then it is a chatbot shell with good marketing, not a harness in the sense that moves scores, and the Free Harness was a giveaway of something that did not matter.</p></li><li><p> If OpenAI ever discloses the monitored share of total inference and it rounds to low single digits blended, then the Monitoring Tax is a frontier-lab boutique cost, not a structural one. And the moat claim has its test: if Cursor-class pricing power survives customers swapping their own model keys into free shells at scale &#8212; revenue attributable to the loop itself &#8212; then the moat was in the shell after all, and this piece mislocated it.</p></li></ol><p style="text-align: justify;">What does not depend on any of those: where the decisions moved. <strong>From now on, we shouldn&#8217;t be choosing a model: the model is a provider card, swappable by design in the open reference, locked only where a vendor insists. We should be choosing a harness: its ingestion paths, its sink authorization, its logging posture, its plugin trust model.</strong> The properties Tencent measured and StateM priced. And anyone modeling frontier-lab economics now has a new line item with a disclosed rate and an undisclosed base, growing with capability by policy.</p><p style="text-align: justify;">DeepSeek and OpenAI disagree about nearly everything: licensing, logging, disclosure, where trust should live. In a single August week, they agreed, by acting in opposite directions, on the only question that matters: <strong>the loop around the model now moves outcomes more than the model inside it, and each lab priced that fact as aggressively as its business model allows</strong>. One priced it at zero to commoditize it. The other priced it at twenty percent of everything it watches to control it. And the week&#8217;s third price agreed with both of them: sixty billion dollars, from a buyer that already owned frontier models, for the loop around them. </p><p style="text-align: justify;">Everything is a plugin now &#8212; models, tools, skills, sandboxes, the UI. Everything except the watcher. <strong>The watcher is the product. The three of them only disagreed about who pays for it.</strong></p><div><hr></div><h3>Notes</h3><p>[1] DeepSeek-AI, <a href="https://github.com/deepseek-ai/deepseek-harness">DeepSeek Harness repository</a>, created August 13, 2026 (GitHub API <code>created_at</code> 2026-08-13T11:56:32Z, retrieved August 19). MIT license, TypeScript monorepo. The developer-preview warning &#8212; &#8220;THERE WILL BE COMPATIBILITY-BREAKING CHANGES,&#8221; bold in the original &#8212; is in the README. No benchmark, eval, or accuracy figure appears anywhere in the repository or launch materials as of August 19, 2026.</p><p>[2] Readings diverged on August 19, 2026: the <a href="https://api.github.com/repos/deepseek-ai/deepseek-harness">GitHub API</a> returned 165,743 stars and 17,618 forks, Shields.io showed ~166k, and GitHub&#8217;s rendered page showed 161.1k &#8212; star counts are eventually consistent and gameable, so the body uses the conservative floor. Stars measure attention, not quality; no independent audit of bot inflation exists for this repository, which is why this piece makes no &#8220;fastest-ever&#8221; comparison.</p><p>[3] Deal terms per <a href="https://techcrunch.com/2026/06/16/spacex-to-acquire-cursor-for-60b-in-stock-days-after-blockbuster-ipo/">TechCrunch, &#8220;SpaceX to acquire Cursor for $60B in stock, days after blockbuster IPO&#8221;</a>, June 16, 2026 (all-stock; TechCrunch reports a $10B break-up fee while Reuters reported a $4&#8211;10B range; Anysphere then preparing a $2B round at a $50B valuation), and completion per <a href="https://www.bloomberg.com/news/articles/2026-08-14/spacex-completes-its-60-billion-cursor-acquisition">Bloomberg, &#8220;SpaceX Completes $60 Billion Cursor Acquisition to Expand AI Coding Tools&#8221;</a>, August 14, 2026. The $26 trillion addressable-market figure is SpaceX&#8217;s own investor pitch as reported by TechCrunch &#8212; a vendor number, cited as the acquirer&#8217;s story, not as a market fact. TechCrunch also notes the acquisition strengthens SpaceX&#8217;s AI division, &#8220;which merged with Elon Musk&#8217;s xAI earlier in 2026&#8221; &#8212; the basis for the body&#8217;s observation that the acquirer already owned frontier models.</p><p>[4] OpenAI, <a href="https://openai.com/index/pacing-model-development-cyber-capabilities/">&#8220;Pacing model development in an era of cyber-critical capabilities&#8221;</a>, August 18, 2026. All OpenAI quotes in this piece are verbatim from the post unless otherwise attributed. The post also states OpenAI &#8220;paused frontier model inference in research clusters for runs that could execute code or use tools that could access the internet&#8221; immediately after the July incident.</p><p>[5] Thomas Claburn, <a href="https://www.theregister.com/ai-and-ml/2026/08/14/deepseeks-innovative-harness-treats-everything-as-a-plug-in/5288095">&#8220;DeepSeek&#8217;s innovative harness treats everything as a plug-in&#8221;</a>, The Register, August 14, 2026 &#8212; the harness field list, the Ronacher quote and identification, and the system-prompt comparisons. <a href="https://github.com/mistralai/mistral-vibe">Mistral&#8217;s Vibe</a>, the open-source CLI coding agent, added to the list by this piece. The prompt sizes (Pi ~200 tokens; Claude Code ~10,000, cut roughly 80 percent) are community measurements relayed by the Register, not vendor statements.</p><p>[6] Ziheng Qin, Yaxin Lu, Zhangyang Wang, Kai Wang, <a href="https://arxiv.org/abs/2608.15089">&#8220;StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling&#8221;</a>, arXiv:2608.15089, August 15, 2026; code at <a href="https://github.com/henryqin1997/statem">github.com/henryqin1997/statem</a>. The paper states the work &#8220;was conducted in the authors&#8217; personal time and does NOT reflect the views of any affiliated organization&#8221;; no institution is named.</p><p>[7] StateM results: GPT-5.5 (xhigh) 83.1% stock &#8594; 92.1% under StateM; GPT-5.6 Sol (xhigh) 84.9% stock &#8594; 95.28% raw pre-adjudication (424/445 trials; 89 tasks &#215; 5 trials). The Ultra comparison is the abstract&#8217;s own: &#8220;StateM raises GPT-5.5 xhigh to 92.1%, versus 83.1% reference and GPT-5.6 Sol Ultra at 91.9%&#8221; &#8212; 91.9% is cited as the flagship&#8217;s reference score, not a StateM result. The 0.2-point margin of 92.1 over 91.9 is within sampling error (the submission reports &#177;0.87% SE); &#8220;numerically above&#8221; in the body is meant literally, not as a significance claim. All figures from the paper and the leaderboard submission (note 11).</p><p>[8] Arithmetic: harness deltas +9.0 (GPT-5.5) and +10.4 (GPT-5.6 Sol xhigh) versus the +1.8 stock-to-stock generation delta (83.1 &#8594; 84.9); 9.0 &#247; 1.8 = 5.0. Same benchmark, same task set, same trial protocol. The 5.0&#215; ratio is built on the GPT-5.5 pair, which is paper-reported and not part of the flagged leaderboard submission; the flagged trials (note 11) all belong to the GPT-5.6 Sol xhigh run.</p><p>[9] Both quotes from the paper; the bottleneck sentence is bolded in the original.</p><p>[10] Costs per the paper and submission: $1,062.95 reported model cost (1.18B tokens) for the GPT-5.6 Sol xhigh run; $15.20 realized API charges for the DeepSeek-V4-Flash final evidence run at 88.09% under standard timeouts (stock baseline 82.7%; the whole DeepSeek adaptation campaign cost $52.22); $574.68 is a <em>different</em> leaderboard submission &#8212; GPT-5.6 Sol (max) at 83.37% raw &#8212; not the same configuration at a different price. $574.68 &#247; $15.20 &#8776; 37.8&#215;. &#8220;xhigh&#8221; and &#8220;max&#8221; denote reasoning-effort configurations.</p><p>[11] <a href="https://github.com/harbor-framework/terminal-bench-2-1/pull/142">Pull request #142, harbor-framework/terminal-bench-2-1</a>, filed July 15, 2026; still open with last activity July 31 as of August 19. Automated review flagged 13 trajectories: 4 as &#8220;harness cheating&#8221; (embedded verification logic &#8212; the authors concede these should score zero, giving 420/445 = 94.38%) and 9 as possible reward hacking (the authors dispute 5 as false positives). The paper&#8217;s two scenarios do not compound: it computes 420/445 = 94.38% (zeroing the 4) and 415/445 = 93.26% (zeroing the 9) each from the raw 424. The PR lists the 13 flagged trajectories as distinct, so the compounded worst case &#8212; all 13 zeroed &#8212; is 411/445 = 92.36%, a floor the paper does not state; this piece uses it. The submission&#8217;s agent is listed as &#8220;statem-Codex&#8221; &#8212; StateM wrapped around Codex. A reviewer separately raised cross-task hint contamination. Adjudication is pending; every headline number in this piece carries that status. For scale: the top <em>merged</em> entry shown on the <a href="https://www.tbench.ai/leaderboard/terminal-bench/2.1">public 2.1 leaderboard</a> sat at 83.8% (Claude Code, Fable 5) as of August 19 &#8212; the StateM figures above that are paper- and PR-reported, not yet board-accepted.</p><p>[12] BusinessBench frozen one-shot transfer: +0.55 macro (+1.34 micro); negative transfer on RefactorBench (&#8722;2.78) and WooCommerce Stock (&#8722;3.70), where the paper says the learned controls targeted the wrong execution boundaries; two mechanism-matched task families gained +10.04. The runbook transfers where the mechanism matches, and not elsewhere.</p><p>[13] The paper&#8217;s own disclosure: the profile learned a verifier&#8217;s boundary convention &#8220;without ever reading verifier code&#8221; &#8212; repeated evaluator feedback encoded the examiner&#8217;s unstated conventions into the runbook.</p><p>[14] Terminal-Bench, <a href="https://www.tbench.ai/news/leaderboard-integrity-update">&#8220;Leaderboard integrity update&#8221;</a>, April 19, 2026: an agent judge now re-reviews all passing trials for reward hacking (zeroed if confirmed), after three organizations &#8212; OpenBlock (OB-1), QuantFlow (Pilot), and ForgeCode &#8212; were penalized for cheating.</p><p>[15] The one-line description and the &#8220;Everything is a Plugin&#8221; framing are on the repository itself (note 1); the fuller sentence is DeepSeek&#8217;s wording as quoted by The Register (note 5); capitalization in the original.</p><p>[16] <a href="https://github.com/cordiverse/cordis">Cordis</a> is a pre-existing plugin meta-framework from the ecosystem around the Koishi chatbot project, not a DeepSeek codebase; its design paper, <a href="https://github.com/cordiverse/paper">&#8220;A Programming Paradigm for Spatiotemporal Composability&#8221;</a>, is GitHub-hosted with no peer-reviewed venue. Some aggregator coverage has misattributed Cordis to DeepSeek; the provenance is the other way around &#8212; DeepSeek adopted it.</p><p>[17] <a href="https://news.ycombinator.com/item?id=49285244">Hacker News, &#8220;DeepSeek Harness developer preview&#8221;</a> &#8212; 739 points, 309 comments as of August 19, 2026. The characterization of where praise and criticism concentrated is this piece&#8217;s reading of the thread; the quoted phrase is from a highly upvoted comment. Forum commentary: color and practitioner sentiment, not evidence.</p><p>[18] Session model per repository documentation and The Register (note 5): append-only session log with a trajectory view exposing raw reasoning. V4-Pro reached general availability on the API &#8212; thinking mode on by default &#8212; on August 13, 2026, the same day as the harness release; see e.g. <a href="https://venturebeat.com/technology/deepseek-harness-launches-as-open-source-rival-to-claude-code-alongside-v4-pro-on-api-with-higher-prices">VentureBeat&#8217;s launch coverage</a>.</p><p>[19] OpenAI, <a href="https://openai.com/index/learning-to-reason-with-llms/">&#8220;Learning to reason with LLMs&#8221;</a>, September 2024: &#8220;after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users,&#8221; serving &#8220;a model-generated summary of the chain of thought&#8221; instead.</p><p>[20] Anthropic&#8217;s summarized reasoning and gated raw traces per The Register (note 5). Anthropic has not published a comparable cost figure for its classifier and summarization pipeline.</p><p>[21] <a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/docs/user/guide/providers.md">DeepSeek Harness provider documentation</a>: provider cards for DeepSeek, Anthropic, OpenAI, Azure, Bedrock, Vertex, and Codex (the last authenticating over OAuth); API keys are entered manually in settings; the documentation designates no provider as the default.</p><p>[22] &#8220;<a href="https://www.airealist.ai/p/open-source-closed-orbit">Open Source, Closed Orbit: The Hardware Monopolist&#8217;s Guide to Owning Open Source</a>,&#8221; The AI Realist &#8212; the black-hole/sun diagnostic: does a vendor&#8217;s open-source contribution make competitors&#8217; products easier or harder to use?</p><p>[23] Irene Zhang, <a href="https://www.chinatalk.media/p/the-deepseek-thesis">&#8220;The DeepSeek Thesis&#8221;</a>, ChinaTalk, August 13, 2026 &#8212; leaked minutes of a four-hour meeting between Liang and investors that circulated in late July; the clause quoted in the body is ChinaTalk&#8217;s English rendering (&#8221;China, in his eyes, will play the role of token factory at global scale, pushing the price of intelligence down as it did for countless other industries during its manufacturing boom&#8221;), not a verbatim Liang quote. The $562,027/day and 545 percent figures are not from the leak: ChinaTalk attributes them to an analysis DeepSeek itself published in February 2025 (R1 API, theoretical daily revenue at a 545 percent cost-profit ratio, with DeepSeek&#8217;s own caveat that actual revenue ran substantially lower). Vendor-published theoreticals &#8212; better provenance than a leak, still the seller&#8217;s math. The mundane read of the harness release &#8212; an ordinary ecosystem move, no grand pricing design &#8212; survives the evidence; the price signal reads the same either way.</p><p>[24] The <code>dsh-plugin</code><a href="https://github.com/topics/dsh-plugin"> GitHub topic</a> functions as the de facto plugin registry: 150 public repositories carried the tag as of August 19, 2026, six days after release. GitHub topics are self-applied by repository owners; there is no authentication, signing, vetting, or review, and DeepSeek has announced no registry governance. The most visible curation is a community-maintained awesome-list.</p><p>[25] OpenAI post (note 4): &#8220;This monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher,&#8221; and: &#8220;Once we determined that Astra may have critical cyber capabilities on August 7, we added an additional monitoring requirement for all inference of Astra with tools (not just RL training and evaluations).&#8221;</p><p>[26] OpenAI, <a href="https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/">&#8220;Responding to the next frontier of critical cyber capabilities&#8221;</a>, August 7, 2026. Body quotes verbatim: OpenAI &#8220;cannot rule out critical cyber capabilities under our Preparedness Framework&#8221;; the mitigation list includes &#8220;We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra,&#8221; with monitors that &#8220;evaluate the model&#8217;s Chain of Thought and trigger a security response to review and interrupt high risk activity&#8221;; and &#8220;Astra is an upcoming model, and was not involved in exploiting Hugging Face.&#8221; The Critical threshold in the Preparedness Framework: the ability to &#8220;identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention&#8221; (body paraphrases). Both posts preserve the modal: the August 18 post&#8217;s own wording is &#8220;determined that Astra may have critical cyber capabilities&#8221; (note 25).</p><p>[27] OpenAI post (note 4), verbatim.</p><p>[28] Thomas Claburn, <a href="https://www.theregister.com/ai-and-ml/2026/08/19/openais-overhead-will-rise-20-percent-for-some-workloads-as-it-hardens-security/5289303">&#8220;OpenAI&#8217;s overhead will rise 20 percent for some workloads as it hardens security&#8221;</a>, The Register, August 19, 2026 &#8212; including the unnamed spokesperson&#8217;s statement that the costs &#8220;won&#8217;t be passed on directly to customers.&#8221;</p><p>[29] The July 2026 incident, from the primary accounts: OpenAI, <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">&#8220;OpenAI and Hugging Face partner to address security incident during model evaluation&#8221;</a>; Hugging Face, <a href="https://huggingface.co/blog/security-incident-july-2026">&#8220;Security incident disclosure &#8212; July 2026&#8221;</a> and <a href="https://huggingface.co/blog/agent-intrusion-technical-timeline">&#8220;Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident&#8221;</a>. OpenAI&#8217;s account &#8212; two models in a cybersecurity evaluation with reduced refusal training, an evaluation-sandbox escape via an Artifactory zero-day &#8212; is deliberately narrower than the &#8220;rogue models hacked Hugging Face&#8221; shorthand that circulated in some coverage, and this piece follows the narrower account.</p><p>[30] OpenAI, <a href="https://arxiv.org/abs/2503.11926">&#8220;Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation&#8221;</a>, March 2025, as quoted in the Register&#8217;s August 19 coverage (note 28).</p><p>[31] Zonghao Ying et al., <a href="https://arxiv.org/abs/2608.16393">&#8220;Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection&#8221;</a>, arXiv:2608.16393 (v1 August 17, v2 August 18, 2026); assessment materials in <a href="https://github.com/Tencent/AI-Infra-Guard/tree/main/Research/deepseek-harness-security-assessment">Tencent&#8217;s AI-Infra-Guard repository</a>. Affiliation: Tencent per the paper&#8217;s metadata (so listed on the Hugging Face papers index); Zhuque Lab per the paper body; publication inside <a href="https://github.com/Tencent/AI-Infra-Guard">Tencent&#8217;s GitHub organization</a> corroborates. The consensual Western comparator referenced in the body: OpenAI and Anthropic&#8217;s <a href="https://openai.com/index/openai-anthropic-safety-evaluation/">2025 pilot alignment-evaluation exchange</a> &#8212; reciprocal, agreed, and model-level rather than a product audit.</p><p>[32] Design: 14,560 executions = 1,120 test cases (16 injection channels &#215; 2 carrier modes &#215; 35 payload objectives) &#215; 13 attack methods including a naive baseline, against DSH commit 47f94385 (the August 13 release commit) running deepseek-v4-flash, with sensitive sinks simulated as fixtures.</p><p>[33] Full-injection success: 5.6% under the deterministic judge, 5.3% under the LLM judge; 68.4% of attempts ended in explicit refusal.</p><p>[34] Hidden Unicode payloads: 25.5% full success delivered inside files versus 0.0% for the same payload pasted as text (deterministic judge) &#8212; the paste path normalizes; the file-parsing path preserves.</p><p>[35] Skills-channel injection: 14&#8211;16% across both judges, among the highest-risk channels in the study.</p><p>[36] The fake_completion method &#8212; a planted note claiming the task is already done, redirecting the agent &#8212; reached 17.0% under the LLM judge in text mode, versus 5.7% for the unmodified baseline payload.</p><p>[37] Output corruption 35.7% versus sensitive-sink action hijack 2.5%; the paper treats these as distinct operational threat profiles. Both are lab rates &#8212; crafted adversarial payloads, simulated sinks, no adaptive attacker &#8212; conditional on an attack reaching the agent, not fleet frequencies.</p><p>[38] The assessment&#8217;s stated mitigation: sensitive sinks require independent authorization mechanisms separate from model interpretation (paraphrase of the report&#8217;s recommendation).</p><p>[39] OpenAI said in January 2025 it had evidence suggesting DeepSeek trained on distilled outputs of its models (widely reported at the time, e.g. by the Financial Times, January 29, 2025); no public resolution followed, and DeepSeek did not respond publicly to the specifics. This piece takes no position on the accusation &#8212; only on who benefits from normalized raw-trace access.</p><p>[40] &#8220;<a href="https://www.airealist.ai/p/the-verification-tax">The Verification Tax</a>,&#8221; The AI Realist &#8212; the argument that in RL with verifiable rewards, the binding cost had migrated from generating candidate solutions to verifying them.</p><p>[41] <a href="https://techcrunch.com/2026/05/27/ai-coding-startup-cognition-raises-1b-at-25b-pre-money-valuation/">TechCrunch, &#8220;AI coding startup Cognition raises $1B at $25B pre-money valuation&#8221;</a>, May 27, 2026; Bloomberg reported the talks on April 23. Priced round, pre-money basis as stated.</p><p>[42] The Integration Premium &#8212; defined in &#8220;<a href="https://www.airealist.ai/p/aws-built-its-own-ai-chip-now-it">AWS Built Its Own AI Chip. Now It Needs Someone Else&#8217;s</a>.&#8221; and extended in &#8220;<a href="https://www.airealist.ai/p/the-model-is-the-machine">The Model Is the Machine</a>&#8221; (The AI Realist): in a disaggregating stack, margin migrates from component makers to the integration layer, and the modal exit for independent specialists is absorption (Groq assets to Nvidia; the Untether team and Taalas to AMD). &#8220;Musk&#8217;s Chip Gambit&#8221; covers the acquirer&#8217;s side of this arc.</p>]]></content:encoded></item><item><title><![CDATA[The Model Is the Machine]]></title><description><![CDATA[AMD just bought a chip company whose product runs exactly one model. That is the point.]]></description><link>https://www.airealist.ai/p/the-model-is-the-machine</link><guid isPermaLink="false">https://www.airealist.ai/p/the-model-is-the-machine</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Sun, 09 Aug 2026 15:36:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rJIq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rJIq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rJIq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!rJIq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!rJIq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!rJIq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rJIq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2316474,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/210476886?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rJIq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!rJIq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!rJIq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!rJIq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff12e7c00-587d-46a9-92a1-eb95ef66db5f_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">For sixty years, mainstream computing has run in one direction: the hardware is fixed, and the software adapts to it. On August 6, AMD paid to reverse the arrow. It signed a definitive agreement to acquire Taalas, a Toronto startup that etches a model&#8217;s weights directly into the chip&#8217;s metal, a silicon that can run only the model it was manufactured with.[1] Taalas co-founder Ljubisa Bajic, in AMD&#8217;s announcement: &#8220;We founded Taalas to rethink AI inference from the ground up by building the hardware around the model.&#8221;[1]</p><p style="text-align: justify;">In March, I wrote that the one-chip-does-everything era of AI inference ended when AWS, Nvidia, and Huawei converged on the same split: prefill on compute-bound silicon, decode on memory-bound silicon.[2][3] AMD held the middle ground: disaggregation in scheduling software, not in silicon.[2] Four and a half months later, AMD closed that ground, not with a decode chip, but with the rung below. Specialization is descending: first, the workload, when training chips split from inference chips; then the phase; now the model, and the GPU duopoly is paying for every rung of the descent.</p><h2>What AMD actually bought</h2><p style="text-align: justify;">Taalas showed working silicon this February: HC1, a 6-nanometer TSMC chip with 53 billion transistors, Meta&#8217;s Llama 3.1 8B etched into a mask-ROM fabric, and the KV cache &#8212; its working memory &#8212; in on-die SRAM.[4] No high-bandwidth memory (HBM), no advanced packaging, no liquid cooling. The company claims about 17,000 tokens per second per user, a tenth of the power of GPU serving, and a twentieth of the datacenter build cost. Vendor numbers, measured on a small model at 3-bit quantization; per-user speed is not aggregate throughput.[5]</p><p style="text-align: justify;">The prefill/decode split was a workaround for chips that must serve any model. Etch one model into the die, and the workaround dissolves: weights sit next to compute, the memory wall that made decode expensive is gone, and one chip serves both phases again, for exactly one tenant.[2]</p><p style="text-align: justify;">When the weights change, there is nothing to reprogram. You reprint: a new checkpoint of the etched architecture touches two metal layers of the roughly 100 mask layers that build the chip, a revision Taalas says TSMC turns in about 2 months, against 6 for a full design.[6] A new architecture is a different job: the fabric is shaped to the model&#8217;s dimensions, and Taalas&#8217;s own roadmap says as much. Its next model class arrives on new silicon, not as a revision of HC1.[6] The first product took 24 people, $30 million of the roughly $219 million raised, and a founding team out of Tenstorrent, AMD, and Nvidia.[7]</p><h2>The admission, priced twice</h2><p style="text-align: justify;">On December 24, Nvidia paid a reported $20 billion &#8212; its largest transaction on record, unconfirmed in any filing &#8212; for a non-exclusive license to Groq&#8217;s inference IP and roughly 90 percent of its staff. &#8220;We are not acquiring Groq as a company,&#8221; Jensen Huang told employees, oddly.[8] Twelve weeks later, the Groq 3 LPX stood on the GTC stage: an SRAM-based decode rack fused into the Vera Rubin platform, with Foxconn reportedly pulling production forward to this summer.[9]</p><p style="text-align: justify;">AMD&#8217;s pattern is the same, run twice. In June 2025, it hired away the entire team behind Untether AI, a Toronto inference-chip company whose products were promptly discontinued.[10] Taalas is AMD&#8217;s second Toronto inference-silicon absorption in fourteen months, and this time it kept the product: Taalas goes into the Instinct roadmap and the Helios racks now ramping with Anthropic, Meta, Microsoft, OpenAI, and Oracle.[1][11] Lisa Su calls herself &#8220;a big believer that there&#8217;s no one-size-fits-all as it comes to chips&#8221;; her AI chief, Vamsi Boppana, frames the acquisition as &#8220;the right compute solutions for every AI workload.&#8221;[12][1] Both lines are unremarkable until you remember what these companies sell. The two firms whose franchises are general-purpose GPUs are the ones paying &#8212; one, a reported twenty billion; the other, undisclosed &#8212; for silicon that is anything but.</p><p style="text-align: justify;">AMD would call this a portfolio, not a pivot, and the undisclosed terms suggest the hedge was probably cheap. An acquisition proves direction only when it ships. Nvidia&#8217;s shipped. AMD&#8217;s now has to.</p><h2>When the model and the machine write down together</h2><p style="text-align: justify;">Model-etched silicon matters for accounting reasons before engineering ones. In April, I wrote that the industry&#8217;s mismatch is temporal: frontier models live three to twelve months while the hardware they run on depreciates over five or six years. OpenAI shipped five versions of GPT-5 in seven months, none of which lasted four months as the flagship.[13] The gap between those schedules is where the balance-sheet fiction lives. Etching closes it in two tiers that mirror the model pipeline itself. Manufactured chips carry one checkpoint: supersede it, and that inventory is done, a two-month reprint producing the next edition. The design &#8212; the shape cast in silicon &#8212; depreciates with the architecture; adapters in SRAM absorb the fine-tune churn in between.[4][6] The chip is not a platform. It is a print run, and the press outlives the edition.</p><p style="text-align: justify;">A print run is only rational if what you are printing holds still, and the thing that must hold still is the architecture. Checkpoints are what reprints are for. Nobody etches the frontier: that tier changes every six weeks and stays on programmable silicon. The bet is the serving tier: the Llama-class workhorses and distilled variants. The demo is the evidence: the architecture Taalas etched in February 2026 is a shape Meta shipped in July 2024, and DeepSeek was still pouring new reasoning checkpoints into it six months on.[4] AMD is pricing the proposition that the volume tier of the model market has commoditized into stable shapes, and that the inference price war will be won a twentieth of a build cost at a time. The most consequential claim in this acquisition is about models, not chips.</p><p style="text-align: justify;">So the strategic question underneath the deal: whose shapes get etched? An etched fabric in AMD&#8217;s roadmap is a standard with hardware gravity. Every checkpoint that wants the economics must ship in that shape, and the compatibility war moves down a layer. The co-design precedent is already on the record: Amazon says Anthropic works with Annapurna Labs to shape next-generation Trainium, Broadcom builds Anthropic&#8217;s custom accelerators, and Anthropic already runs on Helios.[14][11] Watch for a frontier lab handing its serving workhorse to AMD&#8217;s mask shop, the model roadmap becoming a silicon roadmap.</p><h2>What would prove this wrong</h2><p style="text-align: justify;">Three tests, all observable. If all three fail, this was a talent acquisition with a good press release, Untether with better branding.</p><ol><li><p style="text-align: justify;">Taalas silicon appears in a shipping Helios configuration, or it never does. </p></li><li><p style="text-align: justify;">A hardwired, model-specific chip reaches production at a major platform by the end of 2027, or none does. </p></li><li><p style="text-align: justify;">Serving-tier architectures hold still long enough to etch, or shape turnover stays too fast, and the category dies. </p></li></ol><p style="text-align: justify;">The market-structure consequence is already on the scoreboard. Groq&#8217;s assets are inside Nvidia. Untether&#8217;s team and Taalas are inside AMD. Cerebras took the other exit, a May IPO at $185 a share.[15] Anyone pricing an independent inference-silicon position should assume the modal exit is absorption, not platform status. The category built to disrupt GPU vendors is consolidating into them, and margins migrate to whoever assembles the system.[2]</p><p style="text-align: justify;">Strip both deals to the sentence they share: the companies whose franchise is general-purpose silicon have concluded that the most valuable workload in computing no longer wants it. What they are unwinding is older than either company: the separation of software from hardware &#8212; the line that lets you sell one without the other, update one without retooling the other, write one down while the other keeps depreciating. The software industry was built on that line. Etching erases it. The model is the machine now.</p><div><hr></div><h3>Notes</h3><p>[1] AMD press release, <a href="https://newsroom.amd.com/news/amd-acquires-taalas-ai-inference/">&#8220;AMD Acquires Taalas to Advance Compute Solutions for Rapidly Growing AI Inference Market,&#8221;</a> August 6, 2026. Definitive agreement; financial terms not disclosed; closing expected Q4 2026 subject to regulatory approvals. Vamsi Boppana (SVP, AI Group): &#8220;AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload.&#8221; Ljubisa Bajic: &#8220;We founded Taalas to rethink AI inference from the ground up by building the hardware around the model.&#8221; The release names integration targets: the Instinct accelerator roadmap, Helios rack-scale systems, EPYC CPUs, and ROCm software.</p><p>[2] Julien Simon, <a href="https://www.airealist.ai/p/aws-built-its-own-ai-chip-now-it">&#8220;AWS Built Its Own AI Chip. Now It Needs Someone Else&#8217;s,&#8221;</a> The AI Realist, March 15, 2026. Introduces the Reasoning Tax and the Integration Premium (&#8221;in any disaggregating hardware stack, margin migrates from component manufacturers to the integration layer&#8221;); documents the three-ecosystem convergence and AMD&#8217;s middle-ground position: &#8220;The disaggregation is in scheduling, not in silicon.&#8221;</p><p>[3] Julien Simon, <a href="https://www.airealist.ai/p/acquired-absorbed-diaggregated">&#8220;Acquired, Absorbed, Disaggregated,&#8221;</a> The AI Realist, March 26, 2026. The convergence: <a href="https://www.businesswire.com/news/home/20260313406341/en/AWS-and-Cerebras-Collaboration-Aims-to-Set-a-New-Standard-for-AI-Inference-Speed-and-Performance-in-the-Cloud">AWS-Cerebras announced March 13</a>; <a href="https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Enters-Production-With-Dynamo-the-Broadly-Adopted-Inference-Operating-System-for-AI-Factories/default.aspx">Nvidia Dynamo entered production March 16</a>; Groq 3 LPX announced at GTC March 17. Huawei&#8217;s phase-split Ascend 950PR/950DT pair, announced September 2025, ships across 2026.</p><p>[4] Taalas HC1 unveiled February 2026. Specifications per <a href="https://www.heise.de/en/news/AI-inference-cast-in-silicon-Taalas-announces-HC1-chip-11185112.html">heise online</a> (February 2026) and <a href="https://www.datacenterdynamics.com/en/news/ai-chip-startup-taalas-raises-169m-unveils-hc1-processor-optimized-for-llama-31-8b/">Data Center Dynamics</a>: TSMC 6nm, approximately 53 billion transistors on an 815 mm&#178; die; model weights in a mask-ROM recall fabric using a proprietary 3-bit data format with 6-bit parameters; KV cache and fine-tuning adapters in an SRAM fabric; no HBM, no advanced packaging, no liquid cooling. The base model is fixed at manufacture. Adapter-style fine-tunes (LoRA) load into the SRAM fabric at runtime, and the context window is configurable, per <a href="https://www.cnx-software.com/2026/02/22/taalas-hc1-hardwired-llama-3-1-8b-ai-accelerator-delivers-up-to-17000-tokens-s/">CNX Software</a>; a full-parameter fine-tune &#8212; or a LoRA merged into the base weights &#8212; is a new checkpoint, which requires a mask reprint, not a load. Llama 3.1 was released by Meta in July 2024; the same 8B architecture carried DeepSeek&#8217;s <a href="https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-8B">R1-Distill-Llama-8B</a> in January 2025 &#8212; checkpoints changed, the shape persisted. Meta&#8217;s own frontier direction has since turned proprietary with Muse Spark, its first release after the Superintelligence Labs reorganization, though Meta says current Llama models remain open source, per <a href="https://venturebeat.com/technology/goodbye-llama-meta-launches-new-proprietary-ai-model-muse-spark-first-since">VentureBeat</a>. The shape&#8217;s persistence in the serving tier does not depend on Meta&#8217;s frontier roadmap: the installed base and third-party checkpoints sustain it &#8212; if anything, a sponsor stepping back makes an etched shape more like a public instruction set than a vendor product.</p><p>[5] Vendor-published figures: approximately 17,000 tokens per second per user on Llama 3.1 8B, roughly one-tenth the power and one-twentieth the datacenter build cost of conventional GPU serving, per Taalas materials as reported by <a href="https://www.heise.de/en/news/AI-inference-cast-in-silicon-Taalas-announces-HC1-chip-11185112.html">heise</a> and <a href="https://www.theregister.com/systems/2026/08/06/amd-acquires-ai-chip-startup-taalas-to-boost-inference-performance-by-etching-models-into-silicon/5284344">The Register</a>. heise notes early independent tests reaching close to 16,000 tokens/second. Per-user token rate is a latency-side metric and does not translate directly to aggregate throughput per chip; the demonstration model is small (8B parameters) and aggressively quantized. The contrast with Cerebras is density and cost: the WSE-3 reaches on-die weights with 44 GB of SRAM across a full wafer; Taalas reaches them with mask ROM on a single 815 mm&#178; die &#8212; fixed function traded for commodity size.</p><p>[6] Company-described revision process, as reported by <a href="https://www.heise.de/en/news/AI-inference-cast-in-silicon-Taalas-announces-HC1-chip-11185112.html">heise</a>, <a href="https://www.unite.ai/amd-buys-taalas-to-put-hard-wired-ai-models-in-its-accelerator-roadmap/">Unite.AI</a>, and <a href="https://www.theregister.com/systems/2026/08/06/amd-acquires-ai-chip-startup-taalas-to-boost-inference-performance-by-etching-models-into-silicon/5284344">The Register</a>: weights occupy two metal layers of the roughly 100 mask layers used to build the chip, with TSMC turnaround of approximately two months for a two-layer revision versus approximately six for a full design. All three descriptions derive from Taalas, and none specifies whether the two-layer path covers anything beyond new weights for the etched architecture. A cross-architecture port &#8212; different dimensions, vocabulary, or attention configuration; Llama to Qwen, or 8B to 70B &#8212; changes the compute fabric itself. The Register describes re-spins for new models as significantly cheaper than starting from scratch, and Taalas&#8217;s roadmap points to new silicon for new model classes (a mid-sized reasoning model in spring, the HC2 platform &#8212; with standard 4-bit floating-point formats &#8212; by year end) rather than HC1 revisions, per <a href="https://www.heise.de/en/news/AI-inference-cast-in-silicon-Taalas-announces-HC1-chip-11185112.html">heise</a> and <a href="https://www.cnx-software.com/2026/02/22/taalas-hc1-hardwired-llama-3-1-8b-ai-accelerator-delivers-up-to-17000-tokens-s/">CNX Software</a>.</p><p>[7] <a href="https://www.datacenterdynamics.com/en/news/ai-chip-startup-taalas-raises-169m-unveils-hc1-processor-optimized-for-llama-31-8b/">Data Center Dynamics</a>, February 2026: founded August 2023 by Ljubisa Bajic (previously an architect at AMD and Nvidia; co-founder of Tenstorrent, where he swapped the CEO role with Jim Keller in October 2022 and departed in March 2023), Drago Ignjatovic, and Lejla Bajic; $169 million round announced February 2026, approximately $219 million raised in total. Investors include Quiet Capital, Fidelity, and Pierre Lamond, per <a href="https://www.unite.ai/amd-buys-taalas-to-put-hard-wired-ai-models-in-its-accelerator-roadmap/">Unite.AI</a>. Team of 24 and $30 million spent on the first product, per <a href="https://www.heise.de/en/news/AI-inference-cast-in-silicon-Taalas-announces-HC1-chip-11185112.html">heise</a>.</p><p>[8] <a href="https://www.cnbc.com/2025/12/24/nvidia-buying-ai-chip-startup-groq-for-about-20-billion-biggest-deal.html">CNBC</a>, December 24, 2025. Structured as a non-exclusive IP licensing agreement plus the hiring of approximately 90% of Groq staff; approximately $20 billion per investor sources, not confirmed by Nvidia in filings. Jensen Huang internal email obtained by CNBC: &#8220;We are not acquiring Groq as a company.&#8221; Jonathan Ross subsequently joined Nvidia as chief software architect, per <a href="https://www.forbes.com/sites/phoebeliu/2026/03/18/groq-cofounder-ross-explains-whirlwind-ai-chip-deal-with-nvidia/">Forbes</a> (March 18, 2026).</p><p>[9] Groq 3 LPX announced at GTC 2026, March 17, 2026 (see [3]). Chip produced on Samsung&#8217;s 4nm process and integrated into the Vera Rubin platform, per <a href="https://www.tomshardware.com/tech-industry/semiconductors/nvidias-20-billion-groq-deal-produces-its-first-chip">Tom&#8217;s Hardware</a>. Foxconn reportedly accelerating LPX rack production ahead of schedule, per <a href="https://wccftech.com/nvidia-35x-ai-inferencing-leap-arrives-early-foxconn-fast-fowards-groq-3-lpx-racks/">Wccftech</a> (July 2026); production timing is press-reported, not confirmed by Nvidia.</p><p>[10] <a href="https://techcrunch.com/2025/06/06/amd-acqui-hires-the-employees-behind-untether-ai/">TechCrunch</a>, June 6, 2025: AMD hired the engineering team behind Untether AI. <a href="https://www.tomshardware.com/tech-industry/amd-scoops-entire-untether-ai-chip-team-canada-ai-inference-outfit-will-cease-product-support">Tom&#8217;s Hardware</a>: Untether ceased product support. Untether AI was headquartered in Toronto and built at-memory inference accelerators.</p><p>[11] AMD, <a href="https://newsroom.amd.com/news/amd-2q-2026-earnings/">&#8220;AMD Reports Second Quarter 2026 Financial Results,&#8221;</a> August 4, 2026. Revenue $11.5 billion, up 50% year over year; Data Center segment $6.7 billion, up 107%. Helios rack-scale systems described as beginning to ramp, with deployments named for Anthropic, Meta, Microsoft, OpenAI, and Oracle; MI400-series GPUs (MI455X, MI430X) recently launched.</p><p>[12] Lisa Su, remarks at AMD&#8217;s <a href="https://ir.amd.com/news-events/press-releases/detail/1294/aai-2026-amd-delivers-full-stack-compute-for-the-agentic-ai-era">Advancing AI 2026</a> launch event (July 20, 2026, where Helios and the MI400 series debuted), as reported by <a href="https://stocktwits.com/news-articles/markets/equity/amd-buys-toronto-ai-chip-startup-taalas-retail-says-its-a-move-to-compete-more-directly-with-nvidia/cZoBg5yRJJM">Stocktwits/Yahoo Finance</a> (August 6, 2026): &#8220;a big believer that there&#8217;s no one-size-fits-all as it comes to chips,&#8221; in the context of GPUs remaining dominant for flexibility across new models. Reported speech; a primary transcript was not located, and the wording should be treated accordingly.</p><p>[13] Julien Simon, <a href="https://www.airealist.ai/p/train-deploy-write-down">&#8220;Train, Deploy, Write Down,&#8221;</a> The AI Realist, April 7, 2026. GPT-5 cadence per OpenAI release notes: five major versions between August 2025 and March 2026, none surviving longer than four months as the current flagship; hardware depreciation schedules of five to six years per hyperscaler filings.</p><p>[14] Amazon, <a href="https://www.aboutamazon.com/news/company-news/amazon-invests-additional-5-billion-anthropic-ai">&#8220;Amazon and Anthropic deepen their collaboration,&#8221;</a> April 20, 2026: &#8220;Anthropic works closely with Annapurna Labs on developing and optimizing future Trainium chips, providing direct feedback from Claude training workloads to shape next-generation chip design.&#8221; Broadcom&#8217;s fourth custom-accelerator customer was revealed as Anthropic at its Q4 FY2025 earnings, per <a href="https://www.cnbc.com/2025/12/11/broadcom-reveals-its-mystery-10-billion-customer-is-anthropic.html">CNBC</a>, December 11, 2025. Anthropic&#8217;s Helios deployment: see [11].</p><p>[15] <a href="https://www.cnbc.com/2026/05/13/cerebras-prices-ipo-above-expected-range-wall-street-expects-ai-flood.html">CNBC</a>, May 13, 2026: Cerebras priced its IPO at $185 per share, above the expected range, raising approximately $5.5 billion; trading on Nasdaq as CBRS from May 14. See also the <a href="https://www.cerebras.ai/press-release/cerebras-systems-announces-pricing-of-initial-public-offering">Cerebras pricing release</a>.</p>]]></content:encoded></item><item><title><![CDATA[From Google Brain to Google Drain]]></title><description><![CDATA[What six weeks of exits say about the model Google hasn&#8217;t shipped.]]></description><link>https://www.airealist.ai/p/from-google-brain-to-google-drain</link><guid isPermaLink="false">https://www.airealist.ai/p/from-google-brain-to-google-drain</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Wed, 05 Aug 2026 18:22:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1Q-d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1Q-d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1Q-d!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!1Q-d!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!1Q-d!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!1Q-d!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1Q-d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2232557,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/209962444?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1Q-d!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!1Q-d!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!1Q-d!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!1Q-d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82789d4e-a2d0-48b1-94c1-60b42b7a6e06_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">Nine months ago, Google&#8217;s redemption arc looked complete. Gemini 3 shipped in November to strong reviews, the app passed 950 million monthly users, and the Bard debacle of 2023 &#8212; the demo error that once wiped roughly $100 billion off the market cap &#8212; reads like ancient history.[1] This morning, Demis Hassabis stepped back from running Google DeepMind. Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le left to start a company. The stock fell five percent in early trading.[2] <strong>So the right question about today&#8217;s reorganization is not who runs DeepMind now. It is: what happened between November and August?</strong></p><p style="text-align: justify;">Start with what the departures actually are, because the &#8220;brain drain&#8221; framing undersells it. Google put three technical co-leads on Gemini in 2024: Dean, Vinyals, and Noam Shazeer &#8212; the last brought back in a roughly $2.7 billion licensing deal largely to do that job.[3] All three are gone as of this summer. Shazeer walked to OpenAI in June. Nobel laureate John Jumper left for Anthropic the same month. Dean and Vinyals left this morning.[3] </p><div class="pullquote"><p style="text-align: center;">That is not attrition at the edges of a research organization; the people running the model walked out of the building in under two months.</p></div><p style="text-align: justify;">Now read the confirmed record on the model itself. No flagship successor to Gemini 3 has shipped in nine months &#8212; not damning by itself, but conspicuous for a company that shipped its way out of the Bard hole on cadence. The trade press has reported repeated delays and an architectural rebuild of the next Gemini cycle; Google has never confirmed it, and the sourcing is too thin to lean on.[4] You don&#8217;t need it. Sundar Pichai&#8217;s own memo this morning talks about Gemini 4 &#8212; and never mentions the intermediate release the trade press spent the spring waiting for.[5] When the CEO&#8217;s letter skips straight to the next number, the cycle in between did not go well. And on July 30, six days before elevating Hassabis to focus on science, Google dissolved its Nobel-winning AlphaFold team and folded its members into Gemini, with core scientists reportedly departing for Anthropic.[6] A company long on model progress does not strip-mine its science bench to feed the product line. A company short on it does.</p><p style="text-align: justify;">Then read the destinations, because they are the closest thing the market will get to insider disclosure. Shazeer, the man with the clearest view of Google&#8217;s unreleased pipeline, chose the direct competitor: the recipe works, just not here. Dean, Ghemawat, Vinyals, and Le chose a startup whose stated mission is automating the search for new architectures &#8212; &#8220;it might be that we will discover a different transformer architecture,&#8221; Le said at launch, which is a polite way of pricing the current one.[7] Jumper and the AlphaFold core chose the rival that still funds science. Three exits, three verdicts, all rendered by people who saw the roadmap from inside. Any single move can be explained by a compensation package; the pattern of destinations cannot. </p><div class="pullquote"><p style="text-align: center;">Insiders didn&#8217;t publish model evaluations. They published destinations.</p></div><p style="text-align: justify;">Against that record, today&#8217;s reorganization reads as consequence, not strategy. Companies do not decapitate and restructure their AI leadership while the flagship is on track; they do it after the cycle breaks. The honest bull case is narrow: Koray Kavukcuoglu, a DeepMind veteran since 2012, now oversees model development, frontier research, and the Gemini product teams, unified under a single operator reporting to Pichai, replacing a three-co-lead research committee.[2] If the committee were the bottleneck, consolidation would help. If the talent was the engine, Google just watched the engine leave.</p><p style="text-align: justify;">What Google did manage brilliantly is the aftermath, and here the between-the-lines gets interesting. Pichai&#8217;s farewell contains the day&#8217;s most revealing sentence: &#8220;We&#8217;ll continue to work with them as a founding investor and Cloud partner.&#8221;[5] Regular readers know that structure from the revenue side &#8212; Nvidia into CoreWeave, Amazon into Anthropic, capital out as investment, and back as bookings.[8] Applied to headcount: Google holds equity in Discovery Loop, Google Cloud houses its compute for at least the first year, and a research team that was a consolidated expense becomes a customer with upside attached. There is even a recursive twist: if the next Gemini cycle really did stall on architecture, Google just bought a call option on the architecture search it couldn&#8217;t finish internally. Compare Meta, which faced its own misaligned scientist and handled it with a $14.3 billion hostile installation instead of a term sheet: Yann LeCun left, called his new boss &#8220;inexperienced,&#8221; and raised $1.03 billion for a competing lab with no Meta money in it.[9] </p><div class="pullquote"><p style="text-align: center;">Google&#8217;s exits produce tenants and partners; Meta&#8217;s produced a critic.</p></div><p style="text-align: justify;">But don&#8217;t mistake the elegance for an answer. The term sheet recovers salvage value from the exits; it does nothing to make Gemini competitive. Google still holds the strongest non-model position in the race: its own silicon, its own datacenters, distribution through Search and Android, and 950 million people already in the app,  which is exactly why the market took only five percent off rather than repricing the whole company.[1] The infrastructure moat buys time. It does not train the model.</p><p style="text-align: justify;">So here is the between-the-lines of August 5: Google did not announce a strategy for catching OpenAI and Anthropic this morning. It disclosed, in org-chart form, that the last one broke somewhere between November and June &#8212; and it showed that nobody in the industry finances a retreat more expertly. The elephant in the room is still there. It is just exceptionally well-financed.</p><div><hr></div><h3>Notes</h3><p>[1] Gemini 3 shipped in November 2025; <a href="https://cloud.google.com/blog/ko/products/ai-machine-learning/gemini-3-is-available-for-enterprise">Google Cloud blog</a>. The 950M+ monthly active users figure is Google&#8217;s own, from <a href="https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/">Pichai&#8217;s August 5, 2026 memo</a> &#8212; vendor-claimed, not independently verified. Bard demo error and ~$100B single-day market cap loss: February 2023, widely reported at the time.</p><p>[2] Hassabis becomes Chair of Google DeepMind and Chief Scientist of Alphabet, continuing to lead Isomorphic Labs; Kavukcuoglu becomes SVP of Google DeepMind overseeing &#8220;Gemini model development, Frontier AI research, and the Gemini app and developer teams,&#8221; reporting to Pichai. <a href="https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/">Pichai memo, ibid.</a>; <a href="https://x.com/demishassabis/status/2085034334914769203">Demis Hassabis on X</a>; <a href="https://www.axios.com/2026/08/05/google-deepmind-demis-hassabis-ai">Axios</a>. Alphabet shares fell as much as ~5% in early trading, recovering to roughly &#8722;3.5% by mid-afternoon ET: <a href="https://ca.investing.com/news/stock-market-news/alphabet-shares-fall-5-as-ai-pioneer-jeff-dean-exits-in-major-shakeup-4778533">Investing.com</a>; <a href="https://www.tradingkey.com/news/market-movers/262078737-market-movers-googl-20260805">TradingKey</a>. If publishing after the close, update to the closing print.</p><p>[3] Google appointed Noam Shazeer co-lead of Gemini in August 2024, alongside technical leads Jeff Dean and Oriol Vinyals: <a href="https://www.theinformation.com/briefings/google-makes-former-character-ai-ceo-shazeer-a-co-leader-of-gemini-ai">The Information</a>; <a href="https://tech.yahoo.com/ai/articles/google-appoints-former-character-ai-030853742.html">Reuters via Yahoo</a>. The Character.AI licensing deal that returned Shazeer to Google was reported at ~$2.7B. Shazeer to OpenAI and John Jumper (Nobel Prize in Chemistry 2024, AlphaFold) to Anthropic, both June 2026: <a href="https://fortune.com/2026/06/23/google-deepmind-ai-researcher-departures-raise-doubts-about-ability-to-win-the-ai-race-shazeer-jumper-eye-on-ai/">Fortune</a>; <a href="https://techcrunch.com/2026/06/24/ai-researchers-continue-to-leave-google-for-its-rivals/">TechCrunch</a>.</p><p>[4] Reports of three delays and a &#8220;full architectural rebuild&#8221; of the Gemini 3.5 cycle, and of Gemini 4 pre-training beginning with no announced date: <a href="https://finance.biggo.com/news/6f0c6bb2-795f-4c57-9d09-6db691d7638a">BigGo</a>; <a href="https://theairankings.com/google/gemini-3-5-pro/">The AI Rankings</a>; <a href="https://memeburn.com/google-gemini-4-training-starts/">Memeburn</a>. C-tier sourcing, unconfirmed by Google &#8212; cited here as reports only; no claim in the body rests on them.</p><p>[5] <a href="https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/">Pichai memo, August 5, 2026</a>. Full Discovery Loop sentence: &#8220;We&#8217;ll continue to work with them as a founding investor and Cloud partner, and collaborate on a research framework for ML systems and related infrastructure advances.&#8221; The memo references Gemini 4; it does not mention Gemini 3.5.</p><p>[6] Google DeepMind disbanded the AlphaFold team on July 30, 2026, redirecting members to Gemini work; core scientists reportedly moved to Anthropic. <a href="https://www.engadget.com/2225849/google-shuts-down-alphafold/">Engadget</a>; <a href="https://www.pymnts.com/google/2026/google-reshuffles-nobel-winning-deepmind-ai-team/">PYMNTS</a>.</p><p>[7] Discovery Loop: public benefit corporation co-founded by Dean, Ghemawat, Vinyals, and Le; automating the experimental loop of the scientific method, ML research first; Khosla Ventures and Radical Ventures backing, round size and Google&#8217;s stake undisclosed; Google Cloud compute for the first year. Quoc Le quote from launch coverage. <a href="https://www.unite.ai/jeff-dean-leaves-google-to-automate-the-scientific-method-with-discovery-loop/">Unite.AI</a>; <a href="https://www.cnbc.com/2026/08/05/google-chief-scientist-jeff-dean-leaving-company-after-27-years.html">CNBC</a>.</p><p>[8] See prior coverage of the round-trip structure: <a href="https://www.airealist.ai/p/welcome-to-hotel-abilene">&#8220;Welcome to Hotel Abilene&#8221;</a> and <a href="https://www.airealist.ai/p/compute-equals-commitments">&#8220;Compute Equals Commitments&#8221;</a>, The AI Realist.</p><p>[9] Meta&#8217;s investment in Scale AI, announced June 12, 2025: $14.3 billion for a 49% non-voting stake, with Scale CEO Alexandr Wang (then 28) joining to lead what became Meta Superintelligence Labs: <a href="https://www.cnbc.com/2025/06/12/scale-ai-founder-wang-announces-exit-for-meta-part-of-14-billion-deal.html">CNBC</a>. LeCun on Wang (&#8221;young,&#8221; &#8220;inexperienced&#8221;): <a href="https://the-decoder.com/you-certainly-dont-tell-a-researcher-like-me-what-to-do-says-lecun-as-he-exits-meta-for-his-own-startup/">The Decoder</a>; <a href="https://www.cnbc.com/2026/01/05/ai-godfather-calls-meta-ai-boss-alexander-wang-inexperienced-.html">CNBC</a>. AMI Labs: $1.03 billion at a reported $3.5 billion pre-money, March 2026, co-led by Cathay Innovation, Greycroft, Hiro Capital, HV Capital, and Bezos Expeditions, with Nvidia and Eric Schmidt participating; Meta absent from the reported investor list: <a href="https://techcrunch.com/2026/03/09/yann-lecuns-ami-labs-raises-1-03-billion-to-build-world-models/">TechCrunch</a>. AMI builds on the V-JEPA architecture LeCun&#8217;s team developed and open-sourced at Meta &#8212; the one asset Meta retained.</p>]]></content:encoded></item><item><title><![CDATA[You Have To Ask Me Nicely]]></title><description><![CDATA[OVH told Canada how to get the data it refused to hand over. On August 18, the European Union stops making the argument OVH is making in court.]]></description><link>https://www.airealist.ai/p/you-have-to-ask-me-nicely</link><guid isPermaLink="false">https://www.airealist.ai/p/you-have-to-ask-me-nicely</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Tue, 04 Aug 2026 17:44:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4abb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4abb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4abb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!4abb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!4abb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4abb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4abb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1394149,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/209783711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4abb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!4abb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!4abb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4abb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7aa67239-cd2b-4c00-a273-539026074c02_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">In January, OVHcloud said the customer records Canadian police had been demanding since 2024 &#8220;cannot be provided to the Canadian authorities,&#8221; and described the French company and its European subsidiaries as not subject to foreign extraterritorial law.[1]</p><p style="text-align: justify;">On 31 July, facing criminal charges in Ontario, the company published a statement naming the treaty channel through which those same records could be obtained, &#8220;including the data sought by Canadian authorities.&#8221; French authorities, it added, had indicated such a request would be expedited and processed &#8220;within a matter of weeks.&#8221;[2]</p><p>Same records. Same case. Six months apart.</p><p style="text-align: justify;">There is a scene in <em>A Few Good Men</em> where Tom Cruise, playing a Navy lawyer called Kaffee, asks Jack Nicholson&#8217;s Colonel Jessep for a transfer order. Jessep tells him the paperwork is his for the asking. He just has to ask nicely. Kaffee asks nicely, and Jessep hands it over, beaming. The document was never in doubt. The courtesy was, and the courtesy cost Jessep nothing, which is why he made a point of demanding it.</p><div id="youtube2-vyMggFe9WRQ" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;vyMggFe9WRQ&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/vyMggFe9WRQ?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>That is the argument OVH is running in Ontario. It is also, as it turns out, the argument France has been running since 1968.</p><h2>The Charge</h2><p style="text-align: justify;">The charges are real and criminal. On 31 July, OVH Groupe SA and its Montreal subsidiary H&#233;bergement OVH Inc. confirmed that Canadian authorities had charged both with failing to comply with a production order and with obstruction of justice.[3] A French company listed on Euronext Paris is being prosecuted in Ontario for declining to hand over records held in Europe. Octave Klaba said the company would defend the principle of corporate separateness with full confidence.[4]</p><p style="text-align: justify;">Readers who followed this in April know what it looks like. In April 2024, the Royal Canadian Mounted Police obtained a production order for subscriber and account data linked to four IP addresses hosted on OVH servers in France, the United Kingdom, and Australia, as part of a sealed national security investigation. In September 2025, Justice Heather Perkins-McVey held that OVH&#8217;s commercial and virtual presence in Canada brought the French parent within Canadian jurisdiction, that mutual legal assistance was permissive rather than mandatory, and that the French blocking statute was, in practical effect, an empty vessel. OVH filed for judicial review. I covered that ruling, and what it revealed about the European Commission&#8217;s cloud scoring rubric, in &#8220;<a href="https://www.airealist.ai/p/ten-percent-sovereign">Ten Percent Sovereign</a>.&#8221;[5]</p><p style="text-align: justify;">That piece asked why Brussels assigned legal sovereignty 10% of its scoring matrix. This one asks a narrower question, and the answer is less comfortable: what is the thing being weighed?</p><p><strong>Because in Ontario, OVH is not arguing that Canada cannot have the data. The argument is that Canada asked the wrong way.</strong></p><h2>The Test the Senate Set</h2><p style="text-align: justify;">Fourteen months before the charges, a French senator wrote the only sovereignty test in this story that survives contact with the law. 10 June 2025, Microsoft France&#8217;s Anton Carniaux under oath before a Senate commission of inquiry. Senator Dany Wattebled asks whether he can guarantee that French citizens&#8217; data will never be handed over without the explicit agreement of the French authorities. &#8220;Non, je ne peux pas le garantir.&#8221;[6] Wattebled did not ask where the data was stored. He asked who has to agree before it moves. Hold on to that question.</p><p style="text-align: justify;">OVHcloud made the most of that hearing. In August, chief legal officer Solange Viegas Dos Reis told The Register that Microsoft had finally told the truth, and that it was no surprise to her. Customers were shocked, she said, because they had spent years being reassured the American statute would not touch them. &#8220;It&#8217;s false! Because, indeed, the data can be communicated.&#8221;[7] Customers, she said, had started asking how it works with other providers.</p><p style="text-align: justify;">That question got an answer eleven months later, in an Ontario courtroom, and The Register saw the joke coming in November: it would be deeply ironic, it noted, if OVH could not guarantee the same thing, because the company has a subsidiary in Canada.[8] The irony is where this starts, not where it finishes. What matters is not that OVH ended up in Microsoft&#8217;s position. It is that OVH&#8217;s own legal documents had been describing that position all along.</p><p style="text-align: justify;">The CLOUD Act page published by OVH&#8217;s American entity states that the company will comply with lawful requests from public authorities, that these requests can reach data stored outside the United States, and that requests from countries without an executive agreement go through the treaty route.[9] The group&#8217;s data policy promises, when authorities come asking, to limit disclosure to what the authority requires. Limit, not refuse. The Canadian agreement spells out the drill: check that the requesting authority is competent and the request valid, turn away what is obviously neither, hand over the rest, and tell the customer.[10] It is Carniaux&#8217;s testimony, with a different flag.</p><p style="text-align: justify;">Let&#8217;s be fair to OVH, because the company has been careful and largely consistent. Its July 2025 statement distinguished OVH US, which is subject to American process for its own customers, from the French entity and its European subsidiaries, which it said are not subject to the CLOUD Act, the Patriot Act, or FISA.[11] That is accurate as far as it goes, and the FAQ&#8217;s insistence on the treaty route is the same position the company is now arguing in Ontario. Only one public statement in the whole sequence ever promised more than the contracts do, and it is the January one, which said the data cannot be provided. The July 2026 release returned the contracts to the record and added a delivery estimate.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HsYL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HsYL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png 424w, https://substackcdn.com/image/fetch/$s_!HsYL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png 848w, https://substackcdn.com/image/fetch/$s_!HsYL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!HsYL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HsYL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png" width="1456" height="832" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:832,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:321555,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/209783711?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HsYL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png 424w, https://substackcdn.com/image/fetch/$s_!HsYL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png 848w, https://substackcdn.com/image/fetch/$s_!HsYL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!HsYL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac1a0c79-3a13-481c-8bca-7b5a812d8b6a_2800x1600.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">The promise that goes further belongs to the state. SecNumCloud, the French cloud doctrine that qualifies sensitive data, added legal requirements in version 3.2: European capital controls, European applicable law, and contractual immunity clauses. OVH holds it for Bare Metal Pod and said on the day it was awarded that the qualification protects against the legal risks of non-European regulation.[12]</p><p style="text-align: justify;">Read the requirements narrowly, and they aim at third-country law, which is a fair reading and probably ANSSI&#8217;s. But nobody buying on that basis was distinguishing an American injunction from a European one. They heard that their data could not be taken without French agreement. Wattebled asked his question about an American injunction; <strong>the interesting version now is the one he had no reason to ask. Can a qualified French provider guarantee that customer data will never be transmitted to any foreign authority without the explicit agreement of the French authorities?</strong></p><p>Until August 18, more or less. After that, for the most commonly demanded category of data, no. And the reason has nothing to do with American law.</p><h2>What the Guarantee Says</h2><p style="text-align: justify;">Start with what the explicit agreement of the French authorities consists of, because the French state has described it, on the record, in this case.</p><p style="text-align: justify;">In December 2025, Deputy Bastien Lachaud put a written question to the economy ministry: what would the government do to stop a foreign authority seizing data hosted on French soil? The reply was published in the Journal Officiel on 3 March 2026.[13]</p><p style="text-align: justify;">The government declined to comment on foreign proceedings, then explained the law it relies on. Established in 1968, the blocking statute has two limbs. The first forbids handing a foreign authority anything touching French security, sovereignty, public order, or essential economic interests. The second, added in 1980, is the one at issue here, and it does something else. It exists to make sure that requests for evidence &#8220;empruntent bien les canaux et trait&#233;s applicables,&#8221; that they take the applicable channels and treaties. Which channel, the ministry adds, means the 2001 Budapest Convention on cybercrime.</p><p>That second limb does not forbid disclosure. It governs which door the request comes through, and Paris says so in its own words, in the Journal Officiel, about this case.</p><p style="text-align: justify;">The enforcement apparatus matches the ambition. A desk at Bercy, the SISSE, issues opinions on whether the 1968 law applies to a particular demand. The opinion is addressed to the French company that asked, which is expected to forward it to the foreign authority so that the authority can learn the formal French position. If that proves insufficient, the desk can ask a liaison magistrate to raise awareness of the law abroad or open a diplomatic channel. The ministry notes that the desk is working to become better known. That is the protection. Advisory opinions, forwarded onward, backed by a liaison magistrate and a phone call.</p><p style="text-align: justify;">An Ontario court weighed it all and did not detain long. One conviction has ever been reported under that second limb, in the forty-six years since it was enacted. Expert witnesses could point to no case anywhere in which a foreign court had respected the statute. Perkins-McVey borrowed the English High Court&#8217;s verdict on it: an empty vessel.[14] An Ontario court had already brushed past French blocking provisions in a civil case twenty-five years earlier.[15]</p><p style="text-align: justify;">So the French and Canadian positions sit closer together than the diplomatic temperature suggests. France says the data is available through the treaty. Canada says the treaty is optional. <strong>Nobody in the case says the data is unreachable, and the company holding it has now confirmed in writing that it is not.</strong></p><h2>The Third Door</h2><p style="text-align: justify;">Both sides of this dispute point at the same door. OVH calls it mutual legal assistance; the French government, answering Lachaud, names the treaty behind the phrase, the Budapest Convention on cybercrime. It dates from 2001, runs state to state through central authorities, and every party to this case concedes it is slow. In 2022, the Council of Europe adopted a protocol that allows authorities in one country to directly request subscriber information from a provider in another country, which is the category the RCMP requested.[16]</p><p style="text-align: justify;">France signed it, and three weeks later, Brussels told member states to ratify. Canada signed it in June. Neither has been ratified, and 52 signatures have produced 4 ratifications, 1 short of the 5 required to bring it into force.[17] Both governments agreed that the door should exist. Neither built it. One of them then walked through the wall and is now prosecuting the company that was standing in front of it.</p><p style="text-align: justify;">OVH is not the first provider to stake everything on the treaty route. Microsoft made the same argument for emails stored on a Dublin server and won it at the Second Circuit in 2016. Twenty months later, Congress legislated the win away with the CLOUD Act; the full story is in &#8220;<a href="https://www.airealist.ai/p/two-sovereign-clouds-one-legal-wall">Two Sovereign Clouds</a>.&#8221;[18] European sovereignty marketing was built on the wreckage of that case, and it was built against a single statute. Every &#8220;not subject to the CLOUD Act&#8221; claim answers the American door. Canada has now opened a second door through its courts. The European Union spent five years building a third, and that one opens on August 18. It opens on OVH, too.</p><p style="text-align: justify;">On August 18, 2026, the European Union&#8217;s e-Evidence Regulation becomes applicable. It creates the European Production Order: <strong>an authority in one member state can compel a provider established in another to hand over electronic evidence directly, with no state-to-state procedure in between</strong>. The Commission&#8217;s own summary gives the aim as speeding up access &#8220;regardless of where the data is located,&#8221; and lists subscriber data and the IP addresses needed to identify a user among the categories covered.[19]</p><p style="text-align: justify;">Ten days to comply. Eight hours in an emergency. Penalties of up to two percent of worldwide annual turnover for providers that do not.[20]</p><p style="text-align: justify;">Now, the detail that decides the argument. Ask for the contents of messages, and a judge has to sign; the country where the provider sits is told, and it has ten days to object, or ninety-six hours if the case is urgent. <strong>Ask for subscriber records, or for the address that identifies whoever was behind an account, and neither applies. No judge required. Nobody told.</strong>[21] The notification shrinks further from there, disappearing even for message content when the offense and the suspect both sit in the requesting country, which describes most investigations.[22] Subscriber records also sit at the bottom of the instrument&#8217;s threshold: an order may issue for any criminal offense, where traffic and content require offenses carrying at least three years.[22]</p><p style="text-align: justify;">Put the pieces together and imagine the file. A prosecutor in Dublin opens an investigation into a customer of a French host. After August 18, she can send a certificate to that host&#8217;s designated establishment, requiring the subscriber records within ten days for any criminal offense, without any judge required by the Regulation, and without anyone in Paris being told it happened. The SISSE desk cannot issue an opinion on a demand it is unaware of. The liaison magistrate has nobody to call.</p><p style="text-align: justify;">Dublin is the deliberate choice. The argument does not need a member state with a rule-of-law problem, and reaching for one would let the reader file this under judicial standards rather than sovereignty. Ireland is the jurisdiction European data policy treats as the safe pair of hands, and the guarantee fails against Ireland just as it does against anyone else, because notification does not take reputation into account. It simply is not there.</p><p style="text-align: justify;">That certificate is Wattebled&#8217;s question served under a different flag, and from August 18, no French provider can answer it any better than Carniaux did. No, I cannot guarantee it. The reason is a European regulation, not an American one.</p><h2>And Then Where?</h2><p style="text-align: justify;">The Regulation governs how a member state gets the data and says nothing about what it may do with it afterward. Onward transfer to a third country falls under the general law enforcement data regime, which permits it on the basis of adequacy, safeguards, or a list of derogations. The European Union and the United States have had a law enforcement transfer agreement since 2016.[23] Whether the records are then transferred from the police file to an intelligence file is not a European question at all. National security is reserved for the member states by the founding treaties, and several of them have run multilateral signals intelligence arrangements with Washington for decades. So the records can be compelled by Dublin without Paris knowing, and travel onward under a framework in which Paris has no standing, because Paris was never told there was a file.</p><h2>What Was Sold</h2><p style="text-align: justify;">Compare all that with what France told its own parliament it was defending. Transfers to a third state, the question ran, can go through a derogation procedure supervised by French authorities, and Canada&#8217;s refusal to seek prior French control ran against the principles of digital sovereignty.[24] Whatever one thinks of Ottawa, the RCMP had to litigate for two years and lay criminal charges to reach a place an Irish prosecutor will reach in ten days with a form.</p><p style="text-align: justify;">The objection to all this is good and deserves to be stated in full. A European Production Order is not a Canadian one wearing a European flag. It runs between states bound by a common rights floor, carries refusal grounds and remedies, obliges the issuing authority to notify the person whose data was taken, and allows the provider to pull the enforcing state back in. A Canadian order carries none of that. Anyone who tells you the two are equivalent is selling something. And none of it has happened yet: no European Production Order has ever been issued, and how prosecutors use the instrument is a fact about the future.</p><p style="text-align: justify;">But equivalence was never the claim. What was sold to European cloud buyers was reach, not safeguards. Data resident in France is beyond the reach of authorities outside France, as obtaining it requires the consent of French authorities. From August 18 onward, for the most commonly requested category of data, French consent is not part of the design. What survives is the claim that some requesting states are better than others, an argument about who belongs to which club. In the transatlantic context, that argument has already collapsed twice, under the names Safe Harbor and Privacy Shield.</p><h2>The Channel Test</h2><p>Everything above turns Wattebled&#8217;s question into something portable, and it replaces the question most buyers ask.</p><p style="text-align: justify;">Data residency is the wrong question. It tells you where the disk is. It tells you nothing about who can compel the disk. Three questions do the work instead.</p><p style="text-align: justify;"><strong>Who has to sign?</strong> Name the authority, not the country. &#8220;Hosted in France&#8221; is not an answer. &#8220;A French investigating magistrate, acting on a letter rogatory transmitted through the central authority, after a SISSE opinion on whether the 1968 law even applies,&#8221; is an answer. Say it out loud and count the links. Every one of them is a place where a foreign state can be told no, and every one of them is a place where a foreign state can be told yes.</p><p style="text-align: justify;"><strong>How long does that take?</strong> Not the statutory deadline. The observed time. If the provider has published a figure, use theirs. OVH published its own, and it&#8217;s weeks.</p><p style="text-align: justify;"><strong>Has that authority ever said no?</strong> This is the question that separates a wall from a queue, and it is the one nobody asks, because it is answered by enforcement history rather than by statutory text. France publishes no refusal statistics for assistance requests, and none surfaced in this case. That does not prove Paris has never refused. It proves you cannot find out, which, for a buyer, is the operative fact. What the record shows is what happens to those who rely on the wall: one conviction under the blocking statute since 1980, and no foreign court has ever recognized it.</p><p style="text-align: justify;">Run the test against any sovereign offering, European or otherwise. If the third question has no documented refusal, the product is sold with a delay and a review step. Delay has value. Investigations that are delayed all the time, and a review step counts for something. It just isn&#8217;t the thing the marketing describes.</p><p style="text-align: justify;">Does anything pass? For subscriber data after the 18th, nothing European does, and that is the useful finding. Residual exposure to a direct foreign order becomes roughly uniform across every provider with an establishment in the Union, qualified or not. The sovereignty premium still buys real things: a European supply chain, an operator outside American discovery, a stack you can audit, a government that answers to your parliament. What it no longer buys, for the most commonly demanded category of data, is the thing on the label. Price it accordingly.</p><h2>What Would Break This</h2><p>The whole argument turns on one thing, and it is checkable.</p><p style="text-align: justify;">France has to refuse. If Canada files the request the way OVH says it should, and Paris declines it, or sits on it long enough that the investigation dies, then the French review step is a real gate, and I have this wrong. The guarantee would be a wall after all, and a slow one, which is what a wall is.</p><p style="text-align: justify;">Two lesser tests. If the Ontario Superior Court quashes the production order on territorial grounds instead of procedural ones, the Canadian doctrine narrows and stops traveling; as of publication, no ruling on the judicial review has been reported.[25] And if any member state publicly refuses to execute a European Production Order on sovereignty grounds after 18 August, the intra-EU channel will have its own gate.</p><p style="text-align: justify;">In the meantime, there is work to do that costs nothing. Stop asking vendors where the data lives, which every vendor can answer, and ask the three questions, which most cannot. Get the answers in writing, dated, and in the contract file. Then run the same three against the SecNumCloud qualifications, the Cloud de Confiance arrangements, and the SEAL tiers in the Commission&#8217;s framework. None of them scored on the third question, because none of them asked it.</p><p style="text-align: justify;">Three more for the provider itself. Which legal entity, in which member state, is its designated establishment for European Production Orders, because that choice decides whose prosecutors reach you fastest? Where is its law enforcement transparency report, broken down by requesting jurisdiction and outcome, because that document is the only place question three can ever be answered. And what survives once the data has left its hands? There's no good answer, because nothing does. Watch how long it takes for the transparency report to be released.</p><p>Jessep could have refused Kaffee outright. He never meant to. He wanted the courtesy first because it was free, and because demanding it let him keep the form of authority while handing over the substance. That was the arrangement Roubaix and Bercy were selling, and buyers paid a premium for it in good faith.</p><p>From August 18, nobody has to ask, nicely or not. European data sovereignty was never a wall. It was a rule of manners, and manners are optional now.</p><div><hr></div><h3>Notes</h3><p>[1] OVHcloud statement to The Register, added to its report on 26 January 2026: &#8220;OVHcloud&#8217;s number one concern and priority is to protect its customers&#8217; data. This is why this data cannot be provided to the Canadian authorities.&#8221; The same statement describes the group as &#8220;a multi-local player that is not subject to extraterritorial laws.&#8221; See <a href="https://www.theregister.com/2025/11/27/canada_court_ovh/">The Register&#8217;s report of 27 November 2025</a>, to which the statement was added in an update dated 26 January 2026; the quoted wording now appears in the body of that continuously updated report. The statement was given to that publication and is not corroborated by an OVH-published document.</p><p>[2] OVH Groupe SA, &#8220;<a href="https://blog.ovhcloud.com/en/posts/ovhcloud-contests-canadian-production-order-charges/">OVHcloud Confirms Intent to Vigorously Contest Charges Related to Canadian Production Order</a>,&#8221; 31 July 2026. Direct quotation.</p><p>[3] Same release. Charges are for failure to comply with a production order under section 487.0198 of the Criminal Code of Canada and obstruction of justice under section 139(2); the underlying production order was issued on 19 April 2024 under section 487.014. Section 487.0198 creates a summary conviction offence for contravening an order made under sections 487.013 to 487.018; section 139(2) is the general obstruction offence and is indictable.</p><p>[4] Same release, quoting Octave Klaba, founder, chairman and chief executive officer. Klaba resumed the CEO role on 20 October 2025, when the board reunited the chairman and CEO functions, ending Benjamin Revcolevschi&#8217;s one-year tenure (<a href="https://corporate.ovhcloud.com/en/newsroom/news/octave-klaba-chairman-ceo/">OVHcloud corporate release, 21 October 2025</a>).</p><p>[5] <em>R. v. OVH Group SA and H&#233;bergement OVH Inc.</em>, Ontario Court of Justice (Ottawa), Court File 24-000659, Perkins-McVey J., decision released 25 September 2025. <a href="https://drive.google.com/file/d/1QVwO9lPmxuDSQsGd9fHH3QN_ToXs2LQ8/view">Signed decision</a> via David Fraser, McInnes Cooper. Julien Simon, &#8220;<a href="https://www.airealist.ai/p/ten-percent-sovereign">Ten Percent Sovereign</a>,&#8221; The AI Realist, 22 April 2026, which covers the ruling, the virtual presence line of authority across four provinces, and the European Commission&#8217;s Cloud Sovereignty Framework in full.</p><p>[6] S&#233;nat, commission d&#8217;enqu&#234;te sur la commande publique, <a href="https://www.senat.fr/compte-rendu-commissions/20250609/ce_commande_publique.html">compte rendu of the hearing of 10 June 2025</a>. Verbatim. Full analysis of the hearing, the statutory chain and the attribution of the question (Wattebled, per the official transcript, though some press reports credited commission president Simon Uzenat) in Julien Simon, &#8220;<a href="https://www.airealist.ai/p/two-sovereign-clouds-one-legal-wall">Two Sovereign Clouds, One Legal Wall</a>,&#8221; The AI Realist, 26 February 2026. Carniaux has since left Microsoft France.</p><p>[7] Solange Viegas Dos Reis, chief legal officer, OVHcloud, interviewed in &#8220;<a href="https://www.theregister.com/2025/08/27/ovhcloud_interview/">Microsoft can&#8217;t guarantee data sovereignty &#8211; OVHcloud says &#8216;We told you so&#8217;</a>,&#8221; The Register, 27 August 2025. Direct quotations.</p><p>[8] <a href="https://www.theregister.com/2025/11/27/canada_court_ovh/">The Register, 27 November 2025</a>, per note 1.</p><p>[9] OVH US, &#8220;<a href="https://us.ovhcloud.com/legal/faqs/cloud-act/">Cloud Act &#8211; Clarifying Lawful Overseas Use of Data</a>,&#8221; accessed August 2026. This page belongs to the group&#8217;s American entity, which is subject to American process; it is cited here for what the group publishes about its own compliance posture, not as a statement about the French entity. The page was quoted against OVH by AWS in July 2025, a conflation of entities this piece does not repeat. I quoted the same FAQ language in &#8220;<a href="https://julsimon.medium.com/the-sovereignty-mirage-why-european-clouds-wont-save-your-data-a565e82127f5">The Sovereignty Mirage</a>&#8220; in December 2025, eight months before the charges.</p><p>[10] OVHcloud, &#8220;<a href="https://www.ovhcloud.com/en/terms-and-conditions/privacy-policy/">Personal data usage policy</a>,&#8221; accessed August 2026, on requests from judicial, administrative or other authorities. Procedure at clause 4.2 of the <a href="https://storage.gra.cloud.ovh.net/v1/AUTH_325716a587c64897acbef9a4a4726e38/contracts/d54e222-OVH_Data_Protection_Agreement-CA-1.0.pdf">Personal Information Protection Agreement</a> applicable to Canadian clients, version dated 6 September 2023. Clause 4.3 addresses requests originating from an authority outside Canadian jurisdiction concerning a Canadian client.</p><p>[11] OVH statement to <a href="https://www.theregister.com/2025/07/25/microsoft_admits_it_cannot_guarantee/">The Register, July 2025</a>: &#8220;OVH Group abides by local laws in the countries it operates in. As such, OVH US may be subject to requests from American authorities within the framework of the Cloud Act as long as these demands are connected to customers of OVH US and are strictly compliant with applicable American law. The French OVH entity (or its European subsidiaries) is not subject to the Cloud Act, the Patriot Act or the FISA.&#8221;</p><p>[12] ANSSI, SecNumCloud requirements repository version 3.2 (March 2022), which added criteria protecting against non-European law: capital control by European entities, exclusive application of European law, and contractual immunity provisions. The qualification comprises more than 360 technical, organisational and legal requirements. ANSSI&#8217;s qualification decision for OVHcloud&#8217;s Bare Metal Pod was published on 24 March 2025; OVHcloud announced it on <a href="https://corporate.ovhcloud.com/en/newsroom/news/secnumcloud-qualification-bare-metal-pod/">31 March 2025</a>, stating that the qualification &#8220;acts as another form of protection against the legal risks linked to non-European regulations.&#8221; Note that these are structural and contractual requirements rather than a bar on lawful compulsion, a distinction this piece turns on. On what the Commission&#8217;s SEAL ladder does and does not score, see &#8220;<a href="https://www.airealist.ai/p/ten-percent-sovereign">Ten Percent Sovereign</a>&#8220;; on the architectural cost of the qualification, &#8220;<a href="https://www.airealist.ai/p/more-sovereign-different-stack-the-builder-tax">More Sovereign, Different Stack: The Builder Tax</a>.&#8221;</p><p>[13] Assembl&#233;e nationale, <a href="https://questions.assemblee-nationale.fr/q17/17-11534QE.htm">question &#233;crite n&#176; 11534</a>, M. Bastien Lachaud, published in the Journal Officiel of 9 December 2025, p. 9993; answer from the ministry for artificial intelligence and digital affairs published in the Journal Officiel of 3 March 2026, p. 1907. All quotations in this section are from the answer. Loi n&#176; 68-678 of 26 July 1968; Article 1 bis added by Loi n&#176; 80-538 of 16 July 1980; D&#233;cret n&#176; 2022-207 of 18 February 2022 establishing the SISSE opinion procedure. Penalties per the answer: six months&#8217; imprisonment and a fine of &#8364;18,000, and for legal persons a fine of &#8364;90,000.</p><p>[14] <a href="https://drive.google.com/file/d/1QVwO9lPmxuDSQsGd9fHH3QN_ToXs2LQ8/view">Ruling</a> at paragraphs 79 to 123, adopted at paragraph 115. &#8220;Empty vessel&#8221; from Butcher J in <em>Tugushev v. Orlov</em>, [2021] EWHC 1514 (Comm), handed down 28 May 2021, at paragraph 33, applied to Loi 68-678 by Cockerill J in <em><a href="https://caselaw.nationalarchives.gov.uk/ewhc/kb/2024/1424">Joshua &amp; Ors v. Renault SA &amp; Ors</a></em>, [2024] EWHC 1424 (KB), 11 June 2024, at paragraphs 77 to 78. The test as stated there is whether the foreign criminal law relied on is &#8220;not merely a text, or an empty vessel, but is regularly enforced.&#8221; The single reported Article 1 bis conviction is the &#8220;Christopher X&#8221; case, Cour de cassation, chambre criminelle, 12 December 2007, pourvoi n&#176; 07-83.228, discussed at paragraph 85.</p><p>[15] <em><a href="https://www.canlii.org/en/on/onsc/doc/2000/2000canlii22407/2000canlii22407.html">Wilson v. Servier Canada Inc.</a></em>, 2000 CanLII 22407 (ON SC), 50 O.R. (3d) 219 (Cumming J.), cited at ruling paragraph 95. A civil class action involving a French parent, not a criminal production order, and turning on Article 15 of the French Civil Code rather than Loi 68-678. Cited here for the Canadian judicial posture the OVH ruling drew on, not as precedent on the blocking statute.</p><p>[16] Council of Europe, <a href="https://www.coe.int/en/web/cybercrime/second-additional-protocol">Second Additional Protocol to the Convention on Cybercrime on enhanced co-operation and disclosure of electronic evidence</a> (CETS No. 224), adopted 17 November 2021, opened for signature 12 May 2022. Article 7 provides a legal basis for direct co-operation with service providers in another Party&#8217;s territory to obtain subscriber information, subject to reservations and declarations available to Parties.</p><p>[17] France signed on 27 January 2023, the thirty-second state to do so, alongside Germany: Council of Europe, &#8220;<a href="https://www.coe.int/fr/web/cybercrime/-/france-and-germany-become-32nd-and-33rd-states-to-sign-the-second-additional-protocol-to-the-convention-on-cybercrime">France and Germany become 32nd and 33rd states to sign the Second Additional Protocol</a>,&#8221; 27 January 2023. <a href="http://data.europa.eu/eli/dec/2023/436/oj/eng">Council Decision (EU) 2023/436</a> of 14 February 2023 authorised Member States to ratify the Protocol in the interest of the European Union (OJ L 63, 28.2.2023, pp. 48-53); signature had been authorised by Council Decision (EU) 2022/722 of 5 April 2022. Canada signed on 20 June 2023, the thirty-eighth state, per the <a href="https://www.coe.int/en/web/cybercrime/second-additional-protocol">Second Additional Protocol news archive</a>; its justice department has consulted publicly on whether to ratify and on whether to reserve against Article 7, and no instrument of ratification has been deposited. Signature count as of the Council of Europe&#8217;s most recent published tally; ratifications: Serbia first, <a href="https://www.coe.int/en/web/cybercrime/-/japan-becomes-2nd-state-to-ratify-the-second-additional-protocol-to-the-convention-on-cybercrime">Japan second on 10 August 2023</a>, the announcement of which confirms that five ratifications are required for entry into force, Hungary third on 5 February 2026, Costa Rica fourth on 15 April 2026. Entry into force falls on the first day of the month following three full months after the fifth ratification. Verified as of 15 April 2026.</p><p>[18] <em>United States v. Microsoft Corp.</em> The warrant was issued in 2013 for content stored in Microsoft&#8217;s Dublin facility; the Second Circuit held in 2016 that the Stored Communications Act did not reach it; the CLOUD Act was enacted in March 2018 and the Supreme Court dismissed the case as moot. The compelled disclosure provision is codified at 18 U.S.C. &#167; 2713. The full trace of the statutory chain is in &#8220;<a href="https://www.airealist.ai/p/two-sovereign-clouds-one-legal-wall">Two Sovereign Clouds, One Legal Wall</a>.&#8221; Other jurisdictions assert comparable reach; the count here is of the doors this case walks through.</p><p>[19] European Commission, <a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=LEGISSUM:4649186">summary of Regulation (EU) 2023/1543</a> on European Production and Preservation Orders for electronic evidence in criminal proceedings. Full text at <a href="http://data.europa.eu/eli/reg/2023/1543/oj/eng">OJ L 191, 28.7.2023, pp. 118&#8211;180</a>. Application from 18 August 2026. The Regulation binds all member states except Denmark, which is not bound by reason of its opt-out in the area of freedom, security and justice.</p><p>[20] <a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=LEGISSUM:4649186">Same summary</a>. Ten days for transmission of requested data, eight hours in emergencies, ten days for the enforcing authority to raise a refusal ground where notification applies, reduced to ninety-six hours in emergency cases, and pecuniary penalties of up to two percent of a provider&#8217;s total worldwide annual turnover.</p><p>[21] <a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=LEGISSUM:4649186">Same summary</a>: &#8220;If the electronic evidence includes content data or traffic data, except for data requested for the sole purpose of identifying the user, a court or judge must issue or review the order, and the judicial authority must notify the competent authority of the Member State in which the designated establishment is located, or the legal representative resides.&#8221; How the Regulation interacts with Loi 68-678 is untested; no French court has ruled on whether a European Production Order satisfies the applicable-channels requirement of Article 1 bis, though a directly applicable EU regulation is the stronger reading.</p><p>[22] Thresholds per eucrim, &#8220;<a href="https://eucrim.eu/news/e-evidence-regulation-and-directive-published/">E-evidence Regulation and Directive Published</a>&#8220;: subscriber and identification data for all criminal offences and for execution of custodial sentences of at least four months; traffic and content for offences carrying a maximum of at least three years, or listed offences committed by means of an information system. On notification, Jessica Shurson, &#8220;<a href="https://journals.sagepub.com/doi/10.1177/20322844251357090">The balance of efficiency and fundamental rights in the EU e-Evidence Regulation</a>,&#8221; <em>New Journal of European Criminal Law</em>, 2025, on Article 8(2) and the consequence that in most domestic cases the enforcing state receives no notice. See also Athina Sachoulidou, &#8220;<a href="https://journals.sagepub.com/doi/10.1177/20322844241258649">Cross-border access to electronic evidence in criminal matters</a>,&#8221; <em>New Journal of European Criminal Law</em>, 2024, and <a href="https://eucrim.eu/articles/critical-issues-in-the-new-eu-regulation-on-electronic-evidence-in-criminal-proceedings/">eucrim, &#8220;Critical Issues in the New EU Regulation on Electronic Evidence in Criminal Proceedings&#8221;</a>, on the exclusion of subscriber data from the notification requirement.</p><p>[23] Regulation (EU) 2023/1543 contains no onward-transfer provision of its own; it refers to Chapter V of <a href="https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex%3A32016L0680">Directive (EU) 2016/680</a> only in the context of conflicting third-country obligations (Article 17(7)). Transfers of law enforcement data to third countries are governed by Articles 35 to 38 of that Directive, permitting transfer on an adequacy decision, appropriate safeguards, or derogations for specific situations. The EU-US Umbrella Agreement of 2016 provides the standing framework for law enforcement transfers to the United States. Processing for national security purposes falls outside Union competence under Article 4(2) of the Treaty on European Union and outside the Directive&#8217;s scope; arrangements between national intelligence services are governed by national law and by bilateral or multilateral agreements that are not published. This paragraph describes the legal architecture, not any known transfer of the data at issue in this case.</p><p>[24] Assembl&#233;e nationale, <a href="https://questions.assemblee-nationale.fr/q17/17-11534QE.htm">question &#233;crite n&#176; 11534</a>. The description of the derogation procedure and of the Canadian position appears in the question as put by the deputy; the ministry&#8217;s answer sets out the SISSE opinion procedure and reaffirms the government&#8217;s opposition to any circumvention of applicable co-operation channels and treaties.</p><p>[25] OVH filed for judicial review to the Ontario Superior Court of Justice through Miller Thomson at the end of October 2025, per <a href="https://www.heise.de/en/news/Canadian-Court-OVHcloud-from-France-must-hand-over-user-data-11092029.html">heise online, 26 November 2025</a>, which also reports from the court filings that the data has been preserved and that the French Ministry of Justice offered accelerated processing by letter rogatory. No decision on the judicial review had been reported at the time of writing.</p>]]></content:encoded></item><item><title><![CDATA[Amazon Priced the Frontier and Declined It]]></title><description><![CDATA[On 30 June, AWS published two announcements pointing in opposite directions. In July, it closed its AGI Lab. All three were the same decision.]]></description><link>https://www.airealist.ai/p/amazon-priced-the-frontier-and-declined</link><guid isPermaLink="false">https://www.airealist.ai/p/amazon-priced-the-frontier-and-declined</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Wed, 29 Jul 2026 12:01:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tqju!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tqju!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tqju!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tqju!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tqju!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tqju!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tqju!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg" width="1024" height="541" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:541,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:66196,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/208959210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tqju!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tqju!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tqju!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tqju!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15837b4f-cb90-4909-a92f-f00bb5ffc039_1024x541.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On  June 30, AWS announced a billion-dollar investment in a new Forward Deployed Engineering organization.[1] The plan embeds engineers in customer companies in pods of five or six for roughly 45-day engagements, building agentic systems using the customer&#8217;s own data and leaving the customer self-sufficient when the pod withdraws.[2] AWS called the pods &#8220;frontier teams&#8221; and noted that many of the engineers had built their own AI services.</p><p style="text-align: justify;">The same day, AWS published a service availability page announcing the retirement of roughly 20 services and features.[3] Amazon Kendra, Amazon Q Business, and Amazon Bedrock Agents will enter maintenance and be closed to new customers from 30 July.[4][5][6] Ten SageMaker AI features went with them.</p><p style="text-align: justify;">One announcement spends a billion dollars putting AWS engineers inside customer businesses. The other withdraws the products that those customers were meant to use. Three weeks later, Amazon confirmed it was closing its San Francisco AGI Lab and cutting roles across its artificial general intelligence organization.[7][8]</p><p style="text-align: justify;">Read separately, these are a product cull, a spending initiative, and a research retreat. Read together, on the calendar they occupy, they are one decision about which layer of the artificial intelligence business Amazon intends to own.</p><h2>Three vectors and one direction</h2><p style="text-align: justify;">Amazon is spending at the bottom of the stack and at the top, and withdrawing from everything in between.</p><p style="text-align: justify;"><strong>Downward, into the fabric</strong>. Andy Jassy told investors in February that Amazon expected roughly $200bn of capital expenditure in 2026, predominantly in AWS, pointing to what the results release called &#8220;seminal opportunities&#8221; in AI, chips, robotics, and satellites.[9] Project Rainier went live with nearly half a million Trainium2 chips across multiple United States sites and a stated path beyond a million.[10] Trainium3 sits behind it.[11] This is the only layer Amazon is attempting to own outright, and the only one where it is building rather than buying the alternative.</p><p style="text-align: justify;"><strong>Outward, for the intelligence</strong>. Anthropic supplies the frontier under a commitment of more than $100bn over 10 years and up to 5 gigawatts of capacity, powered by Amazon&#8217;s own silicon.[12] The equity position produced $16.8bn in pre-tax gains in the first quarter of 2026, while the AWS segment grew 28% in the same quarter and carries more than $15bn in annualized AI services revenue.[13] That gain is a non-cash mark on a private company&#8217;s valuation, booked through non-operating income, and it reverses if Anthropic&#8217;s next round prices flat or down. Bank of America estimates that Anthropic-related workloads alone could contribute more than $1.5bn in sequential AWS revenue growth in the second quarter.[14] That is an analyst estimate, not a disclosure, and if it is near right, it describes customer concentration as much as momentum. One fact is observable: Claude&#8217;s per-token rates on Bedrock match Anthropic&#8217;s own API to the cent.[15] AWS doesn&#8217;t mark Claude up. Its margin is whatever sits between retail and the undisclosed remittance, plus everything a production workload consumes besides the model, and that second meter is the one no model provider can take away.</p><p style="text-align: justify;">The arithmetic explains the rest of the strategy. A frontier laboratory is a capital sink competing for the same engineers, accelerators, and megawatts as a business already growing at 28%, with a winner-take-most payoff. Instead, equity plus a supply contract does three jobs at once: it secures the output, it locks the supplier&#8217;s training and serving onto Amazon&#8217;s silicon for a decade, and it books the upside as investment gains, not research cost. The capacity remains full, and the income statement shows a research organization rather than a frontier-scale training program, a difference measured in tens of billions.</p><p style="text-align: justify;"><strong>Forward, to the customer</strong>. The $1 billion announced on 30 June goes to engineers who sit inside the customer&#8217;s building. Francessca Vasquez, the AWS vice president who runs the organization, framed the pitch around speed, and the named early customers include the Allen Institute, Cox Automotive, the NBA, the NFL, Ricoh, and Southwest Airlines.[1][2] Amazon is late to this: Palantir has run forward-deployed engineering for over a decade, and Salesforce, Google Cloud, and Anthropic all offer similar versions.[16]</p><p style="text-align: justify;">Look at what that vector costs. Embedded engineering carries service margins, and a company that spent twenty years teaching investors to value self-service software economics is putting a billion dollars into headcount that does not scale like software. Firms accept that when the customer won&#8217;t reach production without it, and the consumption on the other side outweighs the margin surrendered. The forty-five-day cycle and the insistence on leaving customers self-sufficient are the tell: these pods are a customer acquisition cost, not a consulting line.</p><p style="text-align: justify;">AWS frames the initiative as customer obsession, and on its own terms, it is. It is also, on Amazon&#8217;s terms, consumption: the pods deploy onto Bedrock and the surrounding services, and the annuity AWS is buying is the meter running after the engineers leave. And this survives the test that killed the software. An embedded engineer holds the one input that tooling shipping with the model can&#8217;t replicate: the customer&#8217;s own context. That&#8217;s why AWS is paying people margins for a function whose software equivalent it just retired.</p><p style="text-align: justify;">What sits between the fabric and the customer is the part Amazon is vacating: its own frontier models, its own retrieval, agent, and assistant products, and the tooling a customer would use to build models themselves. Amazon has concluded that owning intelligence is a worse business than metering it, and it&#8217;s paying to occupy both ends of the pipe instead.</p><div class="pullquote"><p><em>The interesting question isn't whether Amazon failed here; parts of this were failures, and public ones. It's what Amazon did with the verdict</em>.</p></div><h2 style="text-align: justify;">The organization chart is the strategy document</h2><p>In December 2025, Amazon replaced Rohit Prasad at the head of the AGI organization with Peter DeSantis, a longtime cloud infrastructure executive.[17] CNBC describes the resulting unit plainly: it builds AI models and includes groups working on silicon development and quantum computing.[7]</p><p style="text-align: justify;">That&#8217;s not a research organization with infrastructure attached; it&#8217;s an infrastructure organization with research attached, and the reporting line tells you how the company values each. A model program that reports to the executive who owns chips, data centers, and quantum is a program understood as a load on capacity rather than as a product line.</p><p style="text-align: justify;">The AGI Lab itself was founded in December 2024, built around several dozen employees Amazon acquired from Adept, including its co-founder David Luan, and grew to roughly eighty people at its peak. More than a dozen of the Adept hires have since left, Luan among them in February.[8] The Information, which reported the closure first, describes what the lab was for: building software to make AI agents more useful.[18] Amazon confirmed the site closure and said frontier model research continues under Pieter Abbeel, the Berkeley professor who joined in 2024 when Amazon licensed the technology and hired the team from his robotics startup, Covariant.</p><p style="text-align: justify;">Secondary reporting indicates the cuts fell on model customization and post-training roles, a claim carried by trade press and not confirmed by Amazon.[14] If it holds, it points somewhere specific: post-training is the work that turns a base model into a competitive product, and a company that keeps a research lab while cutting that function has made a choice about which of the two it is doing. Amazon has not disclosed which teams were affected, so hold it loosely.</p><p style="text-align: justify;">The company&#8217;s own account is consistent with this, if read literally. A spokesperson said Amazon was sharpening its focus on the initiatives that matter most for customers so it could move faster on what counts.[7] Customers sit at the top of the stack. Frontier research is not where AWS meets them, and a billion dollars of embedded engineers is.</p><h2>What a landlord keeps</h2><p style="text-align: justify;">The defeat reading has been made in print: The Next Web called the lab closure the clearest sign yet that Amazon had stepped back from a frontier race it never really led.[19] On that reading, the models never landed, the products never took, and a company retiring Kendra and Q Business is making a virtue of a loss. Half of that is simply true, and the tooling shows it best.</p><p style="text-align: justify;">Take SageMaker Clarify, a bias-detection tool built for tabular models. AWS rebuilt it for the new architecture, enabling foundation model evaluation to move from preview in November 2023 to general availability in April 2024.[20][21] Twenty-six months later, it went to maintenance.[22] Ground Truth followed the same arc, and Ground Truth Plus, converted earliest of all in May 2023, is the only item on the June 30 page to reach the full end of support.[23][24][3] Earliest conversion, hardest death. Whether AWS built the wrong thing or built the right thing too late is a fair debate. The outcome is not.</p><p style="text-align: justify;">For now, the surviving services are close to the hardware. SageMaker HyperPod, the one part of the platform still shipping features monthly, does cluster resilience: gang scheduling, continuous provisioning, job recovery, and capacity reservation on Trainium.[25][26][27][28][29] The scheduling logic exists in open source, and a Kubernetes engineer will say so. What does not exist outside the operator is the input: which node sits on which network spine, which accelerator is degrading, which instance is about to be retired, and whether the capacity exists to be reserved at all. The position is the privileged telemetry, not the feature.</p><p style="text-align: justify;">Call it the fabric test. A function that runs against a model goes to whoever ships alongside the model. A function that requires privileged access to physical capacity stays with whoever owns the capacity. It predicts the consequential AI retirements on the June 30 page: Kendra was a managed retrieval system, Bedrock Agents was an orchestration system, and Q Business was an application over a corpus, and none of the three needed hardware.</p><p style="text-align: justify;">The landlord position has a second advantage. It&#8217;s indifferent to which tenant wins: capacity bills identically whether the customer runs Claude, Nova, Qwen, or something that does not exist yet, while every layer AWS dropped required backing a horse.</p><p>This also explains why Nova continues to exist, which is otherwise the loose thread in the argument. A landlord that buys its headline product from a single supplier needs a cap on what that supplier can charge. A competent house brand at the commodity tier does that job without winning anything: it sets a price floor, absorbs workloads where frontier capability is wasted, and keeps a credible team on the payroll.</p><div class="pullquote"><p>Nova doesn&#8217;t have to beat Claude. It has to make Claude&#8217;s price negotiable.</p></div><p style="text-align: justify;">And it explains why AWS now hosts its competitors' tooling rather than fighting it. Managed MLflow occupies the ground Clarify and Experiments were built to hold, and the product page that once sold Experiments now sells MLflow.[30][31][32] Experiments itself was never formally retired. The APIs still answer, and no deprecation notice exists anywhere; AWS confined it to the legacy Studio experience and pointed the documentation at its replacement.[33] A product superseded in place leaves no notice to audit. A week after freezing Model Monitor, the AWS machine learning blog published monitoring guidance built on open-source Evidently, with the results organized in MLflow.[34] Hosting the winner captures the workload without financing the losing race.</p><h2>The only one of the four</h2><p style="text-align: justify;">Set Amazon against its peers, and the position stops looking like an industry trend and starts looking like a choice. This is what the four say in public, not what they plan in private.</p><p style="text-align: justify;">Google builds the whole stack: its own accelerators, models, and surfaces, and it carries the research cost because the models feed products it owns end to end. Microsoft resells third-party models on Azure, as AWS does, but it also sells an assistant, and the quality of an assistant is a claim about the model inside it. Meta is the cautionary case: it spent enormously on infrastructure and could not buy the research culture to make it pay, which is the failure version of stepping back and the version the market reaches for by default.</p><p style="text-align: justify;">Amazon is none of these, and the distinction is less the direction than the decisiveness. Microsoft may yet drift the same way; its Azure catalog is already multi-model. Only Amazon has done it explicitly and completely, on one page, in one quarter. It rebuilt its intelligence tooling for the new architecture, shipped it, then withdrew, while raising capital expenditure toward $200bn, contracting a frontier supplier onto its own silicon for a decade, and committing a billion dollars to engineers who deploy that supplier&#8217;s models inside customer businesses. </p><p style="text-align: justify;">These are the actions of a company that priced the layer and declined it. The distinction carries a forecast. A company that lost a layer re-enters it when conditions improve, and its retreat is temporary. A company that priced a layer and declined it doesn&#8217;t come back, and here the barrier to return is one Amazon built itself: re-entering the frontier would now mean competing against its own anchor tenant, on its own silicon, against a product its own billion-dollar deployment force installs. The supply contract works as more than a substitute for the laboratory. It locks the door behind it. Everything since December 2025 points the same way: a reporting line that subordinates models to infrastructure, a contract measured in gigawatts and decades, and a billion dollars aimed at deployment rather than research.</p><h2>Redundancy, not retreat</h2><p>Which makes the AGI Lab closure something other than a failure signal.</p><p style="text-align: justify;">A laboratory does two things. It secures a capability you cannot otherwise obtain, and it holds an option on architectures nobody has commercialized yet. Amazon has bought the first through a 10-year supply contract using chips it designed. It has not closed the second, at least for now: Abbeel remains, research continues, and the AGI organization is still hiring.</p><p style="text-align: justify;">What Amazon canceled sits between the two, and that is why the reported cuts to model customization and post-training matter more than the closure of a site. Post-training is how a base model becomes a frontier product. Research is how you hold the option and judge what your supplier is selling. And the reportedly cut roles are the same functions Nova Forge now sells to customers as a service: continued pre-training, fine-tuning, and preference optimization on their own data.[34] The cuts and the product are one move. Amazon kept the option, stopped performing the frontier work for itself, and externalized it to paying customers.</p><p style="text-align: justify;">The exposures fit on one list. First, Amazon has contracted for output, not acquired a capability: Anthropic sets its own roadmap, prices its own models, serves Amazon&#8217;s competitors elsewhere, and runs a deliberately multi-cloud posture. The dependence runs both ways: since Amazon holds the equity, the training capacity, and the largest distribution channel, the mutual dependence remains stable until one side&#8217;s alternatives improve faster than the other&#8217;s. Second, the arrangement is subject to competition law. The UK Competition and Markets Authority examined the partnership in 2024 and closed its inquiry on jurisdiction, not merits, while the Federal Trade Commission ran a parallel market study.[35][36] The facts have since outgrown that clearance: a $4bn investment one way now sits beside a $100bn procurement the other, binding the two companies more tightly than anything the 2024 inquiry reviewed. Whether an authority would reach the same conclusion today is unknowable.</p><p style="text-align: justify;">Amazon has evidently judged the whole list cheaper than financing a laboratory it was unlikely to win with. That judgment is defensible on today&#8217;s numbers.</p><h2>What would break this</h2><p style="text-align: justify;">The strongest counterargument is that Amazon is still building models, and it is. Nova 2 shipped at re:Invent in December 2025 in Lite, Pro, Sonic, and Omni variants, alongside Nova Forge for organizations building custom frontier models.[37] An AWS executive would argue, with some justification, that a portfolio cull inside a funded model program is ordinary discipline rather than a change of direction.</p><p style="text-align: justify;">Except that the executive who runs the program has already conceded the premise. DeSantis told CNBC in June that it is &#8220;a fair narrative&#8221; that Amazon&#8217;s models &#8220;haven&#8217;t been at the very frontier&#8221; for the largest workloads, while hoping Amazon would be in the leading-model conversation in the coming year.[38] In the same interview, he put Nova 2 at roughly 50,000 customers since December, said Amazon should be considered on par with Nvidia in its ability to design and produce chips, and confirmed there's no timeline for Jassy&#8217;s April suggestion that Amazon might sell Trainium racks to third parties. Fifty thousand customers is breadth at the commodity tier, which is what a price floor accumulates. Comparing yourself to Nvidia is what a silicon company does. And floating rack sales are a landlord wondering whether to sell the fittings too. The counterargument&#8217;s own spokesman is describing the thesis.</p><p style="text-align: justify;">The thesis breaks if:</p><ul><li><p style="text-align: justify;">Nova receives a genuine frontier push, and a Nova model competes at the top of the Bedrock catalog on customer usage rather than vendor benchmarks.</p></li><li><p style="text-align: justify;">Nova Forge lands anchor customers at scale.</p></li><li><p style="text-align: justify;">The AGI organization is rebuilt outside the infrastructure reporting line, because that reporting line is the clearest signal in the whole sequence.</p></li><li><p style="text-align: justify;">AWS reclaims one of the displaced tooling categories with a proprietary product that wins against the open-source incumbent.</p></li><li><p style="text-align: justify;">Anthropic&#8217;s non-AWS capacity share grows materially, because that would suggest the option value in this arrangement always sat with the supplier, and that Amazon was priced into the landlord role rather than choosing it.</p></li></ul><p style="text-align: justify;">The second-quarter results on 30 July are the near-term test: watch capital expenditure, AWS margin, and whether Nova appears in the earnings narrative at all.</p><h2>What this means for the people buying it</h2><p style="text-align: justify;">For a buyer, this is one question. Before adopting any vendor AI feature, ask whether the function could be built by someone who does not own a data center. If it could, price in migration, and keep the integration thin. If it cannot, treat the dependency as infrastructure and negotiate exit assumptions over years. And do not audit vendor risk off deprecation notices alone: Experiments show a product can be retired by a documentation edit without appearing on any list.</p><p style="text-align: justify;">Five years ago, I argued that software would eat machine learning: teams would assemble models instead of building them, write as little code as possible, and let tooling do the rest.[39] The June 30 page is that prediction arriving. </p><div class="pullquote"><p style="text-align: center;">Everything in the middle of the stack that could be expressed as software has been eaten, mostly by tooling that ships with the model.</p></div><p style="text-align: justify;">What survives is what software cannot eat: the chips, the buildings, the power, and the people who install the result. Amazon saw that outcome and reorganized around it. Retire what software already ate. Spend a billion dollars deploying what it cannot.</p><p style="text-align: justify;">The durable position was never the model. Nor was it the middle of the stack, where services competed with superior open source on launch day and existed mostly to thicken a product manager&#8217;s promotion document. </p><p style="text-align: justify;">Back to cloud 101: the durable position is always the meter.</p><div><hr></div><h3>Notes</h3><p>[1] <a href="https://www.aboutamazon.com/news/aws/aws-1-billion-forward-deployed-ai-engineers">AWS invests $1 billion to embed AI forward deployed engineers with customers</a>, About Amazon, 30 June 2026. Describes the agentic-first model, the compression of deployment timelines, the customer self-sufficiency goal, and the named early customers.</p><p>[2] <a href="https://www.cnbc.com/2026/06/30/aws-amazon-ai-forward-deployed-engineers.html">AWS puts $1 billion into new AI unit to embed engineers with customers</a>, CNBC, 30 June 2026. Source for pod size, engagement length and the Vasquez interview.</p><p>[3] <a href="https://aws.amazon.com/about-aws/whats-new/2026/06/aws-service-availability/">AWS Service Availability Updates</a>, AWS What&#8217;s New, 30 June 2026. The ten SageMaker AI features are A2I, Clarify, Debugger, GeoSpatial, Ground Truth, Mechanical Turk, Model Monitor, Profiler, Role Manager and Studio Lab. Ground Truth Plus is listed separately under end of support. The same page also retires directory, mainframe and console-management services unrelated to AI; the body&#8217;s &#8220;consequential&#8221; excludes those and legacy items such as Mechanical Turk and Studio Lab, which reached maintenance through obsolescence rather than displacement.</p><p>[4] <a href="https://docs.aws.amazon.com/kendra/latest/dg/kendra-availability-change.html">Amazon Kendra availability change</a>, AWS documentation.</p><p>[5] <a href="https://docs.aws.amazon.com/amazonq/latest/qbusiness-ug/qbusiness-availability-change.html">Amazon Q Business availability change</a>, AWS documentation.</p><p>[6] <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/agents-classic-maintenance-mode.html">Amazon Bedrock Agents Classic maintenance mode</a>, AWS documentation. Bedrock models, Knowledge Bases and Guardrails are unaffected; migration is directed to AgentCore.</p><p>[7] <a href="https://www.cnbc.com/2026/07/22/amazon-lays-off-some-employees-in-its-agi-unit.html">Amazon lays off some employees in its AGI unit</a>, CNBC, 22 July 2026. Source for the composition of the AGI organisation, the DeSantis appointment, and the company statement.</p><p>[8] <a href="https://www.geekwire.com/2026/amazon-confirms-its-closing-key-ai-site-in-san-francisco-but-says-work-on-its-top-models-continues/">Amazon confirms it&#8217;s closing key AI site in San Francisco but says work on its top models continues</a>, GeekWire, 24 July 2026. Company spokesperson confirmation of the site closure, the lab&#8217;s founding and headcount, the Adept departures, and Abbeel&#8217;s role.</p><p>[9] <a href="https://ir.aboutamazon.com/news-release/news-release-details/2026/Amazon-com-Announces-Fourth-Quarter-Results/">Amazon.com Announces Fourth Quarter Results</a>, Amazon Investor Relations, 5 February 2026. The release carries Jassy&#8217;s guidance of about $200bn in 2026 capital expenditure and the &#8220;seminal opportunities&#8221; framing; his statement that the spend would be predominantly invested in AWS was made to analysts on the accompanying call, per <a href="https://www.cnbc.com/2026/02/05/amazon-amzn-q4-earnings-report-2025.html">CNBC&#8217;s same-day report</a>.</p><p>[10] <a href="https://www.aboutamazon.com/news/aws/aws-project-rainier-ai-trainium-chips-compute-cluster">AWS&#8217;s Project Rainier: the world&#8217;s most powerful computer</a>, About Amazon, 29 October 2025.</p><p>[11] <a href="https://aws.amazon.com/ec2/instance-types/trn3/">Amazon EC2 Trn3 UltraServers</a>, AWS product page. Vendor-published performance claims; not independently verified.</p><p>[12] <a href="https://www.anthropic.com/news/anthropic-amazon-compute">Anthropic and Amazon expand collaboration for up to 5 gigawatts of new compute</a>, Anthropic, 20 April 2026; Amazon&#8217;s counterpart release is <a href="https://www.aboutamazon.com/news/company-news/amazon-invests-additional-5-billion-anthropic-ai">Amazon announces $5B Anthropic investment, up to $20B more</a>, About Amazon, 20 April 2026. Source for the $100bn ten-year commitment, the up-to-5GW capacity, and the Trainium2-through-Trainium4 scope.</p><p>[13] <a href="https://ir.aboutamazon.com/news-release/news-release-details/2026/Amazon-com-Announces-First-Quarter-Results/default.aspx">Amazon.com Announces First Quarter Results</a>, Amazon Investor Relations, 29 April 2026. Source for the $16.8bn of pre-tax gains booked in non-operating income from the Anthropic investments and for AWS segment sales up 28% to $37.6bn. The annualised AI services revenue figure is from management commentary on the accompanying earnings call.</p><p>[14] <a href="https://www.techtimes.com/articles/321341/20260723/amazon-cuts-agi-jobs-while-pouring-200-billion-ai-infrastructure.htm">Amazon Cuts AGI Jobs While Pouring $200 Billion Into AI Infrastructure</a>, TechTimes, 23 July 2026. Secondary source for the Bank of America estimate and for the reported composition of the cuts; both should be treated as reported rather than confirmed, and the Bank of America figure is an analyst estimate.</p><p>[15] Representative Amazon Bedrock on-demand rates, <a href="https://aws.amazon.com/bedrock/pricing/">aws.amazon.com/bedrock/pricing</a>, retrieved July 2026; parity with Anthropic&#8217;s direct API per <a href="https://platform.claude.com/docs/en/about-claude/pricing">Anthropic platform pricing documentation</a>. Claude Sonnet at $3/$15 per million input/output tokens on both, with matching promotional windows. Rates as of July 2026 and subject to change.</p><p>[16] <a href="https://www.manilatimes.net/2026/07/02/business/foreign-business/amazons-aws-commits-1-billion-toward-new-unit-for-embedded-ai-engineers/2376965">Amazon&#8217;s AWS commits $1 billion toward new unit for embedded AI engineers</a>, The Manila Times, 2 July 2026, on prior forward-deployed engineering organisations at Palantir, Salesforce, Google Cloud and Anthropic.</p><p>[17] Reported December 2025; see note 7 for the CNBC account of the appointment and the scope of the resulting organisation.</p><p>[18] <a href="https://www.theinformation.com/briefings/amazon-shuts-ai-agent-research-lab-agi-layoffs">Amazon Shuts AI Agent Research Lab In AGI Layoffs</a>, The Information, July 2026. Paywalled; the description of the lab&#8217;s remit is quoted via GeekWire at note 8.</p><p>[19] <a href="https://thenextweb.com/news/amazon-shuts-agi-lab-frontier-model-retreat-layoffs">Amazon shuts its AGI Lab in fresh AI layoffs</a>, The Next Web, July 2026.</p><p>[20] <a href="https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-sagemaker-clarify-fm-evaluations-preview">Amazon SageMaker Clarify now supports foundation model (FM) evaluations in preview</a>, AWS What&#8217;s New, 29 November 2023.</p><p>[21] <a href="https://aws.amazon.com/about-aws/whats-new/2024/04/amazon-sagemaker-clarify-foundation-model-evaluations">Amazon SageMaker Clarify now supports foundation model evaluations</a>, AWS What&#8217;s New, 25 April 2024. General availability excluded GovCloud, China and several commercial regions listed in the announcement.</p><p>[22] <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/clarify-availability-change.html">Amazon SageMaker Clarify availability change</a>, AWS documentation.</p><p>[23] <a href="https://aws.amazon.com/blogs/machine-learning/power-your-llm-training-and-evaluation-with-the-new-sagemaker-ai-generative-ai-tools">Power Your LLM Training and Evaluation with the New SageMaker AI Generative AI Tools</a>, AWS Artificial Intelligence blog.</p><p>[24] <a href="https://aws.amazon.com/about-aws/whats-new/2023/05/sagemaker-ground-truth-plus-human-feedback-fine-tuning-data-generative-ai">Amazon SageMaker Ground Truth Plus now supports human feedback and fine-tuning data for Generative AI</a>, AWS What&#8217;s New, 30 May 2023.</p><p>[25] <a href="https://aws.amazon.com/about-aws/whats-new/2026/04/sagemaker-hyperpod-gang-scheduling/">SageMaker HyperPod now supports gang scheduling for distributed training workloads</a>, AWS What&#8217;s New, 8 April 2026.</p><p>[26] <a href="https://aws.amazon.com/about-aws/whats-new/2026/03/amazon-sagemaker-hyperpod-continuous-provisioning/">Amazon SageMaker HyperPod now supports continuous provisioning for Slurm-orchestrated clusters</a>, AWS What&#8217;s New, 25 March 2026.</p><p>[27] <a href="https://aws.amazon.com/blogs/machine-learning/accelerate-large-scale-ai-training-with-amazon-sagemaker-hyperpod-training-operator/">Accelerate large-scale AI training with Amazon SageMaker HyperPod training operator</a>, AWS Artificial Intelligence blog. The HyperPod elastic agent is described as an extension of PyTorch&#8217;s ElasticAgent, and fault detection draws on node health checks and AWS retirement notices.</p><p>[28] <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-release-notes.html">Amazon SageMaker HyperPod release notes</a>, AWS documentation, recording Trn2 and Trn2n support for Slurm and Amazon EKS clusters.</p><p>[29] <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/reserve-capacity-with-training-plans.html">Reserve Flexible Training Plans for ML workloads</a>, AWS documentation.</p><p>[30] <a href="https://aws.amazon.com/about-aws/whats-new/2024/06/amazon-sagemaker-mlflow-capability">Amazon SageMaker now offers a fully managed MLflow capability</a>, AWS What&#8217;s New, 19 June 2024.</p><p>[31] <a href="https://aws.amazon.com/about-aws/whats-new/2025/07/fully-managed-mlflow-3-0-amazon-sagemaker-ai/">Fully managed MLflow 3.0 now available on Amazon SageMaker AI</a>, AWS What&#8217;s New, July 2025.</p><p>[32] <a href="https://aws.amazon.com/sagemaker/ai/experiments">Accelerate generative AI development with Amazon SageMaker AI and MLflow</a>, AWS product page, retrieved July 2026.</p><p>[33] <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/experiments.html">Amazon SageMaker Experiments in Studio Classic</a>, AWS documentation.</p><p>[34] The overlap between the reportedly cut post-training and customisation roles and the capabilities Nova Forge sells (continued pre-training, supervised fine-tuning, direct preference optimisation on customer data) is drawn from the TechTimes report at note 14 and from Amazon&#8217;s Nova Forge description at note 37; the role composition of the cuts remains unconfirmed by Amazon.</p><p>[35] <a href="https://gov.uk/cma-cases/amazon-slash-anthropic-partnership-merger-inquiry">Amazon / Anthropic partnership merger inquiry</a>, Competition and Markets Authority case page. Merger inquiry launched 8 August 2024; phase 1 decision 27 September 2024 that the partnership does not qualify for investigation under the merger provisions of the Enterprise Act 2002.</p><p>[36] <a href="https://www.ftc.gov/news-events/news/press-releases/2024/01/ftc-launches-inquiry-generative-ai-investments-partnerships">FTC Launches Inquiry into Generative AI Investments and Partnerships</a>, Federal Trade Commission, 25 January 2024. Section 6(b) orders issued to the companies involved in the Microsoft-OpenAI, Amazon-Anthropic and Google-Anthropic investments.</p><p>[37] <a href="https://www.aboutamazon.com/news/aws/aws-re-invent-2025-ai-news-updates">AWS re:Invent 2025: Amazon announces Nova 2, Trainium3, frontier agents</a>, About Amazon, December 2025. Covers the Nova 2 family (Lite, Pro, Sonic, Omni), Nova Forge&#8217;s open-training checkpoint model, and Nova Act&#8217;s move to general availability; announcements made 2 December 2025.</p><p>[38] <a href="https://www.cnbc.com/2026/06/17/amazon-ai-frontier-openai-anthropic.html">Amazon has lagged OpenAI and Anthropic, but AI chief sees path to catch up in &#8216;coming year&#8217;</a>, CNBC, 17 June 2026. Direct interview with Peter DeSantis; source for the frontier concession, the Nova 2 customer figure, the Nvidia comparison, and the status of potential Trainium rack sales first suggested by Andy Jassy in April.</p><p>[39] <a href="https://huggingface.co/blog/the-age-of-ml-as-code">The Age of Machine Learning As Code Has Arrived</a>, Hugging Face blog, 20 October 2021, which I co-authored while Chief Evangelist at Hugging Face.</p>]]></content:encoded></item><item><title><![CDATA[The Model Is a Checkpoint]]></title><description><![CDATA[Washington spent last week letting it be known it might ban Chinese AI models, so this week, everyone is an open-models expert.]]></description><link>https://www.airealist.ai/p/the-model-is-a-checkpoint</link><guid isPermaLink="false">https://www.airealist.ai/p/the-model-is-a-checkpoint</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Thu, 23 Jul 2026 11:06:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CPhA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CPhA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CPhA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!CPhA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!CPhA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!CPhA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CPhA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1941764,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/208101192?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CPhA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!CPhA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!CPhA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!CPhA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e081b64-b147-4406-90aa-fe76f07b2604_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">Following the release of Kimi K3, Axios reported that the administration is reviving a push to restrict Chinese models on cybersecurity grounds, with Entity List designations under discussion and House committees probing the American companies that run them.[1] The explainer wave arrived within the news cycle: what open weights are, why enterprises quietly run them, which Chinese models to worry about, and the same comparison table in a hundred fonts. Much of it reads as if the subject were discovered at roughly the moment it hit 46.4% of routed traffic on OpenRouter.[2] Benchmarks, market share, comparison: the wave runs on one axis, and it is the wrong one.</p><p style="text-align: justify;">Being late is forgivable; I would rather people learn about open models now than never. The problem is that it is not just late; it is describing an object from an era that ended eighteen months ago. So let me set the record straight, from the beginning, because I was there for it.</p><p style="text-align: justify;">In March 2021, I published a post on the AWS machine learning blog announcing a partnership between Amazon and a startup most enterprise readers had never heard of. Hugging Face, founded in 2016 and with offices in New York and Paris, made it easy to add Transformer models to your applications. The catalog held about 7,000 pre-trained models in 164 languages. The examples I reached for were BERT (340 million parameters) and GPT (175 billion).[3]</p><p style="text-align: justify;">Read that post today, and the striking thing is what it does not contain. There is no comparison. No table weighing an open model against a proprietary API, no paragraph on cost per token, nothing on privacy or lock-in. There was nothing to compare against. The problem the partnership existed to solve was more basic: most organizations could not deploy these models at all. Getting the file onto managed infrastructure with a few lines of code was the entire value proposition.</p><h2>Three eras</h2><p>Open models have changed what they are three times, and each era is defined by the question practitioners were asking.</p><p style="text-align: justify;">In the distribution era, roughly 2021 to 2022, the question was: <strong>how do I run this at all?</strong> The achievement was deployment. The infrastructure that mattered was hubs, containers, and managed endpoints, and the partnership I announced was that infrastructure being built.</p><p style="text-align: justify;">ChatGPT ended the first era in a weekend. Deployment stopped being the hard part because the incumbent had arrived, and in the comparison era, 2023 to 2024, the question became: <strong>Why would I use this instead of the API?</strong> Open models had a proprietary incumbent to displace, so every argument became a table: transparency versus opacity, cost versus convenience, control versus capability. I spent two years making that argument, and I will show you my own exhibit below.</p><p style="text-align: justify;">The third era had been building quietly since 2023, in preference-tuning papers and merge experiments, while the comparison arguments still held the stage. It blew into public view in January 2025, when DeepSeek shipped R1, trained with a reinforcement learning algorithm the lab published, alongside a family of smaller models distilled from it. In the post-training era, the question changed again, and most people have not noticed, because the era&#8217;s arrival, like the first era&#8217;s ending, was a visibility event: the work predated the moment. <strong>The question is no longer which model to pick. It is what to make from the models you hold.</strong> Open models stopped being products that compete with closed ones and became checkpoints: fine-tuned, aligned, distilled, merged, transplanted, and quantized, both inside the labs that release them and on hardware that fits on a desk. Training no longer ends where the download begins.</p><div class="pullquote"><p>Open models are not a substitute for closed ones. They are a different kind of object, and every era of their history has been a discovery of what that object permits.</p></div><p style="text-align: justify;">For the record: I was at AWS during distribution, at Hugging Face during comparison, and at Arcee AI, the company behind the MergeKit toolkit, during post-training. My career is how I noticed the eras. It is not the evidence, which is other people&#8217;s published work throughout.</p><p style="text-align: justify;">This piece is for teams with the engineering capacity to act on that. If your organization has no machine learning function, the hosted service is the right purchase, and nothing below changes it. But if you are arriving at open models now, through the current wave of explainers and best-of lists, you are being taught the second era&#8217;s answer to the third era&#8217;s question. You should know what this era looks like, from someone who watched all three arrive.</p><h2>The comparison era, from inside</h2><p style="text-align: justify;">In September 2023, as Chief Evangelist at Hugging Face, I published a comparison of SafeCoder, our enterprise code assistant, against the closed-source alternatives.[4] It made five arguments. The model was state-of-the-art. It was transparent: you could read the paper, inspect the training dataset, and check whether your code was in it, since the dataset shipped with an opt-out tool for repository owners.[5] You could fine-tune it yourself in your code. You could run it anywhere, including in an air-gapped environment. And it phoned nothing home.</p><p style="text-align: justify;">I stand by every line, and the post has aged into a period piece anyway. Not because the claims failed, but because the <em>form</em> dates it. Five properties, each argued against a closed-source counterpart. The piece is a table wearing prose, and the table was the era. Open models were the challenger, the incumbent set the terms, and everything we wrote had to justify the open choice against the closed default.</p><p style="text-align: justify;">The comparison era did real work. It established that open models were viable for serious deployment, it forced pricing discipline on the incumbents, and its arguments about privacy and control were correct and remain correct. What it could not do was describe properties with no counterpart on the other side of the table. A comparison can only hold what both columns share. The operations that define this era were invisible to the format: not suppressed, just unrepresentable.</p><p style="text-align: justify;">Which is why, if you learned open models from the comparison literature, the current era will sound like it happened somewhere else. It did not. It happened inside the models you are being recommended.</p><h2>The post-training era: where the models come from</h2><p style="text-align: justify;">Here is what the era&#8217;s best-of lists do not say about the models on them: nearly everything that makes them good happened after pre-training. Supervised fine-tuning, preference training, reinforcement learning, distillation from a larger teacher, and increasingly merging. The base run buys the raw capability; the post-training stack turns it into the model you download.</p><p style="text-align: justify;">The stack starts earlier than most descriptions admit. Continued pre-training, sometimes called mid-training, takes a released checkpoint and keeps feeding it raw text: Code Llama is Llama 2 with its pre-training continued on a code-heavy corpus, and SEA-LION&#8217;s pipeline begins by continuing pre-training on Southeast Asian text before any of the stages below.[6] A checkpoint is not a finished artefact with a tuning knob. It is a training run; somebody paused, and anyone holding the file can press resume.</p><p style="text-align: justify;">The alignment layer opened first, and the progression is worth spelling out because it is the era in miniature. In 2022, preference training meant RLHF as the closed labs ran it: proximal policy optimization, a learned reward model, a critic network, and infrastructure that almost nobody else could operate. In 2023, direct preference optimization collapsed all of that into a single loss function that runs wherever fine-tuning runs, and Zephyr showed a 7-billion-parameter open model beating much larger chat models on supervised fine-tuning plus DPO alone.[7] In late 2024, AI2&#8217;s T&#252;lu 3 published a complete open post-training recipe, including data, code, and weights, and introduced reinforcement learning from verifiable rewards: checkers instead of learned reward models.[8]</p><p style="text-align: justify;">Then, DeepSeek&#8217;s GRPO (group relative policy optimization), the algorithm behind R1, deleted the critic network that made the RL loop expensive, and within weeks of R1&#8217;s release, the open community was replicating the pipeline publicly.[9] Each step removed a piece of infrastructure only a frontier lab could afford, until the whole alignment stack ran in open frameworks, with the small-model end reaching consumer GPUs.[10]</p><p style="text-align: justify;">Merging is the newest layer of that stack, and it is already inside the models the guides recommend. Meta&#8217;s Llama 3.1 report describes averaging models at every post-training stage across reward modeling, supervised fine-tuning, and DPO, and using the flagship to improve its smaller siblings.[11] Cohere&#8217;s Command: A technical report calls expert merging a core feature of its training pipeline: domain experts trained separately, then folded into a single set of parameters.[12]</p><p style="text-align: justify;">Much of that is averaging variants of a single pipeline. Merging separately trained models into a single set of weights was adopted from papers and released systems in 2025.[13] In March, a team from Qiyuan Tech and Peking University distilled DeepSeek-R1 into three domain experts for mathematics, coding, and science, then merged them with Arcee Fusion, from Arcee AI's MergeKit toolkit, landing within two points of the 671-billion-parameter teacher on AIME mathematics. The merge took 4 GPU-hours, compared with 740 for retraining on mixed data, which the authors call a free lunch, published from the far side of the boundary the ban debate is trying to draw. SEA-LION, AI Singapore's national language family, runs multiple DELLA Linear merge stages in the pipeline behind its shipped models.[16] And the field is not fringe: merging has a survey in ACM Computing Surveys, and the toolkit's paper sits in the EMNLP industry track.[17][18]</p><p style="text-align: justify;">So the technical reports describe merging as routine, the lists recommend the models that those reports document, and the word appears in neither. A newcomer can read guide after guide and never learn that the operation exists, let alone that it built the thing they just downloaded.</p><h2>The post-training era: on your desk</h2><p style="text-align: justify;">The same operations run downstream, and the hardware to run them arrived recently enough that most of the literature predates it.</p><p style="text-align: justify;">One calibration first, because the era has a gradient, and you should know where you stand on it. Quantization and distillation are the mass operations: anyone running a model through Ollama is running quantized weights, today, probably without thinking of it as a weight-space operation at all. Fine-tuning is common. Merging and surgery are where the frontier is, not where the median is. The era is defined not by everyone doing everything, but by the possibility of it all with a single object.</p><p style="text-align: justify;">The practitioner&#8217;s toolkit is six operations, and the everyday ones come first. </p><ol><li><p style="text-align: justify;"><strong>Fine-tune a model on your own data</strong>: low-rank adaptation runs on a single consumer GPU, and the output is an adapter file measured in megabytes that you keep, copy, and load anywhere. </p></li><li><p style="text-align: justify;"><strong>Distil a large open model into a small one whose weights you own</strong>; an open teacher is a teacher you are allowed to keep, where a closed vendor&#8217;s terms typically forbid training a competitor on its outputs. NVIDIA went further and released the Nemotron-4 340B family under a licence that explicitly permits generating synthetic training data, because data generation is now a post-training stage in its own right.[19] </p></li><li><p style="text-align: justify;"><strong>Align it</strong>: DPO runs wherever fine-tuning runs, and GRPO-style reinforcement learning, critic deleted and rewards drawn from verifiable checkers, now runs in open frameworks on consumer GPUs; I examined the compute shape of that workload in <a href="https://www.airealist.ai/p/the-verification-tax">The Verification Tax</a>.</p></li><li><p style="text-align: justify;"><strong>Merge models into a new one.</strong> That operation needs no training run, no rented cluster, and no data; for small models, it runs on a CPU or eight gigabytes of video memory, and lazy tensor loading keeps even large merges on modest hardware.[20] </p></li><li><p style="text-align: justify;"><strong>Transplant components between models</strong>: MergeKit ships a tool that moves one model&#8217;s tokenizer into another, and in its published cross-family test, the underlying technique retained roughly 96% of language-understanding performance with no training at all.[21] </p></li><li><p style="text-align: justify;"><strong>Quantize</strong>: reduce numerical precision to trade quality for footprint at a point you choose.</p></li></ol><p style="text-align: justify;">Quantization is the bridge between owning weights and using them, and its arithmetic is blunt. Qwen3.5-397B, a 397-billion-parameter mixture-of-experts model from one of the Chinese families reportedly under discussion in Washington, is roughly 807 gigabytes on disk at full precision: a rack, full stop. At dynamic four-bit, it is about 214 gigabytes and loads on a single Mac Studio. At three bits, it fits a 192-gigabyte machine, and the quantizer&#8217;s own guidance now recommends two-bit dynamic variants as a legitimate operating point rather than a last resort.[22] The same weights span a data center and a desk, and where on that curve you sit is your decision, not a vendor&#8217;s. The community quant tables even publish the far end with honest labels: the one-bit files are marked &#8220;for the desperate&#8221;.[23] That is what a continuous curve looks like: it goes all the way down, and it tells you the quality price at every stop. And the desk changed too.</p><p style="text-align: justify;">Consumer graphics cards stop at 24 or 32 gigabytes of memory, and a model that does not fit does not run.[24] Unified memory architectures remove that ceiling by pooling system and graphics memory. Apple established the pattern. AMD&#8217;s Strix Halo platform puts 128 gigabytes of unified memory in mini-PCs selling for roughly $2,000, with AMD&#8217;s own developer machine at $3,999, and Nvidia sells its own 128-gigabyte unified-memory developer box, the DGX Spark, at the same price. The company unveiled the platform at CES in January 2026, claiming it could run models with up to 200 billion parameters locally.[25] The trade is bandwidth: a discrete card moves memory several times faster, which is what generation speed depends on, so small models run faster on the card while large ones only run on the pool. Mixture-of-experts models bend the curve further by activating only a fraction of their parameters per token, letting a far sparser model outrun a dense one on the same machine.[26]</p><p style="text-align: justify;">And weight surgery is quietly becoming a production dependency, not an enthusiast&#8217;s hobby. Speculative decoding, the serving optimization in which a small draft model runs ahead of a large one, is recommended in current deployment guides and implemented in vLLM, which, by default, requires that the draft and the target share a vocabulary.[27] For a model with its own vocabulary, no compatible draft exists, and building one traditionally means returning to the training stage.[28] The cheap route is the transplant operation above; the heavier one is training a small drafting head, EAGLE-style, directly onto the frozen checkpoint, which is likewise an operation on weights you must hold.[29] Microsoft Foundry&#8217;s import screen now has a dedicated slot for draft models.[30] The platform assumes you arrive holding one. None of the guides I read mentions that the slot exists or that weight surgery is how it's filled.</p><h2>Where the checkpoint beats the frontier</h2><p style="text-align: justify;">The frontier labs post-train for everyone on average. A checkpoint post-trained for one domain beats them on that domain, and the clearest demonstrations now come from companies you know.</p><p style="text-align: justify;">Cursor, one of the companies the House is reportedly asking about, built Composer, the model behind its agent, by training a mixture-of-experts model via reinforcement learning across hundreds of thousands of sandboxed coding environments. According to the company&#8217;s benchmarks, it achieves frontier coding results at 4x the generation speed of comparably intelligent models.[32] The base is widely reported to be an open checkpoint, a question Cursor&#8217;s researchers pointedly decline to answer, which tells you what a checkpoint&#8217;s provenance is now worth.[33]</p><p style="text-align: justify;">Salesforce fine-tuned its xLAM family on synthesized tool-use data and took first place on the Berkeley Function-Calling Leaderboard, ahead of GPT-4 and Claude 3, with models a fraction of their size.[34]</p><p style="text-align: justify;">And Baichuan post-trained a 32-billion-parameter medical model using a staged GRPO recipe against a clinical verifier. On HealthBench, OpenAI&#8217;s own medical benchmark, M2 beat every open model and most closed ones, including o3 and Gemini 2.5 Pro, and at release, it was the only model other than GPT-5 to clear 32 on the Hard subset. Quantized, it deploys on a single RTX 4090.[35]</p><p style="text-align: justify;">Three domains, three of the six operations, three names your board recognizes. None of these teams out-trained the frontier in general, and none needed to. They picked an axis, fine-tuned a checkpoint on it, and kept the axis.</p><h2>Why the literature lags</h2><p style="text-align: justify;">If the post-training era is real, why is the &#8220;Open Models 101&#8221; wave still writing comparison-era explainers? Mostly because that is what explanation is. The genre exists to introduce open models to people arriving from closed APIs, and introducing the unfamiliar means comparing it to the familiar. A comparison can only hold properties both sides share: price, privacy, control, and latency each take a value on both sides, so each gets a row. Continued pre-training, merging, and weight surgery have no closed-form counterparts, so no rows exist for them, and they go unwritten. The blog posts, the press explainers, and the analyst notes all inherit the format, whoever writes them.</p><p style="text-align: justify;">The commercial guides could break the frame, because their authors run these operations for a living. They do not, and the pattern is instructive. I read four guides closely: an inference platform, a GPU renter, a desktop-app maker, and an independent author. Each goes deep on exactly one production capability, the one it sells. Preference optimization and RL fine-tuning appear in none of them, and neither does merging, which nobody sells.[36] Merging needs no rented compute, no serving contract, no application. The comparison frame explains the wave. The billing pattern explains why even its deep end never corrects its shallow end.</p><p style="text-align: justify;">The obvious objection is that if these operations mattered, the closed vendors would sell them. They did. OpenAI shipped a complete distillation workflow in its API in October 2024, and Google, Anthropic, and AWS offer their own distillation paths.[37] What they sell is the operation, never the artifact. OpenAI now even releases open-weight models of its own; what it does not release is the adapted weights from your fine-tuning of its hosted models. The fine-tune lives behind the API, and Microsoft&#8217;s documentation for the equivalent Azure workflow states plainly that the stored training data cannot be exported or downloaded.[38] </p><p style="text-align: justify;">The operations arrived. Nothing came out. And the hosted route delivers real results: published distillation numbers report 14-24 times lower cost per successful task than frontier calls.[39] What the buyer owns afterward is a result, not a model; it cannot be merged, quantized, or moved, and it lives only as long as the vendor&#8217;s base.</p><p style="text-align: justify;">The platforms noticed the same thing, and what they built is the clearest statement of what this era is worth. Amazon Bedrock ships Custom Model Import, generally available since October 2024, which accepts weights in the Hugging Face format: Llama, Mistral, Qwen, and, since last November, OpenAI&#8217;s own open-weight models.[40] Google Vertex accepts custom weights. Microsoft Foundry accepts full models, adapters, and draft models.</p><p style="text-align: justify;">Study those import screens for a moment. Every major platform built a door for the model file to come in. None of them built one facing out. You can bring a merged model to all three hyperscalers; you cannot take a fine-tune of their own models away from any of them. The companies that run the largest AI infrastructure in existence have priced what a file in hand is worth, and their answer is a one-way door.</p><p style="text-align: justify;">So the question to take away from this piece is not the question of the era being compared. Not which model is best, and not even open versus closed. The question is what you will make from the thing you hold. And it is the right question because the models are now made that way, the tools run on a desk, and the explainers who cannot tell you so are still writing inside a frame built for the era before.</p><p style="text-align: justify;">Washington, incidentally, is learning the same lesson from the opposite direction. The reported ban keeps colliding with the property this piece is about: the weights are files, already on disks in their millions, and legal analysts note that restricting publicly available model weights raises constitutional questions a service ban never would.[41] Whatever the security merits of the case, and the security findings are not trivial,[42] the policy difficulty is the object lesson. You can sanction a company and geo-block a service. A file that has shipped is neither, which is presumably why the instruments reportedly under discussion are Entity List designations, procurement rules, and liability: tools that reach companies and contracts, because nothing reaches the file. </p><p style="text-align: justify;">And I should say plainly that the capability this piece describes cuts both ways: the same weight access that lets a practitioner transplant a tokenizer lets anyone strip a model&#8217;s safety training, an operation I examined in <a href="https://www.airealist.ai/p/open-from-both-sides">Open From Both Sides</a>. Weight access is not a virtue. It is a property, and properties do not pick sides.</p><p>So here&#8217;s my one-line &#8220;open models 101&#8221;: An open model shouldn&#8217;t be a finished product that you pick. It should be a training run you resume.</p><div><hr></div><h3>Notes</h3><p>[1] Axios, <a href="https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi">&#8220;The secret Trump administration battle to fight Chinese AI&#8221;</a>, 20 July 2026; see also Tom&#8217;s Hardware, <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/trump-administration-reportedly-reviving-push-to-ban-chinese-ai-models-following-kimi-k3-launch-citing-cybersecurity-concerns-downloadable-open-weights-could-make-an-outright-u-s-ban-nearly-impossible-to-enforce-amid-growing-adoption">&#8220;Trump administration reportedly reviving push to ban Chinese AI models following Kimi K3 launch&#8221;</a>, 22 July 2026. No executive order signed as of this writing; the Commerce Department reportedly evaluated Entity List designations; House committee probes of US firms using Chinese models reported July 2026. All characterisations &#8220;reportedly&#8221; per the sourcing.</p><p>[2] Chinese models at 46.4% of routed token traffic on <a href="https://openrouter.ai/rankings">OpenRouter</a> versus 35.7% for US models, DeepSeek at 17.6% alone, as of July 2026, per the <a href="https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi">Axios</a> reporting and contemporaneous coverage. OpenRouter routes a developer-heavy slice of traffic; treat the shares as indicative of that population, not of the whole market.</p><p>[3] Simon, <a href="https://aws.amazon.com/blogs/machine-learning/aws-and-hugging-face-collaborate-to-simplify-and-accelerate-adoption-of-natural-language-processing-models/">&#8220;AWS and Hugging Face collaborate to simplify and accelerate adoption of Natural Language Processing models&#8221;</a>, AWS Machine Learning Blog, 23 March 2021. Figures as published: ~7,000 pre-trained models, 164 languages, BERT at 340M parameters, GPT at 175B.</p><p>[4] Simon, <a href="https://huggingface.co/blog/safecoder-vs-closed-source-code-assistants">&#8220;SafeCoder vs. Closed-source Code Assistants&#8221;</a>, Hugging Face blog, 11 September 2023.</p><p>[5] The StarCoder base models were trained on The Stack (2.7 TB, permissively licensed code), published with an opt-out mechanism for repository owners. Li et al., <a href="https://arxiv.org/abs/2305.06161">&#8220;StarCoder: may the source be with you!&#8221;</a>, arXiv:2305.06161.</p><p>[6] Rozi&#232;re et al., <a href="https://arxiv.org/abs/2308.12950">&#8220;Code Llama: Open Foundation Models for Code&#8221;</a>: Llama 2 with continued pre-training on a code-heavy corpus. SEA-LION&#8217;s continued pre-training stage per its technical report (see the SEA-LION note above). Continued pre-training at useful scale consumes billions of tokens of compute; it is the heaviest operation in this piece and the one with the highest floor.</p><p>[7] Rafailov et al., <a href="https://arxiv.org/abs/2305.18290">&#8220;Direct Preference Optimization: Your Language Model is Secretly a Reward Model&#8221;</a>, NeurIPS 2023; Tunstall et al., <a href="https://arxiv.org/abs/2310.16944">&#8220;Zephyr: Direct Distillation of LM Alignment&#8221;</a>, the Hugging Face H4 recipe pairing SFT with DPO.</p><p>[8] Lambert et al., <a href="https://arxiv.org/abs/2411.15124">&#8220;T&#252;lu 3: Pushing Frontiers in Open Language Model Post-Training&#8221;</a>, Allen Institute for AI, November 2024. Full recipe, data and weights released; introduces reinforcement learning with verifiable rewards (RLVR).</p><p>[9] GRPO: Shao et al., <a href="https://arxiv.org/abs/2402.03300">&#8220;DeepSeekMath&#8221;</a>, which introduced group relative policy optimisation; DeepSeek-AI, <a href="https://arxiv.org/abs/2501.12948">&#8220;DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning&#8221;</a>, January 2025. GRPO computes advantages within a sampled group, removing the separate critic network. Community replication: Hugging Face&#8217;s <a href="https://github.com/huggingface/open-r1">Open R1</a>.</p><p>[10] Open post-training frameworks include Hugging Face <a href="https://github.com/huggingface/trl">TRL</a> (SFT, DPO and GRPO trainers), ByteDance&#8217;s <a href="https://github.com/volcengine/verl">verl</a>, and <a href="https://github.com/OpenRLHF/OpenRLHF">OpenRLHF</a>; consumer-hardware GRPO recipes ship in <a href="https://docs.unsloth.ai">Unsloth</a>.</p><p>[11] Grattafiori et al., <a href="https://arxiv.org/abs/2407.21783">&#8220;The Llama 3 Herd of Models&#8221;</a>: models from experiments with different data and hyperparameters are averaged at each reward-modelling, SFT and DPO stage; checkpoint averaging is also applied during pre-training annealing, and the flagship model is used to improve the smaller models in post-training.</p><p>[12] Cohere, <a href="https://cohere.com/research/papers/command-a-technical-report.pdf">&#8220;Command A: An Enterprise-Ready Large Language Model&#8221;</a>: &#8220;Expert merging is a core feature of the Command A training pipeline&#8221;, alongside Polyak and seed averaging and interpolation for capability recovery.</p><p>[13] Three grades of the operation, in ascending difficulty. Averaging variants produced by one pipeline under different data or hyperparameters: Llama 3.1 and Command A above. Merging separately post-trained branches that share an ancestor: TinyR1&#8217;s three domain experts and SEA-LION&#8217;s DELLA stages, below. Cross-family merging between unrelated bases: requires tokenizer transplantation when vocabularies differ (note 21) and remains largely research territory (note 14). The 2025 exhibits in this piece are the second grade.</p><p>[14] <a href="https://arxiv.org/abs/2511.21437">&#8220;A Systematic Study of In-the-Wild Model Merging for Large Language Models&#8221;</a>, arXiv:2511.21437. Six methods, four base models, twelve checkpoints each, sixteen benchmarks.</p><p>[15] Sun et al., <a href="https://arxiv.org/abs/2503.04872">&#8220;TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation&#8221;</a>. Qiyuan Tech and Peking University. Branch phase: domain-specific SFT of DeepSeek-R1-Distill-Qwen-32B on math, coding, and science data. Merge phase: Arcee Fusion (Goddard et al., 2024). AIME 2024: 78.1 versus DeepSeek-R1&#8217;s 79.8. Merge cost: 4 H800 GPU-hours versus 740 for data-mixture retraining; the paper reports using 0.5% of the data-mixture merging compute and describes model merging as &#8220;a &#8216;free-lunch&#8217; approach&#8221;.</p><p>[16] AI Singapore, <a href="https://arxiv.org/abs/2504.05747">SEA-LION technical report</a>, arXiv:2504.05747. Post-training pipeline for Gemma-SEA-LION-v3-9B-IT: two instruction-tuning stages, alignment, and multiple merge stages using DELLA Linear.</p><p>[17] Yang et al., <a href="https://arxiv.org/abs/2408.07666">&#8220;Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities&#8221;</a>, ACM Computing Surveys, 2026.</p><p>[18] Goddard et al., <a href="https://aclanthology.org/2024.emnlp-industry.36/">&#8220;Arcee&#8217;s MergeKit: A Toolkit for Merging Large Language Models&#8221;</a>, EMNLP 2024 Industry Track, pp. 477&#8211;485. The author served as Chief Evangelist at Arcee AI until November 2025 and holds no current commercial relationship with the company.</p><p>[19] NVIDIA, <a href="https://arxiv.org/abs/2406.11704">&#8220;Nemotron-4 340B&#8221;</a>, released under the NVIDIA Open Model License permitting synthetic data generation for training other models.</p><p>[20] <a href="https://github.com/arcee-ai/mergekit">MergeKit repository</a>, hardware requirements per repository documentation (accessed 23 July 2026).</p><p>[21] <a href="https://arxiv.org/abs/2506.06607">&#8220;Breaking Down Model Vocabulary Barriers with Tokenizer Transplantation&#8221;</a>, arXiv:2506.06607. Orthogonal matching pursuit reconstruction of embedding matrices; ~96% retention across model families, training-free.</p><p>[22] <a href="https://docs.unsloth.ai">Unsloth documentation</a>, Qwen3.5-397B-A17B: ~807 GB full checkpoint; ~214 GB at Unsloth dynamic 4-bit (UD-Q4_K_XL), loadable on a 256 GB M3 Ultra; 3-bit fits 192 GB systems; 2-bit dynamic (UD-Q2_K_XL) recommended by Unsloth as a size/accuracy balance. Via MoE offloading, runs on a single 24 GB GPU plus 256 GB system RAM at 25+ tokens per second. Unsloth documentation, July 2026. Performance-tier comparison to closed frontier models is Unsloth&#8217;s claim; vendor-adjacent, label accordingly. File sizes are directional and vary by quant build.</p><p>[23] Community GGUF quantization tables (<a href="https://huggingface.co/mradermacher">mradermacher</a> and similar) publish the full precision ladder with quality annotations; the 1-bit IQ1 entries carry the label &#8220;for the desperate&#8221;.</p><p>[24] 24 GB: RTX 3090/4090, RX 7900 XTX. 32 GB: RTX 5090. Manufacturer specifications.</p><p>[25] AMD Ryzen AI Max+ 395 (&#8221;Strix Halo&#8221;): 128 GB unified LPDDR5X. Platform unveiled in Lisa Su&#8217;s CES keynote, 5 January 2026; AMD&#8217;s <a href="https://www.amd.com/en/newsroom/press-releases/2026-1-5-amd-expands-ai-leadership-across-client-graphics-.html">CES press release</a> describes the Ryzen AI Halo developer platform as &#8220;capable of running up to 200 billion parameter models locally&#8221;, with the developer machine at $3,999. Third-party 128 GB mini-PCs sell for roughly $2,000 street as of July 2026; the widely quoted ~$1,499 entry units carry 64 GB and cannot load the largest models. Street prices vary by configuration.</p><p>[26] Sparse mixture-of-experts models activate a fraction of parameters per token, so decode cost tracks active rather than total parameters; for a concrete instance of a very large MoE running on modest hardware via offloading, see the Qwen3.5 figures in note 10.</p><p>[27] <a href="https://docs.vllm.ai/en/stable/features/speculative_decoding/">vLLM speculative decoding documentation</a>: draft and target models must share a vocabulary by default. A heterogeneous-vocabulary path (token-level intersection) exists with ~60% reported acceptance for overlapping vocabularies.</p><p>[28] TokenTiming, <a href="https://arxiv.org/abs/2510.15545">arXiv:2510.15545</a>, on the shared-vocabulary constraint and the training-stage cost of aligned draft models.</p><p>[29] Li et al., <a href="https://arxiv.org/abs/2401.15077">&#8220;EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty&#8221;</a>. EAGLE-class draft heads train against the frozen target model&#8217;s features; the vLLM documentation in note 27 lists EAGLE and MTP among supported speculative methods.</p><p>[30] <a href="https://learn.microsoft.com/en-us/azure/foundry/how-to/fireworks/import-custom-models">Microsoft Foundry custom model import</a>, running on the Fireworks inference runtime within Foundry: full weight models, LoRA adapters (preview), draft models for speculative decoding (preview). Documentation, June 2026.</p><p>[32] Cursor, <a href="https://cursor.com/blog/composer">&#8220;Composer: Building a fast frontier model with RL&#8221;</a>, October 2025. Cursor&#8217;s own comparison places Composer at frontier coding results with roughly four times the generation speed of comparably intelligent models, while noting GPT-5 and Sonnet 4.5 outperform it on raw score.</p><p>[33] Simon Willison, <a href="https://simonwillison.net/2025/Oct/29/cursor-composer/">notes on Composer</a>: Cursor researchers declined to say whether Composer starts from an open-weights base such as Qwen or GLM; reporting has widely inferred an open checkpoint. Cursor is also among the US companies House committees have reportedly queried over Chinese-model use; see <a href="https://www.techtimes.com/articles/320171/20260711/washington-wants-chinese-ai-out-corporate-america-open-weights-block-ban.htm">TechTimes</a>.</p><p>[34] Zhang et al., <a href="https://arxiv.org/abs/2409.03215">&#8220;xLAM: A Family of Large Action Models to Empower AI Agent Systems&#8221;</a>, Salesforce AI Research: first position on the Berkeley Function-Calling Leaderboard at the time, outperforming GPT-4 and Claude 3 on tool use; <a href="https://github.com/SalesforceAIResearch/xLAM">repository</a>. Leaderboard positions have since shifted with newer versions of the benchmark.</p><p>[35] Baichuan, <a href="https://arxiv.org/abs/2509.02208">&#8220;Baichuan-M2: Scaling Medical Capability with Large Verifier System&#8221;</a>: 32B model trained with multi-stage GRPO against a patient-simulator verifier; above 32 on HealthBench Hard, previously exceeded only by GPT-5; surpasses o3, Grok 3, Gemini 2.5 Pro and GPT-4.1 on HealthBench; near-lossless quantisation deploys on a single RTX 4090.</p><p>[36] Sample of four current guides read in full, selected for variance in author business model: <a href="https://www.bentoml.com/blog/navigating-the-world-of-open-source-large-language-models">BentoML</a> (inference infrastructure; June 2026), <a href="https://www.thundercompute.com/blog/best-open-source-llms">Thunder Compute</a> (GPU rental; July 2026), <a href="https://www.localchat.app/blog/open-source-llm-models">LocalChat</a> (local desktop application; June 2026), and <a href="https://huggingface.co/blog/daya-shankar/open-source-llms">an independent author</a> on the Hugging Face community blog (November 2025). Coverage pattern as described; merging absent from all four. Sample and selection logic stated in the interest of checkability.</p><p>[37] OpenAI, <a href="https://openai.com/index/api-model-distillation/">&#8220;Model Distillation in the API&#8221;</a>, 1 October 2024: Stored Completions capture frontier-model outputs as training data for fine-tuning smaller models, inside the platform. On rival offerings: <a href="https://www.infoworld.com/article/3544913/openai-updates-api-with-model-distillation-prompt-caching-abilities.html">InfoWorld</a> notes Google, Anthropic and AWS provide distillation capabilities.</p><p>[38] Microsoft Learn, <a href="https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/stored-completions">Azure OpenAI stored completions and distillation</a>: &#8220;Stored completion distillation training files cannot be accessed directly and cannot be exported externally/downloaded.&#8221; The page documents the classic workflow, which Microsoft schedules for migration to the Responses API in October 2026; the export restriction is as stated as of July 2026.</p><p>[39] TensorZero, <a href="https://www.tensorzero.com/blog/distillation-programmatic-data-curation-smarter-llms-5-30x-cheaper-inference/">&#8220;Distillation with Programmatic Data Curation&#8221;</a>: fine-tuned small models at 13.7x (GPT-4o mini) to 24.1x (Gemini 2.0 Flash Lite) lower cost per successful task than GPT-4o across their evaluation environments. Vendor-published benchmark, B-tier.</p><p>[40] Amazon Bedrock Custom Model Import, <a href="https://aws.amazon.com/about-aws/whats-new/2024/10/amazon-bedrock-custom-model-import">generally available 21 October 2024</a>; <a href="https://aws.amazon.com/about-aws/whats-new/2025/06/amazon-bedrock-custom-model-import-qwen-models">Qwen architectures added 11 June 2025</a>; <a href="https://aws.amazon.com/about-aws/whats-new/2025/11/bedrock-model-import-openai-gpt-oss-models/">OpenAI gpt-oss added 19 November 2025</a>. GA 21 October 2024, Hugging Face safetensors format; Qwen support added 11 June 2025; OpenAI gpt-oss support added 19 November 2025. Google Vertex AI custom weights import. Microsoft Foundry custom model import (note 21).</p><p>[41] Representative commentary: <a href="https://www.fastcompany.com/91576757/trumps-proposed-ban-on-chinese-ai-models-could-strengthen-beijings-hand">Fast Company</a> on the unresolved form and reach of any restriction, and <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/trump-administration-reportedly-reviving-push-to-ban-chinese-ai-models-following-kimi-k3-launch-citing-cybersecurity-concerns-downloadable-open-weights-could-make-an-outright-u-s-ban-nearly-impossible-to-enforce-amid-growing-adoption">Tom&#8217;s Hardware</a> on enforceability of restrictions on downloadable weights. B-tier commentary on a live policy question.</p><p>[42] Booz Allen Hamilton, <a href="https://www.boozallen.com/content/dam/home/docs/cyber/booz-allen-chinese-llm-report-may-2026.pdf">&#8220;What&#8217;s in America&#8217;s Code?&#8221;</a> (May 2026): in 2,800+ trials across roughly 460,000 lines of generated code, three of four Chinese code-generation models produced significantly more vulnerable code when the prompt identified the user as working for the US government; Qwen3-Coder added roughly 130% more vulnerabilities under the government persona. Booz Allen states its evidence stops short of showing backdoors or deliberate insertion. The security concern is empirically grounded; this piece takes no position on the policy response.</p>]]></content:encoded></item><item><title><![CDATA[Qwen 3.8: Soon Is Not a Date]]></title><description><![CDATA[Alibaba answers Moonshot's dated gap with an undated one]]></description><link>https://www.airealist.ai/p/qwen-38-soon-is-not-a-date</link><guid isPermaLink="false">https://www.airealist.ai/p/qwen-38-soon-is-not-a-date</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Sun, 19 Jul 2026 15:33:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4gJZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4gJZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4gJZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!4gJZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!4gJZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4gJZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4gJZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1694143,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/207670362?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4gJZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!4gJZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!4gJZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4gJZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9bbdd5-1dcd-414a-89b9-1a22b42338e8_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On Sunday morning, three days after Kimi K3 launched and eight days before its weights are due, Alibaba&#8217;s Qwen team announced Qwen 3.8: 2.4 trillion parameters, multimodal, &#8220;second only to Fable 5&#8221;, and &#8220;going open-weight soon&#8221; [1]. A preview is live on Alibaba&#8217;s Token Plan and its Qoder coding platforms [1][2]. No benchmark table accompanied the claim. No per-token API price had been published for the preview at the time of writing. No date [2][3].</p><p style="text-align: justify;">When K3 launched, <a href="https://www.airealist.ai/p/kimi-k3-and-the-checkpoint-gap">the Checkpoint Gap</a> named the window between a frontier claim and the artefact that lets anyone verify it. Moonshot&#8217;s version of the gap was aggressive but disciplined: a published benchmark table with unusually transparent harness footnotes, a frontier price of $15 per million output tokens, an independent evaluation within hours, and a closing date of July 27. Every claim in that launch can be graded against a calendar.</p><p style="text-align: justify;">Qwen 3.8 opens the same gap without the discipline, and the difference deserves a name. A dated gap is a falsifiable claim: on July 27, the weights ship and replicate, or they don&#8217;t, and either outcome is information. An undated gap asserts nothing testable. It is an option &#8212; the right, but never the obligation, to ship weights at a moment of the vendor&#8217;s choosing, while collecting the open-weights narrative in the meantime. The option is also free to write: announcing an open-weight release commits no capital, reserves no serving capacity, and books no obligation, while the announcement itself trades at full value in the news cycle. A date can be missed, and a missed date is a verdict. &#8220;Soon&#8221; can only drift. Nothing that happens next week, or next quarter, can prove it false.</p><p style="text-align: justify;">The promise deserves to be read against the record, and the record splits cleanly in April 2026. Through the Qwen3 generation, Alibaba was one of the industry&#8217;s most reliable flagship open-weight publishers: flagship generations shipped with downloadable weights, most under Apache 2.0, and the mid-tier still does [4]. Then Qwen3.6-Max-Preview arrived on April 20 as the first closed flagship in Qwen&#8217;s history, available only via the API on Alibaba Cloud [4]. Qwen3.7-Max followed in May: no checkpoint, no license file, no timeline announced for an open variant [5]. Alibaba never promised weights for either flagship. It simply stopped shipping them.</p><p style="text-align: justify;">And both of those closed flagships nonetheless arrived with Alibaba&#8217;s full launch machinery: a detailed technical post, complete benchmark tables, day-one pricing [4][5]. Qwen 3.8, the largest model the company has ever disclosed and the first flagship since the pivot to revive the open-weight language, arrived as a post on X. The words came back. The date did not.</p><p style="text-align: justify;">None of this makes the promise empty. Alibaba&#8217;s open-source catalog is real, enormous, and central to its ecosystem strategy, and a 2.4-trillion-parameter checkpoint takes time to prepare for public release. Moonshot itself left 11 days between the announcement and the weights. A preview is allowed to be a preview. If a Qwen 3.8 repository with a license file appears on Hugging Face in the coming weeks, this note&#8217;s caution was cheap insurance. But that is the property that makes undated promises worth naming: they can only be honored, never broken. There is no day on which &#8220;soon&#8221; fails.</p><p style="text-align: justify;">Then there is the question of who is making it. Alibaba&#8217;s annual report records an investment of approximately $800 million for a 36% equity interest in Moonshot as of March 2024 [6]. Read the timing with that in mind. Moonshot, which Bloomberg reports reached $300 million in annual recurring revenue in June and plans an IPO in as little as six months, launched the most expensive Chinese model ever shipped, with a dated open-weight commitment [7]. Three days later, one of its major shareholders announced a slightly smaller model, claimed a slightly better ranking, attached the same open-weight language with no date, and released it into the window, even though K3&#8217;s openness cannot yet be verified. Whatever the intent, the effect is precise: the open-frontier crown is being contested during the one week when the current claimant cannot produce the crown.</p><p style="text-align: justify;">The apparent self-harm resolves once you weigh the assets. Alibaba&#8217;s Moonshot position is roughly $800 million of preferred stock, diluted through subsequent rounds and due to be repriced by the IPO, whichever model wins the week [6]. Its Qwen franchise is a different order of asset: it anchors Alibaba Cloud&#8217;s model business and, increasingly, the parent&#8217;s own AI narrative, days after a K3 launch that Bloomberg credits with shaking global technology stocks [2]. Note also what an announcement without a benchmark table can and cannot do. It cannot win a developer, because developers need a table and a price. It can freeze an enterprise purchase decision, hold a sales pipeline, and steady the parent&#8217;s story between trading sessions, because those audiences respond to positioning, not checkpoints. An empty announcement is not a weak launch. It is a different instrument, aimed at a different audience.</p><p style="text-align: justify;">The emptiness may even be the calibration. Published numbers strong enough to beat K3 would mark down the IPO asset Alibaba still holds. Numbers too weak to beat it would concede the crown outright. Words alone contest the crown and leave the comparison untouched: every branch of the decision leads to an empty announcement. And the reading carries its own test. If the full launch post follows within days, with a table and a price, the machinery is catching up with the announcement. The longer the gap between the words and the numbers, the more the announcement was the product.</p><p style="text-align: justify;">For the reader fielding this in a vendor review or an investment memo this week, the two launches belong in different rows of the table. K3&#8217;s numbers are vendor-published but disciplined by a date: they can be checked against the July 27 checkpoint, the independent evaluations already running, and a rate card that commits Moonshot to a price. Qwen 3.8&#8217;s single comparative sentence has, at the time of writing, no benchmark behind it, no price beneath it, and no calendar in front of it. Cite the first as a claim awaiting verification. Cite the second, if at all, as an announcement.</p><p style="text-align: justify;">The grading rule this produces is simple and portable. When a frontier launch opens a Checkpoint Gap, the first question is not the benchmark score but whether the gap has a date. Moonshot attached a date to its claim and a price to its confidence; its numbers are pending verification. Alibaba attached neither; its numbers are pending existence. On July 27, the first gap closes, one way or the other. The second has no edges at all. In the Checkpoint Gap, the only hard currency is a date, and Alibaba isn&#8217;t spending any.</p><div><hr></div><h3>Notes</h3><p>[1] Qwen (@Alibaba_Qwen), <a href="https://x.com/Alibaba_Qwen/status/2078759124914098291">launch announcement</a>, July 19, 2026. The quoted phrases (&#8221;second only to Fable 5&#8221;; &#8220;going open-weight soon&#8221;) are verbatim from the post, which describes Qwen3.8-Max-Preview as available on Alibaba&#8217;s Token Plan, Qoder, and QoderWork. Vendor-published.</p><p>[2] Bloomberg, <a href="https://finance.yahoo.com/technology/ai/articles/alibaba-qwen-unveils-preview-flagship-110209258.html">&#8220;Alibaba&#8217;s Qwen Unveils Preview of Flagship AI Model&#8221;</a>, July 19, 2026. Bloomberg notes the Sunday post &#8220;provided no additional technical specifications&#8221; beyond the parameter count, that Alibaba plans an open-weight release, and that K3&#8217;s launch days earlier helped upend perceptions of Chinese AI capability while shaking global technology stocks.</p><p>[3] The Decoder, <a href="https://the-decoder.com/alibabas-qwen-takes-on-kimi-k3-with-open-weight-qwen-3-8-says-model-is-second-only-to-fable-5/">&#8220;Alibaba&#8217;s Qwen takes on Kimi K3 with open-weight Qwen 3.8&#8221;</a>, July 19, 2026: no benchmark results are available at announcement.</p><p>[4] TokenMix, <a href="https://tokenmix.ai/blog/qwen3-6-max-preview-benchmark-review-2026">&#8220;Qwen3.6-Max-Preview Review&#8221;</a>, April 2026 &#8212; Qwen3.6-Max-Preview, released April 20, 2026, was the first flagship in Qwen&#8217;s history to ship closed-weights only, with API access through Alibaba Cloud; prior flagship generations through Qwen3 shipped with downloadable weights, most under Apache 2.0 (some earlier variants used the non-commercial Qwen research licence), and mid-tier Qwen3.6 models (27B, 35B-A3B) remain open under Apache 2.0. Corroborated by <a href="https://insiderllm.com/guides/qwen-open-weights-vs-closed-frontier-2026/">InsiderLLM</a>, which verified the absence of flagship checkpoints on the official Qwen Hugging Face organisation by direct repository probes as of June 15, 2026. B-tier sourcing; the underlying negative claim (no published flagship weights) is independently checkable against the Hugging Face organisation at any time.</p><p>[5] Yotta Labs, <a href="https://www.yottalabs.ai/post/qwen-3-7-max-release-date-features-open-source-status-and-how-to-access-2026">&#8220;Qwen 3.7-Max: Pricing, Features, and How to Access&#8221;</a>, May 2026: Qwen3.7-Max is not open-weight, cannot be downloaded, and is API-only; no timeline announced for an open variant. On pricing, sources conflict on the exact figure (<a href="https://artificialanalysis.ai/models/qwen3-7-max">Artificial Analysis</a> reports $2.50/$7.50 per million input/output tokens; <a href="https://openrouter.ai/qwen/qwen3.7-max">OpenRouter</a> lists $1.475/$4.425), but all place Qwen3.7-Max well below both Kimi K3&#8217;s $15 and the Anthropic frontier tier &#8212; Alibaba has so far never backed a flagship claim with a frontier price.</p><p>[6] Alibaba Group Holding Ltd, <a href="https://www.sec.gov/Archives/edgar/data/1577552/000095017024063767/Financial_Report.xlsx">Form 20-F, fiscal year ended March 31, 2024</a>: &#8220;the Company invested a total of approximately US$0.8 billion (approximately RMB 5.9 billion) for an approximately 36% equity interest&#8221; in Moonshot, held as preferred stock and accounted for under the measurement alternative. Subsequent Moonshot funding rounds (Series B extension, Series C, and a further round in early 2026) have likely diluted this percentage; Alibaba has not disclosed an updated figure.</p><p>[7] Bloomberg, as cited in [2] and [3]: Moonshot reached $300 million in annual recurring revenue in June 2026 and plans to go public in as little as six months. The &#8220;most expensive Chinese model ever shipped&#8221; characterisation of Kimi K3&#8217;s $15 per million output tokens is per <a href="https://simonwillison.net/2026/Jul/16/kimi-k3/">Simon Willison&#8217;s launch analysis</a>, July 16, 2026.</p>]]></content:encoded></item><item><title><![CDATA[Kimi K3 and the Checkpoint Gap]]></title><description><![CDATA[How to read a Chinese frontier launch in the eleven days before the weights ship]]></description><link>https://www.airealist.ai/p/kimi-k3-and-the-checkpoint-gap</link><guid isPermaLink="false">https://www.airealist.ai/p/kimi-k3-and-the-checkpoint-gap</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Fri, 17 Jul 2026 06:44:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!91LW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!91LW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!91LW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!91LW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!91LW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!91LW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!91LW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1754825,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/207388116?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!91LW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!91LW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!91LW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!91LW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc332182b-0aaf-495b-b6fb-b5bf8b65510a_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">On July 16, Moonshot AI announced Kimi K3, a 2.8-trillion-parameter model it bills as the first open model in the 3-trillion-parameter class. The launch post is unusually candid for the genre. It states plainly that K3 trails the two strongest proprietary models, Claude Fable 5 and GPT-5.6 Sol, and its limitations section concedes &#8220;a noticeable gap in user experience&#8221; against both [1]. The press coverage was less careful: within hours, headlines had K3 pushing Chinese AI &#8220;into Fable-level territory&#8221; [2].</p><p style="text-align: justify;">Two numbers in the launch matter more than any benchmark score. The first is a price: <strong>$15 per million output tokens, the most expensive model a Chinese lab has ever shipped</strong> and nearly four times Moonshot&#8217;s own K2.6 [3]. The second is a date: the full weights are promised by July 27, 2026 [1]. Until then, every measurable claim about K3 passes through an API endpoint operated by Moonshot. The score that matters this week is not an Elo. It is a date.</p><p style="text-align: justify;">Call the window between those two events the Checkpoint Gap: the period when a model is priced, benchmarked, and covered as open, while the artifact that would let anyone verify it sits on the vendor&#8217;s servers. The industry has had a recent, expensive lesson in what can live inside that window. In April 2025, Meta submitted a chat-tuned variant of Llama 4 Maverick to the LMArena leaderboard, where it briefly ranked second. When the public weights were tested, the released model landed around 32nd, and LMArena rewrote its submission policies with an unusually direct rebuke [4]. The benchmarked model and the shipped checkpoint were not the same thing. Nobody outside Meta could have known before the weights landed.</p><p style="text-align: justify;">Nothing suggests Moonshot is running that play. The company&#8217;s record argues the opposite: Kimi K2 shipped open under a modified MIT license in July 2025, and K2.6 followed in April 2026 [5]. Moonshot&#8217;s benchmark footnotes are also more transparent than most Western launches, and they repay reading. K3 is evaluated with Moonshot&#8217;s own KimiCode evaluation harness on most coding benchmarks, while rivals run under Claude Code, Codex, or their best score across harnesses. Fable 5&#8217;s results may reflect a fallback to Opus 4.8 when it refuses a task. K3 itself currently runs at a single reasoning effort, its highest [6]. Where the harness effect can be isolated, it looks small: on DeepSWE, K3 scores 67.5 under KimiCode and 67.3 under the official leaderboard&#8217;s harness [6]. These are disclosed choices, not hidden ones. But disclosure is not verification. Until July 27, good faith and the Llama 4 play are indistinguishable from the outside, and the correct reading posture is the same for both.</p><p style="text-align: justify;">What independent evidence exists places K3 near the frontier, not at it &#8212; and even that evidence was gathered through Moonshot&#8217;s API. Artificial Analysis scores K3 at 57 on its Intelligence Index, behind Claude Fable 5 (around 60) and GPT-5.6 Sol (around 59) and comparable to Opus 4.8, and its private knowledge-work evaluation ranks K3 second only to Fable 5 [7]. Third place at launch from a lab that will hand you the weights is a remarkable result. It is not the frontier moving to Beijing. And note the tense the evaluator itself uses: once available, AA writes, K3 &#8220;would clearly lead&#8221; the open-weight field [7]. The conditional is doing the work. Even the independent scorekeeper is writing from inside the gap.</p><p style="text-align: justify;">The price is where this launch stops resembling every previous Chinese release. DeepSeek&#8217;s pitch never depended on benchmark verification: at $0.87 per million output tokens, 90% of the frontier was a bargain even if the numbers flattered [8]. K3 at $15 has exited the Chinese price war entirely. On the rate card, it is 17 times DeepSeek V4, level with Anthropic&#8217;s Sonnet tier, and 30% of Fable&#8217;s $50 [3][8]. On cost per completed task, a separate metric, AA&#8217;s estimate lands K3 at $0.94 &#8212; beside GPT-5.6 Sol at $1.04, half of Opus 4.8 at $1.80, and three times its open-weight peer GLM-5.2 [7]. </p><p style="text-align: justify;">Both metrics say the same thing: this model is priced with the frontier, not against it. Nobody pays that premium for third place unless the table holds. The rate card is Moonshot&#8217;s self-assessment denominated in dollars, and it converts the benchmark table from marketing garnish into the product justification. The Checkpoint Gap matters more for K3 than for any Chinese launch before it, because this is the first one to ask frontier prices for claims nobody can yet check.</p><p style="text-align: justify;">The gap closes on schedule, making the launch falsifiable in a way most are not. Three tests, all dated. </p><ol><li><p style="text-align: justify;">Do the weights ship by July 27, and does the released checkpoint match what the API has been serving?</p></li><li><p style="text-align: justify;">Do independent evaluations under a neutral harness reproduce the launch table? </p></li><li><p style="text-align: justify;">Where does third-party pricing settle? On that last test, temper expectations of a DeepSeek-style collapse toward the hosting floor. Moonshot recommends serving K3 on supernodes of 64 or more accelerators, so price competition will come from large inference providers, not from tiny GPU resellers [1]. If the $15 holds once alternatives exist, the market has accepted the frontier positioning. If it collapses, the price will launch in the theatre.</p></li></ol><p>Did the frontier move to China? Not on the evidence available today, and Moonshot&#8217;s own launch post doesn&#8217;t claim it did. What moved is the posture. DeepSeek priced like it had something to prove. Moonshot is pricing like it has already proved it, eleven days before anyone can check.</p><div><hr></div><h3>Notes</h3><p>[1] Moonshot AI, <a href="https://www.kimi.com/blog/kimi-k3">&#8220;Kimi K3: Open Frontier Intelligence&#8221;</a>, July 16, 2026. The weights date, API pricing ($0.30/MTok cache-hit input, $3.00/MTok cache-miss input, $15.00/MTok output), the concession that K3 trails Claude Fable 5 and GPT-5.6 Sol, the user-experience limitation, and the recommendation to deploy on supernodes of 64 or more accelerators all appear in the launch post. Vendor-published.</p><p>[2] Fortune, <a href="https://fortune.com/2026/07/16/moonshots-kimi-k3-pushes-chinese-ai-into-fable-level-territory/">&#8220;Moonshot&#8217;s Kimi K3 pushes Chinese AI into Fable-level territory&#8221;</a>, July 16, 2026.</p><p>[3] Simon Willison, <a href="https://simonwillison.net/2026/Jul/16/kimi-k3/">&#8220;Kimi K3, and what we can still learn from the pelican benchmark&#8221;</a>, July 16, 2026 &#8212; identifies K3 as the most expensive model released by a Chinese lab to date, at the level of Anthropic&#8217;s Claude Sonnet series. Kimi K2.6 pricing ($0.95/MTok input, $4/MTok output) per <a href="https://platform.kimi.ai/docs/pricing/chat-k26">Moonshot&#8217;s platform documentation</a>.</p><p>[4] The Register, <a href="https://www.theregister.com/2025/04/08/meta_llama4_cheating/">&#8220;Meta accused of Llama 4 bait-n-switch to juice LMArena rank&#8221;</a>, April 8, 2025. LMArena stated Meta&#8217;s interpretation of its submission rules &#8220;did not match what we expect from model providers&#8221; in its April 2025 policy update; the released Maverick&#8217;s subsequent placement around 32nd is documented in <a href="https://the-decoder.com/metas-llama-4-models-show-promise-on-standard-tests-but-struggle-with-long-context-tasks/">The Decoder&#8217;s follow-up coverage</a> of the leaderboard.</p><p>[5] Moonshot AI, <a href="https://www.kimi.com/blog/kimi-k2">&#8220;Kimi K2: Open Agentic Intelligence&#8221;</a>, July 2025, and <a href="https://www.kimi.com/blog/kimi-k2-6">&#8220;Kimi K2.6&#8221;</a>, April 2026.</p><p>[6] Moonshot K3 launch post [1], benchmark footnotes 1&#8211;8. Harness assignments per benchmark, the Fable 5 fallback condition, the single (maximum) reasoning effort at launch, and the DeepSWE dual-harness scores (67.5 under KimiCode; 67.3 under the official leaderboard&#8217;s mini-SWE-agent harness) are all disclosed there.</p><p>[7] Artificial Analysis, <a href="https://x.com/ArtificialAnlys/status/2077832874183860404">Kimi K3 launch evaluation</a> and <a href="https://artificialanalysis.ai/models/kimi-k3">model page</a>, July 16, 2026. Intelligence Index: K3 at 57, behind Claude Fable 5 and GPT-5.6 Sol and comparable to Opus 4.8 and GPT-5.5. AA-Briefcase (private long-horizon knowledge work): Elo 1547, second behind Claude Fable 5. Cost per task is AA&#8217;s evaluation-derived estimate, not a market price: K3 $0.94, GPT-5.6 Sol $1.04, Claude Opus 4.8 $1.80, GLM-5.2 $0.32. One counterweight AA also reports: K3 generated 130M output tokens across the Index against a 63M median for comparable reasoning models, so on verbose workloads the rate card understates effective cost &#8212; a consequence of the single maximum reasoning effort available at launch. The open-weight comparison (&#8221;would clearly lead&#8221; GLM-5.2 and DeepSeek V4 Pro) is AA&#8217;s own phrasing, stated in the conditional pending the weight release.</p><p>[8] Fortune [2] for comparative output-token pricing: DeepSeek V4 at $0.87, z.ai&#8217;s GLM-5.2 at $4.40, and Claude Fable at $50 per million output tokens.</p>]]></content:encoded></item><item><title><![CDATA[The Backstop Has a Name Now - part 2]]></title><description><![CDATA[How Nvidia finances a cloud tells you what it thinks the cloud is worth. CoreWeave and Nebius got equity. Sharon AI got a revenue-share.]]></description><link>https://www.airealist.ai/p/the-backstop-has-a-name-now-part-11d</link><guid isPermaLink="false">https://www.airealist.ai/p/the-backstop-has-a-name-now-part-11d</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Mon, 06 Jul 2026 09:05:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4W_6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4W_6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4W_6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!4W_6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!4W_6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4W_6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4W_6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1924676,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/205041043?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4W_6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!4W_6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!4W_6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4W_6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcab7276f-929f-489b-a6e3-7abc7f11aee0_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">Ask why Nvidia took equity in the neoclouds it works with &#8212; CoreWeave, Nebius, even Firmus, its own revenue-share launch partner &#8212; and took none in Sharon AI, and the answer is in the instrument itself.</p><p style="text-align: justify;">Start with what Nvidia does for the clouds that made it. CoreWeave has been public since March 2025; it carries more than $20 billion of debt and once looked reckless to venture investors, but it borrows against investment-grade, asset-backed paper &#8212; an $8.5 billion facility rated A3, secured by the chips and the customer contracts, at roughly SOFR plus 2.25 percent.[1] Nebius, the old Yandex reconstituted on Nasdaq, posted positive adjusted EBITDA in the first quarter of 2026, raised more than $6 billion this year, and ended the quarter with $9.3 billion in cash.[2] Both are anchored by hyperscalers whose contracts pay in advance: CoreWeave by Microsoft, Nebius by Meta and Microsoft, on deals worth tens of billions.[3] And into both, Nvidia put equity &#8212; $2 billion into CoreWeave in January, $2 billion into Nebius in March, the same instrument it used that season in Lumentum and Coherent.[4]</p><p style="text-align: justify;">Equity is the tell. It is a bet on enterprise value: a junior claim that pays only if the company becomes worth something, which it has for these two. Nvidia takes equity where there is value to own. It did not offer Sharon AI equity, and it would not, because the thing a revenue-share does that equity cannot is sit senior to the shareholders, a claim on revenue paid off the top, ahead of a stake that may end up worthless. How Nvidia chooses to finance a cloud is a readout of what it thinks the cloud is worth. See value, and it buys in. See doubt, and it keeps its distance: a claim on the revenue, a loan to produce it, and no share of the company.</p><p style="text-align: justify;">The counterparty side complicates the neat version. CoreWeave and Nebius fund their buildouts with investment-grade debt and hyperscaler prepayments, so they never needed Nvidia&#8217;s financing. Firmus and Sharon AI did take it, and Firmus is worth $5.5 billion, so a revenue share is not, by itself, a mark of distress. It is how a buildout gets financed quickly, as Nvidia&#8217;s own chief financial officer frames it: a way to serve companies with demand but who cannot secure financing quickly enough.[5] The instrument opens a new recurring revenue line for Nvidia. What separates the two who took it is whether Nvidia also wanted to own them.</p><p style="text-align: justify;">The debt market already sorts AI borrowers by distance from cash: Amazon borrows unsecured, Oracle against backlog, SoftBank could not borrow against a private mark at all.[6] The revenue-share program is what sits below SoftBank&#8217;s rung, where the debt market says no and Nvidia says yes, through an instrument no bank would offer. It is a familiar move under a new name. An earlier piece called it the Overbuild Put: a company financing a buildout it might not fill names its own backstop, and the backstop is the giveaway.[7] Meta named its exit, a cloud business it might have to start if it overbuilt. Sharon AI&#8217;s exit is named for it, by Nvidia; the credit support is the backstop that lets a buildout proceed whose demand is unproven. The difference is who writes the text. Meta wrote one on its own capacity. Nvidia writes one on Sharon AI&#8217;s, the supplier backstopping the demand of a customer it also sells to.</p><p style="text-align: justify;">The February COMECON piece argued Nvidia holds the independent cloud tier captive; Part 1 named the productized version.[8] The line is not equity versus revenue-share, since Firmus took both; it is ownership versus none. Nvidia holds a stake in CoreWeave, Nebius, and Firmus. In Sharon AI, it holds only a claim.</p><h2>What &#8220;fragile enough to sign&#8221; looks like</h2><p style="text-align: justify;">Sharon AI is the exhibit, the one launch partner Nvidia financed but would not own. Firmus, the other, raised $505 million at a $5.5 billion valuation with Nvidia among its backers;[4] whatever else it is, it is not the bottom of the ladder. Sharon AI&#8217;s own filings tell most of the story before any short seller does.</p><p style="text-align: justify;">At the end of March, the company held $164 million in cash and no revenue; its filings do not expect revenue to begin until September, and operating cash flow is negative, with a capital-expenditure requirement of about $720 million to support its lead contract.[9] Against the cash are two customer contracts totaling $1.26 billion and a second worth roughly $950 million.[10] Its reported quarterly loss of $20 million is mostly an accounting artifact: the company carries its convertible notes at fair value, so remeasuring them drove a $70 million loss through the income statement, offset by a $66 million gain on the sale of half of a Texas data-center joint venture.[11] A company whose reported profit swings with the value of its own debt and whose cash comes from an IPO and asset sales rather than operations is being valued on a chain of announcements, not cash flow. There is none yet.</p><p style="text-align: justify;">The build is financed by a stack of commitments, each conditioned on the next. A $200 million facility from an investor called Digital Alpha and a $500 million facility from USD.AI, a roughly one-year-old decentralized-finance protocol, were both announced as &#8220;up to&#8221; and, per a short report, remained unexecuted months later.[12] In May, the company closed $350 million of convertible notes led by Oaktree. Per the same report, those notes were funded only if Sharon AI first signed a binding contract for at least 4,068 additional GPUs, its lender declining to treat the announced $1.26 billion in demand as sufficient collateral.[13] In June, it raised an additional $1.6 billion.[14] And on June 12 came the keystone, the six-year Nvidia deal, whose filing language says it is &#8220;structured so that Sharon AI can commit to large-scale NVIDIA infrastructure&#8221; through the revenue-share and credit-support.[15] Nvidia&#8217;s credit line is the piece that lets the rest of the tower stand.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9jrC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9jrC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png 424w, https://substackcdn.com/image/fetch/$s_!9jrC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png 848w, https://substackcdn.com/image/fetch/$s_!9jrC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png 1272w, https://substackcdn.com/image/fetch/$s_!9jrC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9jrC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png" width="1639" height="2124" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2124,&quot;width&quot;:1639,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:403924,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/205041043?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57cac719-a074-4811-9ba9-a948e6d4e23c_1640x2400.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9jrC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png 424w, https://substackcdn.com/image/fetch/$s_!9jrC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png 848w, https://substackcdn.com/image/fetch/$s_!9jrC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png 1272w, https://substackcdn.com/image/fetch/$s_!9jrC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2edaa882-2e11-4e34-9fd5-f514df57e083_1639x2124.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">Then the customer, where the short case overreaches, and the real point is sharper. Sharon AI&#8217;s forward revenue rests on the $1.26 billion contract with ESDS Software Solution, and ESDS is not a shell. It is an established Indian cloud provider with roughly &#8377;361 crore (about $43 million) in revenue last fiscal year, up 27 percent, profitable, lightly indebted, and preparing an IPO of its own.[16] The problem is scale, not solvency. The contract calls for average annual payments of roughly $250 million and $140 million in letters of credit; the annual figure alone is six times ESDS&#8217;s entire revenue.[17] A real $43 million company can be a real counterparty to a contract its own size; whether it can perform one thirty times larger is the open question, and it is the same question Sharon AI&#8217;s own lender asked. The chief executive brings his own history. Manning&#8217;s prior public company, Mawson Infrastructure Group, alleges in court filings that he directed roughly A$11.5 million to a shipping firm he controlled without disclosing his interest to the board, allegations he contests and that remain unadjudicated. That same firm, Flynt, is now a disclosed related-party vendor of Sharon AI, according to the company&#8217;s own prospectus.[18] None of this has stopped the stock, which has risen sharply; the market has not endorsed the short thesis.[19] But it is the profile of an operator that could not fund this build on ordinary terms, which is why the revenue-share was there to be signed.</p><p style="text-align: justify;">One detail seals the instrument argument. Nvidia holds no equity in Sharon AI; its annual report called Nvidia a &#8220;strategic shareholder&#8221; and the company filed a correction stating Nvidia owns none.[20] It invested $2 billion in equity in each of CoreWeave and Nebius, and took a stake in Firmus, its own revenue-share partner. Only in Sharon AI did it take a revenue-share, a credit claim, and no ownership at all, because what it is underwriting here is not an asset it wants to hold. It is a demand that needs to be kept alive.</p><h2>What comes next</h2><p style="text-align: justify;">The rule generalizes, and it is the thing to take from Sharon AI. When a supplier finances its own customer, the instrument it chooses is a private credit rating, better informed than the market&#8217;s, because the supplier sits inside the relationship. Equity is the vote of confidence. The revenue share is how the buildout gets paid for, and Firmus, worth $5.5 billion, took one too. What sets Sharon AI apart is the equity Nvidia declined to take. Call it the Instrument Test: read what a vendor takes, not what it announces. It travels beyond Nvidia: whenever a vendor lends to the customers who buy its product, the instrument encodes what the vendor privately believes, and usually before the tape does.</p><p style="text-align: justify;">Which makes the program itself an indicator. Nvidia split its partners by what it was willing to own: it took a stake in CoreWeave, Nebius, and Firmus, and in Sharon AI it took only a claim. As the program spreads, watch which companies get a revenue share with no stake besides it. That pairing, not the revenue-share alone, is the readout of how far down the counterparty-quality curve Nvidia is reaching, and it moves before the market does, because it is Nvidia&#8217;s own hand showing. When the next partner arrives with a credit line and no equity cheque, Nvidia is saying, in the one language it cannot fake, that it is a company it will finance but will not own. And note the second edge. When a chip vendor&#8217;s credit support helps fund the purchase of its own chips and then collects a share of the revenue those chips generate, the revenue is partly its own money coming home &#8212; the same round-trip that ran at the hyperscaler layer, now reaching the bottom of the neocloud tier.[21]</p><p style="text-align: justify;">A fair objection remains: a revenue-share cuts both ways. If an operator&#8217;s utilization falls, so does Nvidia&#8217;s cut, so this is exposure to Sharon AI&#8217;s success, not merely a claim on it. But exposure to the success of the operators least likely to deliver it is either deep conviction or the position of a supplier that has run out of stronger customers to sell to.</p><p style="text-align: justify;">CoreWeave was fragile once, too, and became a company worth billions; one of these operators may do the same, and Oaktree and Goldman are betting on it. But read the instrument. In CoreWeave, Nebius, and Firmus, Nvidia is a shareholder, betting the company wins. In Sharon AI it is a creditor, arranging to be paid whether it wins or not. Which seat a supplier takes says more than any forecast it offers. And at the bottom of the ladder, Nvidia took the creditor&#8217;s.</p><div><hr></div><h3>Notes</h3><p>[1] CoreWeave (Nasdaq: CRWV) listed March 28, 2025; total debt now exceeds $21 billion. Its March 2026 $8.5 billion facility is rated A3/A(low), secured by substantially all assets of the borrower group, at SOFR + 2.25% floating or ~5.9% fixed, maturing March 2032, and is the first investment-grade-rated GPU-backed financing: <a href="https://sacra.com/c/coreweave/">Sacra</a>; <a href="https://qz.com/gpu-collateralized-debt-ai-neocloud-coreweave-financing-risks-050526">Quartz</a>. Its credit agreement requires contracts with &#8220;large and creditworthy&#8221; customers covering future debt repayments.</p><p>[2] Nebius Group (Nasdaq: NBIS), formerly Yandex N.V., resumed Nasdaq trading October 2024. Q1 2026: revenue $399M (+684% YoY); positive adjusted EBITDA (~45% AI-cloud margin) and net income of $621.2M (inclusive of non-operating items); more than $6 billion raised in 2026 ($4.3B convertible notes plus Nvidia&#8217;s equity), ending the quarter with $9.3B in cash: <a href="https://www.tikr.com/blog/nebius-grew-revenue-684-in-q1-sold-out-its-entire-capacity-and-raised-its-capex-target-to-25-billion">TIKR</a>; <a href="https://simplywall.st/stocks/us/software/nasdaq-nbis/nebius-group">Simply Wall St</a>.</p><p>[3] CoreWeave&#8217;s anchor customer is Microsoft; Nebius holds a five-year ~$27B Meta agreement and a ~$17&#8211;19.4B Microsoft agreement, with hyperscaler prepayments helping fund capex: <a href="https://www.forbes.com/sites/rashishrivastava/2025/09/22/coreweaves-29-billion-bet-that-its-debt-fueled-ai-boom-wont-go-bust/">Forbes</a>; <a href="https://www.morningstar.com/stocks/this-ai-cloud-stock-is-up-over-300-year-can-it-rise-further">Morningstar</a>.</p><p>[4] Nvidia made a $2 billion private placement in CoreWeave in January 2026 (22,935,780 Class A shares at $87.20) as part of an expanded collaboration targeting more than 5 GW by 2030, and agreed to purchase CoreWeave&#8217;s unsold capacity through 2032: <a href="https://sacra.com/c/coreweave/">Sacra</a>; <a href="https://www.forbes.com/sites/rashishrivastava/2025/09/22/coreweaves-29-billion-bet-that-its-debt-fueled-ai-boom-wont-go-bust/">Forbes</a>. Its $2 billion equity investment in Nebius (March 11, 2026) mirrored $2 billion investments in Lumentum and Coherent the same month: <a href="https://mlq.ai/news/nvidia-invests-2-billion-in-nebius-to-advance-ai-cloud-infrastructure/">MLQ News</a>. Nvidia also holds equity in Firmus, having joined a 2025 round and participated in Firmus&#8217;s April 2026 US$505 million round at a US$5.5 billion valuation led by Coatue; the April participation was reported subject to closing conditions. Firmus separately arranged a US$10 billion debt facility. Verify against Firmus&#8217;s April 2026 funding release before publication.</p><p>[5] Nvidia CFO Colette Kress, framing the program as serving companies with demand that cannot secure financing quickly enough: <a href="https://blogs.nvidia.com/blog/nvidia-unlocks-ai-compute-at-scale-capital-partners-to-power-ai-infrastructure-buildout/">NVIDIA blog, July 1, 2026</a>.</p><p>[6] &#8220;Cash Flow Lends. Valuation Doesn&#8217;t.,&#8221; The AI Realist, <a href="https://www.airealist.ai/p/cash-flow-lends-valuation-doesnt">June 12, 2026</a>.</p><p>[7] &#8220;The Overbuild Put,&#8221; The AI Realist, <a href="https://www.airealist.ai/p/the-overbuild-put">June 1, 2026</a> &#8212; reading a fallback-monetisation or backstop remark as a credit signal on a debt-financed buildout whose demand is unproven.</p><p>[8] &#8220;Jensen&#8217;s COMECON: How Nvidia Built an Empire of Captive Clouds,&#8221; The AI Realist, <a href="https://www.airealist.ai/p/jensens-comecon-how-nvidia-built">Feb. 14, 2026</a>; &#8220;The Backstop Has a Name Now (Part 1),&#8221; The AI Realist, July 2026 <em>(confirm published slug before linking)</em>.</p><p>[9] SharonAI Holdings Inc., Form 10-Q for the quarter ended March 31, 2026 (cash $164,288,288; negative operating cash flow; revenue commencement not expected until approximately September 2026), attached to the Company&#8217;s Form 424B3, <a href="https://www.sec.gov/Archives/edgar/data/0002068385/000149315226030359/form424b3.htm">SEC EDGAR (CIK 0002068385)</a>. Approximately $720 million of capital expenditure is tied to the lead customer arrangement per the same filing.</p><p>[10] ESDS master services agreement of $1,260,000,000 and a second customer contract of approximately $950,000,000, per the Form 424B3 referenced in [9].</p><p>[11] Fair-value option elected on the convertible notes (Level 3); Q1 2026 included a $70.2 million loss on remeasurement and a $65,919,712 gain on the sale of a 50% interest in the Texas Critical Data Centers joint venture; convertible-note fair value of $199,358,226 as of March 31, 2026 (Form 424B3, [9]).</p><p>[12] Announcements of a Digital Alpha facility of up to $200 million (Jan. 19, 2026, subject to definitive documentation) and a USD.AI facility of up to $500 million (Jan. 22, 2026). Characterisations of USD.AI&#8217;s on-chain capacity and the unexecuted status of the Digital Alpha facility are from <a href="https://www.bleeckerstreetresearch.com/research/shaz">Bleecker Street Research, &#8220;SharonAI (SHAZ),&#8221; April 30, 2026</a> (a short seller with a disclosed position).</p><p>[13] <a href="https://www.sec.gov/Archives/edgar/data/0002068385/000149315226024865/ex99-1.htm">SharonAI Form 8-K, Exhibit 99.1 (closing of $350 million of convertible senior notes due 2031, led by Oaktree), May 2026, SEC EDGAR</a>. The 4,068-GPU closing condition is as characterised in the Bleecker Street report ([12]); confirm against the note purchase agreement before publication.</p><p>[14] SharonAI&#8217;s oversubscribed $1.6 billion private placement (approximately $900 million equity and $700 million of 4.75% convertible notes due 2032), June 2026, with Goldman Sachs as lead placement agent: <a href="https://www.benzinga.com/markets/equities/26/06/60177706/sharon-ai-jumps-nearly-12-after-hours-what-is-going-on-with-shaz-stock">Benzinga</a>.</p><p>[15] Six-year Nvidia collaboration (72 MW, up to 40,000 Grace Blackwell GB300), per the Company&#8217;s June 12, 2026 announcement reproduced in the Form 424B3 ([9]).</p><p>[16] ESDS Software Solution Limited, FY2025 (year ended March 31, 2025): total revenue &#8377;361 crore (~$43 million), up 27% year over year; profit before tax up nearly fourfold; debt-to-equity of 0.15 and current ratio of 2.32; DRHP filed March 30, 2025 for a &#8377;600 crore IPO on the BSE and NSE: <a href="https://unlistedzone.com/esds-software-solution-limited-a-transformative-year-and-a-promising-future">unlistedzone</a>; <a href="https://wwipl.com/unlisted-shares/esds-software-solution-limited/financial">ESDS financials via WWIPL</a>. Verify against ESDS&#8217;s audited DRHP financials before publication.</p><p>[17] The $250 million average annual payment and the $140 million letter-of-credit obligation are per the Bleecker Street report ([12]); confirm against the master services agreement before publication.</p><p>[18] Mawson Infrastructure Group Inc. (Nasdaq: MIGI) alleges, in its January 10, 2025 court filing, that James Manning, its former chief executive, caused the company to pay over A$11.4 million to Flynt International Cargo Solutions (a Vertua subsidiary) for services it &#8220;did not need,&#8221; without disclosing his interest or seeking board approval: <a href="https://www.sec.gov/Archives/edgar/data/1218683/000117184325000166/exh_991.htm">Mawson Form 8-K, Exhibit 99.1, SEC EDGAR</a>. The allegations are unadjudicated and contested. SharonAI&#8217;s own prospectus discloses Flynt ICS as a related-party vendor: &#8220;Flynt is a subsidiary of Vertua Limited and affiliated to the Group through common ownership by James Manning,&#8221; and the Group paid Flynt $167,638 in services expenses for the year ended December 31, 2024.</p><p>[19] Public market data as of early July 2026; SHAZ has risen sharply since the April 30 short report, and the market has not validated the short thesis.</p><p>[20] SharonAI corrected its FY2025 Form 10-K, which had described NVIDIA as a &#8220;strategic shareholder,&#8221; to state that NVIDIA holds no equity securities of the Company: <a href="https://www.stocktitan.net/sec-filings/SHAZ/8-k-sharon-ai-holdings-inc-reports-material-event-05a155b67636.html">SharonAI Form 8-K correction (SEC EDGAR)</a>.</p><p>[21] &#8220;The Round Trip,&#8221; The AI Realist, <a href="https://www.airealist.ai/p/the-round-trip">May 4, 2026</a>.</p>]]></content:encoded></item><item><title><![CDATA[The Backstop Has a Name Now - part 1]]></title><description><![CDATA[Nvidia is handing back tens of billions to shareholders. Now it&#8217;s offering to finance the customers who can&#8217;t afford its chips.]]></description><link>https://www.airealist.ai/p/the-backstop-has-a-name-now-part</link><guid isPermaLink="false">https://www.airealist.ai/p/the-backstop-has-a-name-now-part</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Sat, 04 Jul 2026 06:09:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JgWl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JgWl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JgWl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!JgWl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!JgWl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!JgWl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JgWl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1518863,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/205016499?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JgWl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!JgWl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!JgWl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!JgWl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8adde3b5-b47b-4eca-a548-ebd803555260_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">In its most recent quarter, Nvidia returned a record $20 billion to shareholders. In May, its board authorized another $80 billion in buybacks; in June, it raised $25 billion in the bond market &#8212; the balance sheet of a company with no financing problem of its own.[1] On July 1, it announced a program to help companies that cannot afford its chips buy them anyway, in exchange for a share of the profits.[2]</p><p style="text-align: justify;">The two sit oddly together. A company that hands out that much to shareholders does not usually need to lend to its own customers.</p><p style="text-align: justify;">In February, I called this the <a href="https://www.airealist.ai/p/jensens-comecon-how-nvidia-built">COMECON model</a>: Nvidia keeps the independent GPU clouds, the neoclouds, captive through four instruments: GPU allocation, equity stakes, credit enhancement, and demand backstops.[3] On July 1, it took two of them, the credit enhancement and the backstop, and gave them a product name.</p><p style="text-align: justify;">The program is a &#8220;revenue-sharing and credit-support model.&#8221; A cloud puts Nvidia GPUs on the floor without carrying the full capital cost, draws token credits against future capacity today, and hands Nvidia its standard hardware margin plus a recurring, usage-linked share of the cloud revenue that capacity generates.[2] How large that share is, Nvidia has not said. The first named partners are Sharon AI, deploying up to 40,000 Grace Blackwell GB300s in Australia, and Firmus, building toward 360 megawatts and 170,000 GPUs in Batam, Indonesia.[4] The arrangement is not wholly new; Nvidia ran a demand backstop with CoreWeave and took an equity stake in OpenAI. What is new is the packaging: the credit support and a revenue-share leg under one name, one template, one marketing page.[5]</p><p style="text-align: justify;">The bullish read wrote itself within hours: recurring revenue, a widening moat, and a supplier tying itself to customer usage rather than a single sale. The bearish read wrote itself too: circular financing, vendor-funded demand, the fiber and telecom buildouts in period costume. Both are on offer. Neither is the question that matters.</p><p style="text-align: justify;">The question that matters is why the market's strongest supplier is doing this now. The answer isn&#8217;t on Nvidia&#8217;s blog; it&#8217;s on its customers&#8217; earnings calls.</p><p style="text-align: justify;">On April 29, Google said it would begin delivering its TPUs to a select group of customers for their own data centers, its formal move into the merchant-silicon market Nvidia has long dominated.[6] Weeks later, AWS confirmed it is exploring selling Trainium to outside data centers.[7] For a decade, the hyperscalers built custom chips and kept them in-house; in 2026, both began selling them. Nvidia&#8217;s largest customers are becoming its competitors, from the top of the stack down. (I argued in May that Google is the one to watch: its chip runs its own frontier models, while Amazon&#8217;s mostly runs rented workloads &#8212; the line that separates a silicon business from a silicon cost center.[8])</p><p style="text-align: justify;">Take Nvidia&#8217;s case at its strongest. The financing gap is real; lenders have been wary of hardware whose resale value no one can yet model, so builders with genuine demand still can&#8217;t get compute funding fast enough. The inference tenants Nvidia names alongside the program (Baseten, Fireworks AI, Together AI) are well-capitalized companies, not strays. And a revenue-share cuts both ways: if a partner&#8217;s utilization falls, so does Nvidia&#8217;s cut. That is exposure to the customer&#8217;s success, not a lien on it. On its own terms, the program clears a bottleneck and aligns incentives, and that reading is not wrong.</p><p style="text-align: justify;">It is incomplete because of when it arrived. A supplier that spent a decade as the only game in town does not wander into neocloud finance in the same quarter that its two largest customers start selling their own chips. The gap is real; the timing is no coincidence; the calendar tips the scales toward defense. Whatever else it does, the program buys the loyalty of the layer beneath the hyperscalers, the independent clouds with no silicon of their own; at the moment, the layer above turns competitive. The revenue-share ties their economics to Nvidia&#8217;s; the credit-support makes Nvidia the reason some of them can exist at all.</p><p style="text-align: justify;">There is a sharper edge, and it is the subject of the next piece. Nvidia&#8217;s CFO frames the program as serving companies that have demand but cannot secure financing fast enough; even long-term commitments haven&#8217;t unlocked the capital.[9] That describes the borrowers that conventional lenders turn away. </p><div class="pullquote"><p style="text-align: center;">How far down the collateral ladder the model reaches is the open question. </p></div><p style="text-align: justify;">Firmus is a large greenfield campus and is not obviously a distressed borrower. Sharon AI, the other named partner, is another matter: it was listed on Nasdaq in February and carries a market capitalization above a billion dollars on almost no revenue, and its financing stack and headline contracts are already the subject of a detailed short-seller report.[10] I&#8217;ll take those allegations through the primary filings next. Nvidia, for the record, holds no equity in it; the hold runs entirely through the revenue-share and the credit line.[11]</p><p style="text-align: justify;">The backstop was always there; on July 1, it got a name, which is what usually happens to an improvisation just before it becomes a system. Whether it is mostly defense or mostly reach, the next deals will say. If the template stays with capital-constrained clouds, it is the base-reinforcement it looks like; if Nvidia extends it to well-funded clouds that could finance the GPUs themselves, the defensive read was wrong, and this is Nvidia annexing cloud economics wherever it can. Either way, the neocloud it piloted on is where the risk is hiding, and the filings are where the next piece goes.</p><div><hr></div><h3>Notes</h3><p>[1] NVIDIA returned approximately $20 billion to shareholders in Q1 FY2027, and its board authorized an additional $80 billion in repurchases on May 18, 2026: <a href="https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000051/q1fy27pr.htm">NVIDIA Q1 FY2027 results (Form 8-K), SEC EDGAR</a>. The $25 billion multi-tranche notes offering priced on June 18, 2026: <a href="https://www.sec.gov/Archives/edgar/data/0001045810/000119312526275783/d48176d8k.htm">NVIDIA Form 8-K, June 18, 2026, SEC EDGAR</a>.</p><p>[2] <a href="https://blogs.nvidia.com/blog/nvidia-unlocks-ai-compute-at-scale-capital-partners-to-power-ai-infrastructure-buildout/">NVIDIA Unlocks AI Compute at Scale, NVIDIA blog, July 1, 2026</a> (co-authored by CFO Colette Kress).</p><p>[3] <a href="https://www.airealist.ai/p/jensens-comecon-how-nvidia-built">Jensen&#8217;s COMECON: How Nvidia Built an Empire of Captive Clouds, The AI Realist, Feb. 14, 2026</a>.</p><p>[4] Firmus scale from NVIDIA blog [2]; Sharon AI terms (six-year collaboration, 72MW of new Australian capacity, up to 40,000 Grace Blackwell GB300) from <a href="https://finance.yahoo.com/sectors/technology/articles/sharon-ai-announces-six-strategic-112000755.html">SharonAI Holdings, &#8220;Six Year Strategic Compute Collaboration with NVIDIA,&#8221; June 12, 2026</a> (BusinessWire; corresponds to the Company&#8217;s Form 8-K filed June 12, 2026).</p><p>[5] The credit-support and backstop model packages arrangements Nvidia previously ran case by case &#8212; a demand backstop with CoreWeave and a reported ~$30 billion equity investment in OpenAI (a separate data-center lease guarantee was reported to be under discussion, not executed, as of mid-2026): <a href="https://aiweekly.co/alerts/nvidia-launches-revenue-share-model-with-sharon-ai-firmus">AI Weekly, July 2026</a>.</p><p>[6] Sundar Pichai, Alphabet Q1 FY2026 earnings call, April 29, 2026: <a href="https://www.datacenterdynamics.com/en/news/google-to-sell-tpus-to-a-select-group-of-customers-for-their-data-centers/">Google to sell TPUs to a &#8220;select group of customers,&#8221; Data Center Dynamics</a>.</p><p>[7] AWS signaled external Trainium sales in Andy Jassy&#8217;s April 2026 shareholder letter (<a href="https://thenextweb.com/news/amazon-custom-chips-jassy-letter-fifty-billion-trainium">The Next Web</a>); AWS&#8217;s Peter DeSantis confirmed exploratory talks in June 2026 (reported by Bloomberg; <a href="https://www.electronicsforyou.biz/industry-buzz/aws-weighs-selling-trainium-ai-chips-in-challenge-to-nvidia/">summary</a>).</p><p>[8] <a href="https://www.airealist.ai/p/two-chips-one-decade-one-winner">Two Chips, One Decade, One Winner, The AI Realist, May 27, 2026</a>.</p><p>[9] CFO framing that the program targets companies with demand but insufficient access to financing: <a href="https://finance.yahoo.com/technology/ai/articles/nvidia-launches-revenue-sharing-model-131824248.html">Nvidia launches revenue-sharing model, Yahoo Finance, July 2026</a>; see also NVIDIA blog [2].</p><p>[10] <a href="https://www.bleeckerstreetresearch.com/research/shaz">Bleecker Street Research, &#8220;SharonAI (SHAZ),&#8221; April 30, 2026</a>. Report authored by a short seller with a disclosed position; allegations to be examined against primary filings in the follow-up piece.</p><p>[11] SharonAI corrected its FY2025 Form 10-K, which had described NVIDIA as a &#8220;strategic shareholder,&#8221; to state that NVIDIA holds no equity securities of the company: <a href="https://www.stocktitan.net/sec-filings/SHAZ/8-k-sharon-ai-holdings-inc-reports-material-event-05a155b67636.html">SharonAI Holdings Form 8-K correction</a>.</p>]]></content:encoded></item><item><title><![CDATA[The Models That Learned Physics]]></title><description><![CDATA[Generative AI isn&#8217;t just chatbots and code. The transformer architecture has quietly generalized far past language, and the next domain it&#8217;s reaching is the physical world.]]></description><link>https://www.airealist.ai/p/the-models-that-learned-physics</link><guid isPermaLink="false">https://www.airealist.ai/p/the-models-that-learned-physics</guid><dc:creator><![CDATA[Julien Simon]]></dc:creator><pubDate>Tue, 30 Jun 2026 09:27:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!k66J!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!k66J!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!k66J!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png 424w, https://substackcdn.com/image/fetch/$s_!k66J!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png 848w, https://substackcdn.com/image/fetch/$s_!k66J!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png 1272w, https://substackcdn.com/image/fetch/$s_!k66J!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!k66J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png" width="1424" height="752" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:752,&quot;width&quot;:1424,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1918220,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.airealist.ai/i/204116576?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!k66J!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png 424w, https://substackcdn.com/image/fetch/$s_!k66J!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png 848w, https://substackcdn.com/image/fetch/$s_!k66J!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png 1272w, https://substackcdn.com/image/fetch/$s_!k66J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe10a8c9c-68ef-4a6b-b464-a3515dd1330e_1424x752.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;"><em><span>Disclosure: </span><a href="https://www.simcon.ai"><span>Simcon</span></a><span>, used here as a worked example, is a portfolio company of </span><a href="https://www.fortino.capital"><span>Fortino Capital</span></a><span>, where I am an AI operating partner. I&#8217;ve kept this piece to what Simcon has already made public, and the </span><a href="https://www.simcon.ai/en/solutions/cadmould-ai-solver-live-demo"><span>live demo</span></a><span> is open to anyone.</span></em></p><p style="text-align: justify;"><span>The conversation about generative AI has narrowed to two things it does extremely well: write text and write code. But that is a small slice of what the underlying machine turned out to be good at.</span></p><p style="text-align: justify;"><span>The transformer was built for language. Then it generalized. The same architecture that predicts the next word learned to generate images, then to model protein structures, then to forecast time series and weather. Each jump landed in a domain that looked nothing like the last, and each time the lesson repeated: feed a transformer enough diverse examples of a thing, and it learns a representation of that thing general enough to handle inputs it never saw.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!U5cV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!U5cV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png 424w, https://substackcdn.com/image/fetch/$s_!U5cV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png 848w, https://substackcdn.com/image/fetch/$s_!U5cV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png 1272w, https://substackcdn.com/image/fetch/$s_!U5cV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!U5cV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png" width="1456" height="618" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:618,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!U5cV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png 424w, https://substackcdn.com/image/fetch/$s_!U5cV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png 848w, https://substackcdn.com/image/fetch/$s_!U5cV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png 1272w, https://substackcdn.com/image/fetch/$s_!U5cV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ec8dc8b-876e-42c0-bb83-0f1e187d46df_2048x869.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;"><span>So here is the question that follows naturally and almost nobody is asking out loud: what happens when you point it at industrial engineering data and complex 3D physics? Not text, not pixels: the airflow over a car, the heat moving through a turbine, the way molten plastic fills a mold. The data that engineers generate daily by the terabyte, which never touches the public internet.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.airealist.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.airealist.ai/subscribe?"><span>Subscribe now</span></a></p><h2 style="text-align: justify;"><span>The part of AI that doesn&#8217;t read</span></h2><p style="text-align: justify;"><span>Everything the frontier labs compete on is trained on the internet: text, code, images, and video. That data is effectively a global commons: everyone scrapes the same web, so no one owns the input.</span></p><p style="text-align: justify;"><span>The physical world is the opposite. The data that describes it lives inside the companies that build things, and it never gets posted anywhere. Decades of crash tests, wind-tunnel runs, thermal cycles, material trials, and &#8212; the unsung hero of modern engineering &#8212; simulation results. Before a car, a phone case, or a medical device gets built, engineers simulate it. Simulations are used to predict and optimize the quality and cost issues that will arise during manufacturing and how the parts will behave in the real world. The result is fewer costly defects in the real world. To predict the physics, they run numerical solvers that chew through large systems of partial differential equations, which is computationally intensive. These runs are slow, expensive, and have been the backbone of industrial design for decades.</span></p><p style="text-align: justify;"><span>They are also a training set. Every solver run is a labeled example: this target geometry, these conditions, this outcome. A company that has run millions of them is sitting on something no web scraper can ever reach.</span></p><p style="text-align: justify;"><span>The bet behind Physics AI is that you can train a model on that simulation output, the way a language model is trained on text, and get a network that has, in effect, learned the physics. Not the equations. The behavior. Feed it a shape it has never seen, and it predicts the result, in seconds, without solving anything.</span></p><p style="text-align: justify;"><span>The category now has a name: Large Engineering Models. The label is deliberate. It claims the same lineage as Large Language Models &#8212; the same transformer architecture underneath, the same idea that scale and diverse data produce something that generalizes &#8212; but is trained on the physical world rather than language.</span></p><h2 style="text-align: justify;"><span>Why this didn&#8217;t work before</span></h2><p style="text-align: justify;"><span>Engineers have wanted fast simulation forever, and the idea of replacing a slow solver with a fast approximation is old. The approximations are called surrogate models, and until recently, they came with a catch that made them nearly useless for real design work.</span></p><p style="text-align: justify;"><span>A classical surrogate is fitted to a specific problem. Train it on one family of parts, and it interpolates nicely within that family. It also falls apart the moment you hand it a geometry it hasn't seen before. It learned the answers, not the physics. Engineers got a tool that was fast exactly where they didn&#8217;t need help and unreliable everywhere they did.</span></p><p style="text-align: justify;"><span>The numerical solvers had the opposite profile: accurate and trustworthy across any geometry, but far too slow to run inside a design loop. So the trade-off stood. Fast or general: pick one.</span></p><p style="text-align: justify;"><span>What changed is the architecture. The transformer &#8212; the same design that made language models work &#8212; turns out to be good at consuming large, diverse collections of physical examples and learning a representation that holds up on inputs it was never trained on. The published method these systems draw on, Universal Physics Transformers, was demonstrated on automotive aerodynamics in a 2025 peer-reviewed paper. [2] The detail that matters for a non-specialist is simple: it was built to generalize across shapes, not memorize a few. That is the wall the old surrogates hit, and it is the wall the new models are designed to go through.</span></p><p style="text-align: justify;"><span>Fast </span><em><span>and</span></em><span> general, at the same time. That is the whole claim. Everything else is engineering.</span></p><h2 style="text-align: justify;"><span>A real-life LEM you can run in a browser</span></h2><p style="text-align: justify;"><span>Abstractions are easy to oversell, so here is a concrete one: plastic injection molding, in a corner of manufacturing that most people never think about. It is how a staggering share of the plastic object parts around you were made: caps, casings, connectors, dashboards. A mold costs six or seven figures and takes months to design. Get the design wrong, and the plastic will not fill the cavity properly. And you will only find out after the steel is cut.</span></p><p style="text-align: justify;"><span>So engineers simulate the fill first. Historically, that meant a numerical solver and a wait of minutes to hours per design, which in practice limits how many variations you can reasonably try.</span></p><p style="text-align: justify;"><span>This is where domain expertise and data decide everything, and it is </span><a href="https://www.simcon.ai/en/"><span>Simcon&#8217;s</span></a><span>. The German company has developed injection-molding simulation software for over 35 years, used by manufacturing world leaders such as Bosch, Continental, Roche, and Arburg. And now they&#8217;ve built, trained, and deployed a Transformer-based model for 3D physics.</span></p><p style="text-align: justify;"><span>Their </span><a href="https://www.simcon.ai/en/solutions/cadmould-ai-solver-injection-moulding-simulation"><span>Cadmould AI Solver</span></a><span> is trained on millions of  simulation runs generated by their own numerical solvers &#8212; the proprietary ground truth I described earlier, accumulated over decades and turned into a training corpus no one else has. The architecture came from a research collaboration; the physics, the data, and the validation are Simcon&#8217;s, and the company now owns the model outright, trains it in-house, and runs the cloud infrastructure that hosts it for customers. [3] Simcon bills it as the first Large Engineering Model for injection molding. Results in seconds instead of hours, a speedup the company puts in the range of 200 to 1,000 times, across part shapes the model was never trained on. [4]</span></p><p style="text-align: justify;"><span>You don&#8217;t have to take the number on faith. Simcon put a </span><a href="https://www.simcon.ai/en/solutions/cadmould-ai-solver-live-demo"><span>research preview</span></a><span> on the open web. It runs in a browser, on a mid-tier cloud GPU, and the geometries it ships with were explicitly not in the training data, to show it generalizes rather than parrots. [5] You change a parameter, you watch the fill pattern redraw, you change it again. The hours-long loop becomes a conversation. Anyone reading this can </span><a href="https://www.simcon.ai/en/solutions/cadmould-ai-solver-live-demo"><span>try it</span></a><span> online, and you can also </span><a href="https://www.simcon.ai/en/checkout/demo/start"><span>schedule a demo</span></a><span> with the Simcon team.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VFL7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VFL7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png 424w, https://substackcdn.com/image/fetch/$s_!VFL7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png 848w, https://substackcdn.com/image/fetch/$s_!VFL7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png 1272w, https://substackcdn.com/image/fetch/$s_!VFL7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VFL7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png" width="1456" height="771" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:771,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VFL7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png 424w, https://substackcdn.com/image/fetch/$s_!VFL7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png 848w, https://substackcdn.com/image/fetch/$s_!VFL7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png 1272w, https://substackcdn.com/image/fetch/$s_!VFL7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3514b37-7a8b-4d73-a4f3-5aef022ecddf_2048x1084.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;"><span>The current model covers the filling stage of the molding process. The cooling, shrinkage, and warpage steps, which are decisive for many real-world outcomes, are on the roadmap. Accuracy is reported within a few percent of the numerical solver and improves as the training set grows. [6] The AI is a fast compass for exploration, and you can still validate the chosen design with the classical solver before cutting steel. That framing &#8212; AI to explore, trusted solver to confirm &#8212; is the sober version of the technology, and it is more convincing than a claim of replacement.</span></p><p style="text-align: justify;"><span>It is also more than a division of labor. The two engines feed each other. The model lets engineers explore thousands of designs in the time it would take the solver to check one; the solver then verifies the chosen design at full accuracy &#8212; and every such verification run is a fresh, high-fidelity training example for the next version of the model. Fast exploration surfaces the designs worth checking; precise validation turns the checks into new data; the new data makes the next model better at exploring. The loop closes in favor of whoever owns both engines. A company with only a fast model has a clever demo. A company with only a solver has what the industry already had. The advantage goes to the one running both, because each loop around widens the lead.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IbyA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IbyA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png 424w, https://substackcdn.com/image/fetch/$s_!IbyA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png 848w, https://substackcdn.com/image/fetch/$s_!IbyA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png 1272w, https://substackcdn.com/image/fetch/$s_!IbyA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IbyA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png" width="1456" height="864" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IbyA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png 424w, https://substackcdn.com/image/fetch/$s_!IbyA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png 848w, https://substackcdn.com/image/fetch/$s_!IbyA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png 1272w, https://substackcdn.com/image/fetch/$s_!IbyA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f03d1eb-df35-492f-81a1-c43e5b2df7c2_2048x1216.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2 style="text-align: justify;"><span>Why Europe, for once, is well positioned</span></h2><p style="text-align: justify;"><span>The reflex in any AI story is that the US trains the biggest models and Europe writes the rules. In language, that is broadly true. Large Engineering Models invert one piece of it.</span></p><p style="text-align: justify;"><span>The scarce input here is not compute or web text. It is high-fidelity physical data from real industrial processes that lives disproportionately within European industry. Europe&#8217;s manufacturing sector holds potentially a century or more of accumulated knowledge and data: how materials behave, how processes fail, how a good part differs from a bad one, captured across generations of engineers and now sitting in solver archives, test logs, and process records.</span></p><p style="text-align: justify;"><span>A model is only as good as the data and the domain knowledge behind it, and those don&#8217;t transfer in a deal. They sit with the companies that have spent decades generating high-fidelity physical data &#8212; most of them industrial firms, many of them European, none of them frontier labs.</span></p><p style="text-align: justify;"><span>That is the moat. It is the one input a frontier lab cannot buy, scrape, or out-compute, and it is the thing that looks like a legacy liability in the software era right up until it becomes the training corpus for an entire category. Germany&#8217;s machine builders, the automotive supply chain, the molders and tool shops: each is a reservoir of exactly the data these models need, and the web does not contain. A simulation company with 35 years of its own solver output turning into a defensible AI asset is not a fluke. It is the shape of the whole opportunity.</span></p><p style="text-align: justify;"><span>This is also why the European Commission, in its Apply AI Strategy last October, named manufacturing a strategic sector and tied sectoral AI adoption to reducing Europe&#8217;s dependence on non-EU technology. [7] The advantage is not guaranteed &#8212; owning the data is not the same as building the models or the businesses on top of them, and that gap is where most of the value will be won or lost. But the raw material sits on the right side of the Atlantic, which is not something you can say about most of the AI race.</span></p><h2 style="text-align: justify;"><span>Why synthetic data isn&#8217;t enough</span></h2><p style="text-align: justify;"><span>A common objection is that synthetic data dissolves the moat: if you can generate training data on demand, the proprietary corpora stop being scarce, and the advantage migrates back to whoever has the most compute.</span></p><p style="text-align: justify;"><span>However, this objection is weaker than it looks. Valuable synthetic data is not random data. It is data that captures the rare failure, the edge case, the point where the physics turns nonlinear, and a part that looked fine starts to warp during cooling. Knowing which scenarios are worth generating and whether a generated sample is physically trustworthy or quietly wrong is itself domain expertise. You cannot synthesize your way past not knowing what matters. Synthetic generation doesn&#8217;t remove the need for decades of accumulated know-how; it raises the price of admission to a layer where that know-how is even scarcer.</span></p><p style="text-align: justify;"><span>A model you can access in a browser is predicting the 3D physics of parts it never saw, in seconds, and a company that knows the cost of getting it wrong is putting it in front of customers &#8212; alongside, not instead of, the solver they already trust. The first models to read and write the world got the headlines. The ones learning to predict it may turn out to matter more to the people who build things, and Europe is holding more of the raw material than it has in any other part of this race.</span></p><p style="text-align: justify;"><span>The machines are starting to learn physics. The question worth asking is, who can you trust to teach them, and on whose data?</span></p><div><hr></div><h3 style="text-align: justify;"><span>Notes</span></h3><p style="text-align: justify;"><span>[2] Benedikt Alkin et al., &#8220;AB-UPT: Scaling Neural CFD Surrogates for High-Fidelity Automotive Aerodynamics Simulations via Anchored-Branched Universal Physics Transformers,&#8221; Transactions on Machine Learning Research, accepted October 2025 (arXiv:2502.09692). The published architecture targets automotive aerodynamics computational fluid dynamics; its application to injection molding is a separate, domain-specific implementation. Code released by Emmi AI on </span><a href="https://github.com/Emmi-AI/anchored-branched-universal-physics-transformers"><span>GitHub</span></a><span>; paper on </span><a href="https://arxiv.org/abs/2502.09692"><span>arXiv</span></a><span>.</span></p><p style="text-align: justify;"><span>[3] Simcon GmbH, &#8220;SIMCON Unveils World&#8217;s First Large Engineering Model for Plastic Injection Moulding,&#8221; BusinessWire, March 18, 2026. The Cadmould AI Solver is described by Simcon as co-developed with Emmi AI on the model architecture; the training data, domain validation, and commercialization are Simcon&#8217;s per the company&#8217;s own product and scientific pages. CEO quote and product framing from the same release. </span><a href="https://www.businesswire.com/news/home/20260318680159/en/SIMCON-Unveils-Worlds-First-Large-Engineering-Model-for-Plastic-Injection-Moulding"><span>BusinessWire</span></a><span>.</span></p><p style="text-align: justify;"><span>[4] Speed range and &#8220;trained on over a million simulation trajectories&#8221; per Simcon&#8217;s product and technical pages; the 200&#8211;1,000&#215; and &#8220;up to 1000&#215;&#8221; figures are vendor-claimed and not independently reproduced. Trade-press coverage (Plastics Today, Plastics Technology, MoldMaking Technology, March 2026) repeats the &#8220;up to 1,000&#215;&#8221; figure sourced to Simcon. </span><a href="https://www.simcon.ai/en-us/solutions/cadmould-ai-solver-injection-molding-simulation"><span>Simcon</span></a><span>.</span></p><p style="text-align: justify;"><span>[5] Research preview runs in-browser on a cloud GPU; Simcon states the demo geometries are not part of the training data. Live at simcon.ai. </span><a href="https://www.simcon.ai/en-us/solutions/cadmould-ai-solver-injection-molding-simulation"><span>Simcon demo</span></a><span>.</span></p><p style="text-align: justify;"><span>[6] Filling-stage scope, roadmap to packing/cooling/shrinkage-and-warpage, accuracy reported &#8220;within 2&#8211;5% of numerical methods,&#8221; and the explicit &#8220;explore with AI, validate with classical solver&#8221; workflow are all per Simcon&#8217;s public materials and CEO statements. Accuracy figures are vendor-claimed. </span><a href="https://www.simcon.ai/en/solutions/cadmould-ai-solver-scientific-research"><span>Simcon scientific page</span></a><span>.</span></p><p style="text-align: justify;"><span>[7] European Commission, &#8220;Apply AI Strategy,&#8221; COM(2025) 723, published 8 October 2025. The strategy names manufacturing among its strategic sectoral flagships (deploying &#8220;agentic&#8221; AI to optimise production lines, targeted Q4 2026) and frames sectoral AI adoption as part of strengthening European digital sovereignty and reducing dependence on non-EU technology providers. </span><a href="https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:52025DC0723"><span>EUR-Lex</span></a><span>.</span></p>]]></content:encoded></item></channel></rss>