Frontier weights went free and too large to self-host, so the value moved to whoever serves them at production throughput.
October 02, 2026

The change has been building all year, but the summer data makes it concrete. Hugging Face's biannual state of open models report, published in mid-August, lays out a picture where the model itself stopped being the point.
The economics inverted. When a lab ships a frontier open-weight model under a permissive license, the model becomes a download rather than an asset. The platform that serves it at scale, with the right hardware, throughput, and latency, is the asset.
The headline is hard to overstate. In almost every month of 2026, the largest and most capable open model came from a Chinese lab, and it was larger than anything an American lab shipped.
China's monthly ceiling ran between 754 billion and 2.78 trillion parameters. American labs stayed under 130 billion parameters in five of seven months. The exceptions, NVIDIA's Nemotron 3 Ultra at 561 billion and Thinking Machines' Inkling at 952 billion, sit closer to the Chinese ceiling but still trail it.
This is a reversal from the old progression, where labs started small and climbed toward the top. Several Chinese labs skipped the climb entirely. Xiaomi, Ant Group, and Meituan all cleared a trillion parameters this year. Twelve months ago, all three were strangers to open weights.
Building large stopped being a differentiator. It became table stakes.
What makes that possible is the community's quantization layer. A lab ships a trillion-parameter model, and within days the ecosystem has GGUF builds and quantization recipes that make it runnable on hardware a single team can afford. The model is large on release day and pragmatic within a week.
This is the finding I keep coming back to. Hugging Face took the top 25 model repositories by downloads and the top 25 by likes over the same window. Exactly one repository appears on both lists.
The two numbers measure different acts. A like records that a release mattered, and it flows to frontier models in the weeks after they ship. A download records that something is wired into a pipeline that runs on a schedule, and it accrues to small, stable models over years.
Treating attention as a proxy for adoption is the single most common mistake in coverage of the model ecosystem.
The data on who moves volume makes the same point. Chinese frontier labs are the only accounts where the large-parameter band carries the downloads. MiniMax records essentially all of its 2026 downloads above 70 billion parameters, Moonshot 88 percent, DeepSeek 55 percent, Z.ai 39 percent. Large American accounts look the opposite, with Google, Microsoft, and IBM recording almost all of their volume below the 70 billion mark.
The counterintuitive winner is Qwen. Its full-spectrum strategy, models from under 1 billion up to the 2.4 trillion Qwen 3.8 Max, recorded 2,045 million downloads, about 55 times Moonshot's 37 million. Breadth wins when the goal is to become the family developers build on.
Downloads tell part of the story. The derivative count tells the rest.
Qwen-based models now account for 151,448 derivatives on Hugging Face, a footprint 2.6 times Meta's total and 4.7 times Llama's specifically. The third-largest source is Unsloth, a community account shipping quantized and fine-tuning-ready builds, many of which extend the Qwen ecosystem further.
The pace matters as much as the total. Qwen derivatives grew at roughly 180 to 210 new repositories per day through the first seven months of 2026. Adoption at that rate comes from a working loop: a broad family attracts developers, developers produce derivatives, and the derivatives pull in the next wave.
Of those 151,448 derivatives, almost all are downstream work from other developers. Even among the 28,531 GGUF conversions of Qwen models, Qwen itself published only 54. The ecosystem built the distribution layer. The lab shipped weights and stayed out of the way.
That is the template for how value migrates. The weights anchor the ecosystem, and the community, plus the platforms that serve it, capture the activity that follows.
The license data is where the strategic point lands. Of 178 Chinese releases above 20 billion parameters this year, 59 percent ship under Apache 2.0 and 22 percent under MIT. Commercial restrictions are rare.
That is the opposite of what a licensing business would do. Chinese labs license their largest models as permissively as their smallest, and more permissively than American labs license theirs. On the American side of the same size band, 29 percent is Apache or MIT, 41 percent sits under custom terms, and 30 percent ships with terms left unspecified.
When the frontier model and the budget model share the same license, the model stops being the thing you pay for. The competition moves downstream.
The recent shift toward heavier terms on the very largest releases, Kimi K3 and Qwen 3.8 2.4T both added revenue-share clauses, marks the beginning of a counter-trend. Watch that space. For everything below the absolute ceiling, permissive is the default.
For all the attention on trillion-parameter releases, the day-to-day workload runs small. Among models that declare a parameter count, those under 1 billion take 83 percent of all-time downloads, while everything above 100 billion takes 1 percent. Restricting to 2026 changes the picture only slightly, with 3 percent of volume going to models above 70 billion.
So how does a trillion-parameter model reach anyone? Through llama.cpp.
The ceiling moved with the runtime. The July snapshot carries GGUF builds of DeepSeek-V4-Flash at roughly 284 billion parameters and Kimi-K3 at roughly 2.8 trillion. Local inference used to mean an 8 billion parameter model on a laptop. It now means a trillion-parameter mixture-of-experts spread across a few machines.
The growth numbers confirm where the energy is. Repositories declaring the gguf library rose 464 percent, Apple's mlx 148 percent, and lerobot 194 percent, against 16 percent for transformers and 21 percent for diffusers. The runtime layer, the part that decides where a model can physically run, is growing three to seven times faster than the modeling core.
The model is the headline. The runtime is the business.
The download distribution repeats the same lesson at the model-family level. Qwen records 39.6 million GGUF downloads a month, nearly twice Gemma's 20.8 million and more than five times Llama's 7.5 million. That gap holds even though Llama-derived GGUF repositories slightly outnumber Qwen's on the Hub. Same shelf space, a fifth of the traffic. Availability alone carries little weight; a family that fills the practical layer moves the volume.
The two organizations publishing the most new open models this year are the companies that make the hardware: AMD and NVIDIA, each with more than 200 new model repositories, far ahead of the rest of the field.
That is a distribution and optimization play. A model tuned for your hardware and freely available is the clearest proof the hardware works. AMD's conversion work makes trillion-parameter models run efficiently on its silicon. NVIDIA's Nemotron family proves its own stack. Open models became a way to move accelerators.
The same competition runs in reverse inside China, where open models increasingly target domestic chips.
For the inference layer, this is the signal. When hardware vendors treat open models as their sales engine, the demand they create lands on someone's infrastructure. Serving that demand is the job someone has to do.
If the weights are free and the models are too large to run locally at production scale, then the scarce resource is the inference stack.
That is the quiet conclusion of the whole report.
Layer | Where value sat in 2025 | Where it sits now | Proof from the report |
|---|---|---|---|
Weights | Licensed asset | Free download | 81% of Chinese releases above 20B ship Apache 2.0 or MIT |
Attention | Benchmarks and likes | Still with frontier labs | Likes flow to releases for a few weeks, then move on |
Adoption | Assumed to follow attention | Small, stable models in pipelines | One repo overlaps the top 25 by likes and by downloads |
Distribution | The lab | The community | Qwen published 54 of 28,531 Qwen GGUF builds |
Runtime | Modeling frameworks | Local and edge runtimes | gguf repos up 464% vs 16% for transformers |
Serving | Afterthought | The product | Trillion-parameter MoE needs a stack, and that stack is the scarce part |
The value that used to live in the model now lives in:
Throughput. Serving a 2.8 trillion parameter model with 104 billion active parameters at production latency is an engineering problem, one that rewards the platforms that solve it well.
Placement. The gap between a model you can download and a model you can call with low latency, autoscaling, and a sensible bill is the product.
Reliability. Once a model is wired into a pipeline that runs on a schedule, uptime and consistency outweigh novelty.
The report frames the race as a marathon rather than a sprint. Scores and likes turn over every week. A model family that becomes embedded in the infrastructure is the durable exit. That happens through inference endpoints rather than downloaded weights sitting in a repository.
The practical takeaway for a team choosing a model today has three parts.
First, separate the decision about the model from the decision about the serving layer. The model landscape is fluid, with frontier capability rotating between labs on a monthly cadence. The serving layer is the part you commit to.
Second, weight the full-spectrum families. If a lab covers the range from edge to frontier, you can standardize on one toolchain while letting individual workloads pick their own size. That is an operational advantage a frontier-only portfolio lacks.
Third, treat throughput as the purchase criterion. When the weights carry zero license cost, the thing you optimize for is tokens per second per dollar of infrastructure. A platform that sustains high utilization on the largest open models delivers value that a permissive license by itself leaves on the table.
This is where a GPU cloud earns its place. The models are open and the licenses are permissive, which means the differentiator is whether you can run them at production scale.
GMI Cloud hosts a broad catalog of open-weight models ready for inference, covering the frontier releases and the full-spectrum families alike, on high-performance GPUs built for consistent throughput. Check the model catalog to see what serves today, read the developer docs for the quickest path to a live endpoint, and scan the engineering blog for the patterns our team uses in production.
The weights moved into a commodity market. The value moved to the people who serve them. Position yourself where the value is.
Join the GMI Cloud builder community on Discord or find us on X at @gmi_cloud.
Roan Weigert
DevRel Lead @ GMI Cloud
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
