Nvidia gives away the software you need to run an inference cloud: TensorRT-LLM, Dynamo and NIM, plus the open-source vLLM and SGLang that it contributes to. The franchisee buys the GPUs and brings the data center. That is Nvidia’s franchise model: Nvidia supplies the technology stack, and someone else supplies the capital, power and operations.
Google, AWS and Microsoft own all three layers of their clouds: the data centers, the compute and the software. Nvidia’s franchisees, such as CoreWeave, Nebius, Crusoe, Lambda, Nscale and GMI Cloud, control and finance the first two and run Nvidia’s software on top.
That split is also why serving software is so hard to beat. Every franchisee gets the same software for free, and it improves every few weeks because chipmakers and model labs pay to make it better. A company selling its own serving software on rented GPUs has to beat that free stack while paying 2–3x more for compute.
How the franchise works
Marriott and McDonald’s show how the model works. Marriott owns or leases 50 of its 10,082 hotels, and McDonald’s franchises about 96% of its 46,028 restaurants. Both supply the brand and the operating manual, and McDonald’s also runs the supply chain. The owners put up the capital and run each hotel or restaurant to the parent’s standards. Every location uses the same manual, so owners compete on location, cost of capital and execution, and the parent is paid on every room or every sale.
Nvidia runs AI clouds the same way. Like McDonald’s, it controls the supply: every franchisee buys Nvidia’s GPUs. It also supplies the operating manual: the serving software. And it supplies some of the demand and credit, such as its agreement to buy $6.3B of CoreWeave’s unused capacity and its $1.5B leaseback of Lambda’s GPUs. The franchisees supply the balance sheets, the power and the operations. Nvidia is paid on every GPU they buy.
The franchisees use the manual. When Nvidia released version 1.0 of Dynamo in March 2026, the cloud partners adopting it included CoreWeave, Crusoe, GMI Cloud, Nebius, Nscale and Together AI.
How Many Neoclouds Can There Be?
People often ask me how many neoclouds the market can support. The count is the wrong question. Nobody sizes Marriott by counting its hotel owners; what matters is how many rooms fly its flag. For Nvidia’s franchise, the questions are how much AI compute runs on Nvidia GPUs outside the big three clouds, and how much revenue and market value that slice captures.
On the buying side, the slice is already big. In Nvidia’s latest quarter, 45% of its $89B of data center revenue came from customers outside the hyperscalers, a group that includes neoclouds, sovereigns and enterprises. Nvidia expects that share to reach about half, and expects its neocloud partners to end 2026 with 8GW of capacity, up from about 3GW a year earlier.
On the revenue and value side, it is still small. Google, AWS and Microsoft took 63% of the $143B cloud market in the second quarter of 2026. Synergy estimates neoclouds earned $25B in all of 2025, about 6% of the market, though their revenue was growing more than 200% a year by the fourth quarter. The largest franchisees are worth about $195B combined, counting Crusoe’s latest private round and Nscale’s IPO target. That is under 4% of Nvidia’s market value.
So the size of the franchise layer comes down to two shares: how much AI compute runs on Nvidia GPUs rather than the big three’s own chips, and how many of those GPUs sit with franchisees rather than inside the big three. TrendForce expects about 70% of AI servers shipped this year to use GPUs and 28% to use custom chips such as Google’s TPUs. The number of franchisees only decides how the slice is divided, and it tends to go to whoever has the cheapest capital. By my estimates, CoreWeave’s revenue run-rate is nearly twice that of the next three franchisees combined.
Why the manual is hard to beat
Serving software sits between chips and models, and both sides want it free. Nvidia sells more GPUs when tokens get cheaper. AMD and Intel need the open engines to run well on their chips to sell them. Model labs need their models running everywhere on launch day. So all of them pay engineers to improve the same code.
The numbers show it. About 3,600 people committed code to the four projects in the past year, roughly 100 commits a day. Engineers with AMD email addresses made 1,425 commits to vLLM and SGLang, nearly three times Nvidia’s 514. No serving startup can replicate that aggregate engineering effort internally.
Advantages don’t last long. In February 2025, DeepSeek published the software behind its own serving system, one of the most efficient in the world. The open projects adopted the pieces within days and weeks. By May, the SGLang team had rebuilt the whole system on 96 H100s and reached 80–93% of DeepSeek’s throughput. New models from OpenAI, DeepSeek and Moonshot now run on the open stack on launch day.
The stack also keeps getting faster on GPUs that are already installed. Software alone made the same Nvidia GPUs 1.7–5x faster within months, and every franchisee gets those gains the day they ship, at no cost.
The renter’s math
Serving software has to overcome the cost of the GPU underneath it, and most serving companies rent. By my estimates, a Vera Rubin GPU costs about $3.71 an hour to own over five years and about $11 to rent. On an H100, a one-year lease costs about 2.3x what an owner pays. So a renter needs 2.3–3.0x the owner’s throughput just to break even on cost per token, against an owner running the same free software.
Assumptions: the own-vs-rent model from “Who Wins in Inference”: about $9M per 72-GPU Vera Rubin rack, fully financed at 9% over five years; colo at about $190 per kW-month; power at about $0.09 per kWh; 25% residual value; rent of $11 for every GPU-hour. H100: SemiAnalysis’s April 2026 one-year lease midpoint ($2.40 an hour) against an owner’s cash cost plus debt service ($1.06).
Proprietary engines have pushed the field forward, and the open projects adopted their best ideas. The most recent published same-hardware comparison I found, on B200, shows a gain of about 10%. A 10% edge on a 3x cost gap still leaves a renter paying 2.7x more per token.
That changes the burden of proof. Proprietary software doesn’t merely need to benchmark faster. On rented GPUs, it needs a sustained multiple of performance before software differentiation overwhelms the economics of ownership. A platform renting from a franchisee pays its margin, then has to out-engineer a manual the franchisee’s own customers already use.
Where software still pays
Proprietary software still wins where the free stack is thin. The value doesn’t disappear; it moves above the raw serving engine, into model customization, workflow, reliability, enterprise features and demand aggregation, or below it, into owned capacity. Fireworks, Together and Baseten have built large businesses helping customers tune and serve their own models. New chips are another opening, since their software is younger and improves fastest. And owners can build on top of the stack: Crusoe built its own cluster-wide caching layer.
Routers face the same logic from the other side. They own the demand, not the supply, and in a shortage the GPU companies serve their own contracted customers first. Jessie Dong’s version:
routers (like openrouter, vercel ai gateway) want guaranteed gpus too. they own the demand, NOT the supply
the gpu companies they use (together, fireworks) will prioritize their own customers over the router’s customers. it’s a great model when compute is easy to get, and a horrible one now when compute is hard to get.
it’s a really dangerous place to be!
What it means
The rule underneath all of this: when software sits between two highly profitable layers, both layers have an incentive to make it free. Chipmakers want serving software cheap because it sells chips. Model labs want it cheap because it distributes their models. When the software in the middle is free, the money moves to whoever owns the GPUs it runs on.
Backing a proprietary engine on rented GPUs is a bet that one company can outrun Nvidia, Nvidia’s rivals and every model lab, on every model, from a 2–3x cost handicap. I would rather back the franchisees.
Everyone gets the same manual. The returns go to whoever owns the GPUs.
Sources
Commit data: Data Gravity analysis of the git histories of vLLM, SGLang, TensorRT-LLM and Dynamo, main branches through September 28, 2026; merges and bots excluded, identities merged by name and email, company counts from email domains.
Franchise model: Marriott Q2 2026 results, CoreWeave capacity agreement, Lambda leaseback, Dynamo 1.0, DGX Cloud Lepton
McDonald’s: Q2 2026 10-Q
DeepSeek and day-0 support: DeepSeek inference system overview, SGLang rebuild, gpt-oss, DeepSeek-V4 on SGLang, Kimi K3 on vLLM
Software gains: Blackwell on InferenceMAX, MoE gains on GB200, MLPerf v6.0 co-design, InferenceX v2, DeepSeek-V4 day 0 to day 43, DeepSeek-V4 on GB300 with SGLang
Proprietary engine benchmarks: Together on B200, Together Inference Engine, Fireworks FireAttention, Crusoe MemoryAlloy
Own vs rent: Data Gravity own-vs-rent model, Nebius price list, SemiAnalysis H100 index
Routers: Jessie Dong on X











The franchise analogy also explains why "sovereign" neoclouds look so similar underneath: the flag goes on the building and the balance sheet, but the operating manual (Dynamo, NIM, TensorRT-LLM) is the same everywhere. Fine for cost, but it leaves cost of capital, power and jurisdiction as the only durable differentiators. Do you see any franchisee building real lock-in above the serving layer, or is that simply ceded to the model labs?