Special thanks to Wayne at Ornn for the data behind this piece.
Thesis. This summer’s GPU shortage split the market in two. Old GPUs are holding their value. The newest GPUs earn a scarcity premium that the market expects to fade. Memory appears to decide which side a chip lands on.
Old GPUs are holding their value. By Ornn’s estimate, an H100 bought a year ago resells for more than it cost and has returned 76% including rent. A five-year A100 rental is priced at 80% of the one-month rate.
The B200 premium is priced as temporary. A B200 now costs 21% more per unit of compute than an H100; in June it cost 18% less. Its five-year rental price is only 54% of its one-month price.
Memory appears to set the price of frontier GPUs. Since July the H200 has rented at about 1.76 times an H100, the ratio of their memory capacity.
Memory makers are taking the margin. Micron’s 84.9% gross margin is now above NVIDIA’s 75%, and NVIDIA is guiding its margin down because of memory costs.
The implication: GPUs no longer lose value along a single curve. Scarcity sets how much extra the newest chips can charge today. Memory appears to set much of what a frontier chip is worth. And an old chip keeps its value as long as it is still the cheapest way to run real work.
Old GPUs Are Holding Their Value
The A100’s 80% five-year rate is the highest of any GPU, for a contract that ends when the chip family is eleven years old. The H100 shows the same thing in dollars: bought for about $19,000 in September 2025, it would resell for $20,500 today and has earned $13,500 in rent after costs.
The likeliest source of that demand is how the AI labs handled the shortage. Instead of raising list prices, they cut usage allowances. On September 14 Anthropic set Claude Code weekly limits 17% below their summer level. OpenAI’s $20 Plus plan now allows 5–45 Codex messages per five hours on GPT-6 Astra, down from 10–100 on GPT-5.6 Sol. What users actually paid still rose. Ornn’s Token Price Index shows closed-model token prices up 13% to 59% in a month and DeepSeek’s open-weight prices down 20%.
That gap pushes cost-sensitive work to open-weight models, which anyone can run on any hardware. Ornn’s September 7 paper, The Economics of Open-Weight Inference, finds the cheapest open model at a given benchmark score costs up to 4.8 times less per task than the cheapest closed model. Closed models still lead at the very top.
On open models, old chips are the cheapest way to serve the work. On gpt-oss-120b, an A100 delivers 78% of an H100’s output for a third of the rent: $0.29 per million tokens at September 24 prices, against $0.66 on an H100 and $0.59 on a B200.
The rental data is consistent with that demand arriving. Between March and September, A100 occupancy rose from 74% to 90% while supply grew 13%, so rented A100s grew about 37%. Ornn’s own paper notes that this data cannot isolate how much of the demand comes from open-weight work.
The B200 Premium Is Priced as Temporary
A B200 does 2.28 times the computing work of an H100. In June it rented for 1.87 times as much; on September 24 it rented for 2.74 times as much, and it has stayed above the 2.28x line every day since September 9. B200 rent rose 79% in three months, to $8.01 an hour on Ornn’s GPU price index.
The market does not expect this to last. Ornn’s forward curves hold the B200 at 98% of its one-month price for a year, then drop it faster than any other GPU. Contract prices match: CoreWeave and Nebius charge about $40 million per megawatt-year for three- to six-month capacity, 1.6 to 2 times what Nebius gets on one- to three-year deals. Buyers are paying extra now because they expect new supply, most likely NVIDIA’s Vera Rubin, which began shipping in August.
Memory Appears to Set the Price of Frontier GPUs
The cleanest test is the H200, which uses the same compute chip as the H100 but has 1.76 times the memory capacity and 1.43 times the memory speed. In late June it rented at 1.45 times the H100, close to its memory-speed ratio. Since July it has rented near 1.76 times, its memory-size ratio. Long-context and AI-agent workloads need large memory to hold their working state, and renters are now paying for it.
That test covers one pair of chips, but the wider pattern points the same way. Per gigabyte of memory, the B200, H200 and H100 now rent for almost the same price, 3.6 to 4.2 cents an hour. The A100 is the exception: it has the same 80 GB as an H100 but rents for a third as much, because it lacks the speed and number formats to put that memory to frontier use. That is the most likely reason the two ends of the fleet diverged. Memory increasingly appears to anchor frontier GPU pricing, while the A100 looks priced on how cheaply it produces tokens.
Memory Makers Are Taking the Margin
DRAM contract prices rose 93–98% in the first quarter, according to TrendForce. Micron’s gross margin reached 84.9%, above NVIDIA’s 75.0%, which NVIDIA expects to fall to 71–72% because memory costs are “headed even higher into next year.” SK hynix earned a 76% operating margin, a level few chipmakers of any kind reach. The biggest repricing may still be ahead: high-bandwidth memory is sold on annual contracts, and TrendForce expects those to reset sharply higher in 2027.
The Shortage Is Broad
Every seller that publishes prices raised them. AWS raised its reserved GPU block rates about 20% on July 1. CoreWeave raised prices about 25% (reported). Nebius lists the B200 at $8.50 from October 1, up from $7.15. Oracle renewed expiring GPU contracts at a 20% premium, with 97.9% of its AI capacity in use. Microsoft, Alphabet and Amazon all say demand exceeds supply, and NVIDIA calls its growth outlook “supply-constrained.”
What It Means
1. Value GPUs by generation, not with one depreciation curve. Expect B200 rent to step down after about a year. Value an A100 on whether it stays the cheapest way to serve open-weight models, not on its age.
2. The biggest memory repricing is still ahead. High-bandwidth memory contracts reset in 2027, which keeps pressure on GPU costs and NVIDIA’s margin.
3. Lock in longer contracts if your demand is steady. Short-term capacity costs up to twice as much, and long-term prices already assume relief after the first year.
4. Run patient workloads on open models and older chips. Agents, batch jobs and reinforcement-learning runs cost less than half as much per token on an A100 as on an H100.
5. Watch three numbers. B200 rent falling below 2.28 times the H100’s means the Blackwell shortage is easing. H200 rent falling below 1.76 times the H100’s means memory is no longer the constraint. Falling A100 occupancy means open-weight demand is no longer absorbing the old fleet.
What Could Break This
New supply could arrive faster than expected, from the Rubin ramp, more chip-packaging capacity or Meta’s reported plan to rent out spare GPUs. The data also has limits: Ornn’s spot index covers on-demand rentals, a small part of a market that runs mostly on long contracts; its forward curves are analyst estimates, not trades; and Ornn rents GPUs and sells this data, which its paper discloses. The H100 return figures are Ornn’s estimates. And cheaper new inference hardware could pull open-weight work off old chips.
Data Gravity covers AI infrastructure, compute economics, and durable software systems. Related coverage: How Long Does a GPU Last?, How Much Does an NVL72 Actually Cost?, Who Makes Money When Inference Gets 10x Cheaper?, and Why the Best GPU Doesn’t Always Win.
Method: Ornn Data GPU price index pulled September 24, 2026 (daily 20:00 UTC settlements, June 25 to September 24; H200 measured from June 26 because its June 25 print is a one-day outlier). Token index compares 15-day averages. Forward marks, occupancy and throughput inputs are from Ornn’s paper of September 7, 2026. Spec ratios use NVIDIA’s published dense BF16 throughput and HBM figures.















Great piece ! And take that Michael Burry .
Thx for this