Subscribe
Sign in
Home
Data Infra
Data Center
Security
Archive
Latest
Top
Discussions
How Nvidia Franchised the Cloud
Nvidia gives away the software to run an inference cloud. Franchises bring the GPUs and the data center. This unbundled the Big 3 clouds.
Sep 28
•
Chris Zeoli
9
1
The GPU Depreciation Curve Is Broken
Old GPUs are holding their value. New ones command a scarcity premium. Memory explains why.
Sep 25
•
Chris Zeoli
33
3
3
How Long Does a GPU Last?
Seven to nine years physically. Six on the books. The bears say three. The whole argument is about the wrong variable.
Sep 21
•
Chris Zeoli
30
1
2
The 10 Companies AI Can't Scale Without
Ten chokepoints in the AI supply chain, from the lithography tool to the landlord. Twelve tickers, $12.5 trillion of market cap, and every supply-relief…
Sep 14
•
Chris Zeoli
32
4
How the GPU Became Collateral
Three years ago nobody would lend billions against GPUs. Now the paper is rated A3 and sits in insurance portfolios. The breakthrough was structuring…
Sep 10
•
Chris Zeoli
24
1
August 2026
How Much Does an NVIDIA NVL72 Cost?
The short answer, from recent purchase orders and analyst teardowns: ~$5.0M to buy a GB300 NVL72, ~$5.7M to deploy, $240–410K a year to run, and…
Aug 31
•
Chris Zeoli
23
2
2
Why the Best GPU Doesn't Always Win
The same GPU got 19x faster between February and May. The same die does 4.4x more work in one rack than in another. FLOPS explain none of it.
Aug 28
•
Chris Zeoli
27
2
Who Makes Money When Inference Gets 10x Cheaper?
A layer-by-layer map of where the dollars land when the price of intelligence collapses.
Aug 19
•
Chris Zeoli
73
7
8
The Rise of Inference Engineering
Squeezing more tokens out of the same silicon has become its own engineering discipline — and it now decides which AI products have margins
Aug 13
•
Chris Zeoli
60
2
9
Why AI Is a Storage Workload
The GPU isn’t the bottleneck anymore — data movement is. Inference runs on accumulated state, and that state needs somewhere to live.
Aug 8
•
Chris Zeoli
50
5
2
The AI Memory Stack
AI's memory problem isn't buying more HBM. It's managing a full hierarchy — GPU cache to cold storage — with a different winner at every tier.
Aug 4
•
Chris Zeoli
45
5
July 2026
How AI Tokens Are Made
Inside the five scheduling, memory, and routing problems that turn GPU compute into output tokens — and why neoclouds keep buying the companies that…
Jul 30
•
Chris Zeoli
38
2
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts