The InfiniBand and Ethernet Wars
The fabric war is over and Ethernet won it on economics. The pattern the ending exposed is about to run again one tier up, and it points at where AI networking value pools next.
The short version
Ethernet didn’t outrun InfiniBand, it absorbed it. RDMA in 2010, adaptive routing in 2023, in-network collectives in 2025, each clone arriving faster than the last, until Ultra Ethernet started shipping capabilities InfiniBand never had. The share flip is done, from over 80 percent InfiniBand in late 2023 to roughly two-thirds Ethernet ever since 2025. NVIDIA’s answer was to sell the winner, and its networking business now runs near $60 billion annualized. Now the war is climbing into the rack, where AMD’s UALink-over-Ethernet systems launched this week, and the durable value is pooling at the endpoint, in the NIC, the rack-scale fabric, and a collectives software layer that nobody owns yet.
What the ending of a standards war teaches
Every large AI cluster is really two computers. The first is the one everyone talks about, the racks of GPUs. The second one decides whether the first is actually earning its keep. That’s the network stitching tens of thousands of GPUs into a single machine. When the second computer fails, the first one stops earning. At frontier scale something in the cluster breaks every few hours, and because training is synchronous, a single flapping optic or misbehaving NIC doesn’t degrade one server, it stalls all of them. On a 100,000-GPU cluster, roughly $4 billion of silicon renting for about $300,000 an hour at market GPU rates, a failure that takes thirty minutes to detect and checkpoint-restart burns $150,000 of compute. And a failure that never quite resolves, a marginal link retraining every few minutes, becomes a tax on every step of every epoch of every run.
That failure math is why the fabric stopped being a plumbing decision and became a boardroom one, and why the fight over it, between InfiniBand, the purpose-built interconnect NVIDIA acquired with Mellanox, and Ethernet, the fifty-year-old open standard, turned into the biggest open-versus-proprietary fight in infrastructure since Linux took the server. Here’s the thing though. That war has effectively been decided. NVIDIA’s own networking segment, the house that InfiniBand built, now runs at nearly $60 billion annualized and its growth is led by Ethernet products. So this piece is not really about who won. It’s about the mechanism by which the winner won, because that mechanism is about to run again one tier up the stack, and this time you can watch it with your eyes open.
One chart carries the whole argument. For twenty-five years, every technical advantage InfiniBand shipped eventually showed up on Ethernet. Track how long each one stayed exclusive and you get the war’s entire logic in five rows.
Three lessons hang off that chart, and they organize everything below. First, in networking, a proprietary advantage is a rental, not a moat, and the rent is set by how hard the feature is to standardize. Second, the clones arrive faster each cycle because the buyers switched sides, and the Ultra Ethernet Consortium is nothing more than the customers collectively funding the cloning of their supplier’s moat. Third, the moment the copier starts originating, the war is over and the front moves. Ultra Ethernet’s spray-native transport is that moment. The new front, the rack fabric and the endpoint, opened this week.
Three moats, three half-lives
InfiniBand held three real technical moats. Each one fell by a different mechanism, and each mechanism teaches you something about the next war, so it’s worth being precise about what actually happened rather than reciting the acronyms.
The first moat was latency, and it was never made of wire. A kernel-mediated TCP round trip costs tens of microseconds in context switches and buffer copies. InfiniBand’s RDMA did the same work in about one, a 40x gap that mattered enormously when a collective operation finishes only when the last message lands. But the moat was software, kernel bypass and a hardware transport, and software travels. When RoCE ported InfiniBand’s own transport onto Ethernet in 2010, the 40x premium collapsed to roughly 2x, and at the megabyte message sizes that dominate gradient exchange, 2x on first-packet latency is noise against bandwidth. The lesson is that physics moats are rare and most latency moats are actually code, which means they are one port away from evaporating.
The second moat was congestion behavior, and its fall is the one that matters most for what comes next. AI traffic is a small number of enormous, identical-looking flows, and classic Ethernet load balancing hashes whole flows onto single paths, so two elephants collide on one spine link while the rest sit idle. Meta’s published RoCE work documents the failure mode at production scale, and NVIDIA’s marketing pegs commodity ECMP at roughly 60 percent delivered bandwidth, a vendor number, but one whose direction nobody who runs these fabrics disputes. InfiniBand avoided this with central scheduling and credits. Ethernet’s first answer, PFC plus DCQCN, worked but ran on operational pain, pause storms and victim flows and heroic tuning. The answer that stuck is packet spraying, splitting every transfer across all paths and reordering at the destination, which is what Spectrum-X, Ultra Ethernet, AWS’s SRD and now MRC all do. Notice where the reordering happens. Not in the switch. In the NIC. The durable fix for Ethernet’s biggest weakness was to move intelligence out of the network core and into the endpoint, and that relocation is the hinge of this entire piece.
The third moat was computing inside the network. SHARP lets InfiniBand switches sum gradients as they flow through, roughly doubling effective allreduce bandwidth, and it survived nine years without an open equivalent because it requires co-design across switch silicon, NIC and the collectives library. That is the longest half-life on the chart, and it explains NVIDIA’s entire subsequent strategy. When your co-designed features are the only ones that hold, you push co-design deeper, which is exactly what the NVL72 rack is. The lesson reads both ways. Co-design buys the longest exclusivity, and even it expires, since UEC 1.0 specified In-Network Collectives and Broadcom’s Tomahawk Ultra shipped collective offload in 2025.
The endgame came fast once the buyers organized. The Ultra Ethernet Consortium, founded in 2023 under the Linux Foundation by AMD, Arista, Broadcom, Cisco, HPE, Intel, Meta and Microsoft among others, shipped its 1.0 spec in June 2025, and it doesn’t imitate InfiniBand so much as leapfrog it. UET sprays packets across every viable path, delivers out of order while completing in order, starts transmitting before handshakes finish, and trims congested packets into a fast-lane loss signal instead of dropping them silently, targeting fabrics of a million endpoints. Then in May 2026 the model labs took the pen themselves. MRC, from OpenAI, Microsoft, Broadcom, AMD and NVIDIA, grafts the same ideas onto existing RoCE across eight parallel data planes, and its geometry is the real tell. It supports 131,072 endpoints in two switching tiers instead of three, three hops instead of up to seven, and double the compute for roughly 20 percent more switches. Meanwhile the proof of practice had already landed, since Meta built twin 24,576-GPU clusters as an explicit A/B test, Quantum-2 InfiniBand against RoCE on Arista, and trained Llama 3 on the Ethernet one.
The receipts are in the market data. In late 2023 InfiniBand carried over 80 percent of AI back-end switch spend. Ethernet crossed half during 2024, took more than two-thirds for full-year 2025, and held that share through 1Q 2026 even as InfiniBand sales more than tripled year on year on the Blackwell Ultra ramp. Both fabrics are growing absolutely. Only one is growing structurally, and the vendor behavior underneath tells you which, since Celestica retook the top Ethernet spot, Cisco posted the quarter’s biggest share gains, and 1.6T switches are sampling for an H2 ramp, the kind of multi-vendor churn InfiniBand’s single-supplier structure cannot produce.
NVIDIA’s counter was to sell the winner
Watch what the incumbent did when it saw the chart above forming, because it’s the cleanest execution of the losing-a-standards-war playbook on record. In 2023, with Meta already on RoCE and AWS never having adopted InfiniBand at all, NVIDIA launched Spectrum-X, its own AI Ethernet, pairing the Spectrum-4 switch with a mandatory BlueField SuperNIC and doing per-packet adaptive routing with reordering at the endpoint. InfiniBand techniques in an Ethernet frame. It claims about 95 percent effective bandwidth against roughly 60 for commodity ECMP Ethernet, vendor numbers worth your skepticism, but the deployments are not marketing. xAI’s Colossus stood up 100,000 H100s in 122 days on Spectrum-X and has since grown past a reported half-million GPUs, still on Ethernet. Within two years Spectrum-X passed a $10 billion annualized run-rate and IDC ranked NVIDIA the number one datacenter Ethernet switch vendor by revenue, ahead of Cisco and Arista, in a market running above a $60 billion annualized pace.
The strategic content of Spectrum-X is not the throughput claim. It’s the attach. You can leave InfiniBand and NVIDIA still sells you the switch, the software, and above all the SuperNIC, because the performance features only light up end to end on NVIDIA endpoints. The skeptic’s line is that Spectrum-X is open the way a hotel minibar is convenient, and the skeptics are right, which is the point. NVIDIA read where the intelligence was moving and made sure that even its concession to open standards kept the endpoint proprietary. When the same company then co-authored MRC and opened NVLink Fusion to partners like Marvell, the pattern became unmistakable. Every NVIDIA response to openness protects the same square on the board, the endpoint.
The war doesn’t end, it moves up a tier
Here’s the map that explains why nobody gets to relax. An AI datacenter runs three networks, and they differ by roughly 40x in bandwidth per GPU at each boundary. Inside the rack, the scale-up domain moves 1.8 TB/s per GPU over copper NVLink and the rack behaves like one accelerator. One tier out, the scale-out back-end gives each GPU a 400 or 800G port into the Clos fabric, and that middle band is where the entire war just described was fought. The front-end tier below was always Ethernet and never contested.
Whoever owns a tier taxes everything that crosses it. Every fabric vendor is now trying to move up one band.
The second front opened on schedule, and this week it went live. AMD launched Helios on July 20 with Microsoft committing to deploy it at scale on Azure, a 72-GPU MI455X rack whose scale-up fabric tunnels UALink over Ethernet on switch silicon co-designed by Broadcom and HPE’s Juniper, with Pensando Vulcano UEC NICs handling scale-out and mass production ramping toward Q2 2027. Run the absorption chart’s logic against it. NVLink’s scale-up exclusivity dates to 2014-era SLI lineage but its modern form to 2016, the open clone arrived as a spec in April 2025 and as shipping hardware in July 2026, and NVIDIA pre-emptively opened NVLink Fusion to partners rather than wait to be cloned. The half-lives aren’t just shrinking. The incumbent now starts the absorption itself, which tells you it believes the pattern too.
The tier above the rack opened at the same time. Frontier training now spans buildings and metros, so NVIDIA shipped Spectrum-XGS for what it calls scale-across, with CoreWeave first in line, and Broadcom built Jericho4 for the same job. The corroboration that matters comes from Arista, which expects scale-across to contribute at least a third of its $3.5 billion 2026 AI target after describing the segment as virtually nonexistent a year ago. Inference pushes in the same direction, since disaggregated prefill and decode keep the heaviest traffic inside the rack and turn the middle tier into KV-cache plumbing that plain Ethernet handles fine. Both ends of the stack are pulling value away from the tier InfiniBand just spent three years defending.
Follow the money to the endpoint
Put every public company selling into this war on the same annualized basis and the growth ranking is itself an argument. NVIDIA’s networking segment runs near $60 billion annualized, up 199 percent, growing faster than its compute business. Broadcom’s AI semiconductors run at roughly $43 billion annualized, up 143 percent, with guidance implying a $64 billion pace this quarter and a backlog above $30 billion. Credo, which sells reliability-first copper cables into hyperscale racks, tripled fiscal 2026 revenue to $1.3 billion. Arista grew 35 percent and guided 2026 to $11.5 billion. Cisco’s datacenter switching, the pre-AI networking economy, grew 9. Sum the disclosed pieces, NVIDIA’s networking line, the roughly 40 percent of Broadcom’s AI revenue that is networking, Arista’s AI business and Credo, and AI networking grosses roughly $80 billion annualized before counting optics, cables, private vendors or the white-box builders. Two honest deductions apply, since the fiscal calendars don’t align perfectly and some of Broadcom’s switch silicon ships inside Arista and white-box systems and gets counted twice, so call the clean number $70 billion and change, still well over double what the entire datacenter switch industry grossed in a year as recently as three years ago. A component category became an industry in thirty months.
Now look at composition, because two details in those reports carry the thesis. Broadcom says networking is about 40 percent of its AI revenue and guides that mix toward 30 as custom accelerators ramp, which means the switch is becoming the smaller half of the AI story at the merchant switch champion itself. And the arithmetic of MRC makes the same point at the fabric level. When an MRC-style network doubles its GPUs, NICs double with them, one per accelerator, while the flattened two-tier topology needs only about 20 percent more switches. The endpoint scales with compute. The core deliberately doesn’t. That’s not a side effect, it’s the design goal of every post-2023 fabric, and it redraws where the revenue pools.
Size the two prizes and the migration stops being abstract. Call accelerator shipments 6 to 8 million units a year by 2027, across NVIDIA, AMD and the custom XPU programs, with one AI NIC apiece at $1,500 to $2,500 as 800G becomes the floor. Those are my assumptions, not a vendor’s, and they put the AI NIC pool at $10 to $20 billion a year and compounding linearly with compute, against a back-end switch market that Dell’Oro’s $100 billion five-year forecast averages to roughly $20 billion a year growing sublinearly by design. The endpoint is on course to match the core in dollars before the decade ends, and that’s before scale-up endpoint silicon, which Helios just turned into a merchant category.
The same migration is visible in the physical layer. A scale-out fabric consumes up to six optical transceivers per GPU across its tiers, Ethernet optics grew about 70 percent in 2025 on top of a doubling in 2024, and LightCounting sees a path to $100 billion a year in AI-cluster optics by 2030. Reliability, not just bandwidth, is what’s being priced, since at one interruption every few hours a flapping transceiver is a six-figure event. That’s why Credo can triple revenue selling copper cables whose pitch is a claimed thousandfold reliability edge over pluggable optics, and why NVIDIA is moving optics into the switch package entirely with co-packaged Quantum-X and Spectrum-X Photonics parts arriving through 2026, claiming 3.5x better power per bit. Everyone in the physical layer is selling the same two things now, joules and uptime.
Which brings the argument to its destination, the M&A record, because the buyers have been marking this thesis to market for fifteen years.
Every layer of that table has consolidated except one. The collectives software sitting above the NIC, the layer NCCL occupies, has no independent owner, and understanding why it’s empty tells you why it won’t stay that way. Until roughly last year there was nothing to be neutral about. Fleets were homogeneous, the fabric vendor shipped the software as a control point, NCCL from NVIDIA, RCCL as AMD’s port, and the only buyers rich enough to fund an alternative built private ones instead, NCCLX at Meta, MSCCL at Microsoft. Great abstraction layers only emerge when the hardware underneath fragments. VMware needed commodity server sprawl, Terraform needed multicloud. AI fabrics crossed that line in the last twelve months, with UEC NICs shipping from three vendors, UALink racks arriving alongside NVL72s, and neoclouds running InfiniBand and Ethernet estates side by side. An operator buying Helios and Vera Rubin systems in the same year now owns a fabric zoo with no common software plane. That is a company waiting to exist. Sizing it honestly requires an analogy rather than a bottom-up, because the category doesn’t exist yet. Infrastructure software that manages a hardware layer has historically captured low single digits of the spend beneath it, Datadog runs at roughly one to two percent of the cloud bills it watches, VMware at its peak took closer to ten percent of the server market it abstracted. Two to four percent of software attach on a $70 billion fabric complex is a $1.5 to $3 billion annual revenue pool by the decade’s end, concentrated in whoever wins the neocloud and enterprise tier first, and the founder pool to watch is the short list of people who have actually shipped NCCL, NCCLX or MSCCL.
The layer comes with a data business attached, which is what makes it venture-scale rather than a feature. A frontier cluster takes an interruption every few hours, Meta logged 466 across 54 days of Llama 3 training, and attributing a stalled step to a flapping link, a mistuned rail or a failing NIC is still done with vendor tools that stop at each vendor’s border. The collectives library is the one component that observes every transfer on every fabric in the job. Whoever owns it owns stall attribution, and in a market where an hour of cluster time costs $300,000, stall attribution isn’t a dashboard, it’s a claims adjuster for the most expensive machines ever built.
Now stress-test that thesis the way the steelman below stress-tests InfiniBand, because it can fail three ways. Hyperscalers might in-house the layer, the way they in-housed the NIC itself, since Annapurna became EFA, Google built Falcon, and NCCLX and MSCCL already exist, which would shrink the independent opportunity to neoclouds and enterprises. NVIDIA might integrate it away, folding enough fabric-awareness into NCCL and enough openness into NVLink Fusion that neutrality stops being worth paying for. Or the heterogeneity could prove temporary, since if Rubin-era racks simply win the next two years, fleets re-homogenize and the abstraction moment passes the way OpenStack’s did, and silicon cycles run longer than venture timelines either way. My honest read is that this layer gets built three times, once in-house, once by NVIDIA, and once independently, and the independent version matters only if open endpoints ship at volume on schedule. Which is exactly what the watchlist is for.
The steelman and the scoreboard
The bull case for InfiniBand from here rests on four legs, and it’s a real case. Quantum ships in lockstep with each GPU generation, which is why its sales tripled on the Blackwell Ultra ramp two years into Ethernet’s supposed rout. In-network collectives are specified in Ultra Ethernet but proven only on Quantum, and frontier training pays for that maturity gap today, not in 2027. The next million buyers are enterprises and sovereigns without elite network teams, a population that buys reference architectures that work out of the box, which is how CoreWeave built a business on Quantum fabrics. And co-packaged optics lands on Quantum first in NVIDIA’s own roadmap. The reference-architecture demand is measurable, not anecdotal, since of the 45 systems new to the November 2025 Top500, 31 chose InfiniBand and only 4 chose RoCE. Declining to a third of a market that quintuples is still a growth business. The share data says the war is decided, not that the loser disappears.
A thesis you can’t falsify is a mood. Six markers over the next eighteen months, each phrased so it can prove me wrong.
InfiniBand won the opening battle because it was the only fabric that treated the network as part of the computer. Ethernet won the war for the oldest reason in infrastructure. Given enough volume, open and good-enough runs down proprietary and perfect, and then stops being merely good enough. The next war, for the rack and the endpoint, started this week with the same armies in the same formation. This time, at least, we have the pattern.
Sources & further reading
Market data · Dell’Oro, 1Q 2026 AI scale-out report (Jun 2026) · Dell’Oro, “Ethernet is Winning the War Against InfiniBand” (Jul 2025) · Dell’Oro, 2025 full-year · Dell’Oro, $100B five-year forecast · IDC on NVIDIA taking #1 in DC Ethernet switching
Public financials · NVIDIA Q1 FY2027 results (May 2026) · Broadcom Q2 FY2026, $10.8B AI revenue (Jun 2026) · Arista Q1 2026 results (May 2026) · Credo FY2026 results, revenue triples (Jun 2026)
Standards and transports · UEC 1.0 announcement · VIAVI on UE 1.0 internals · The Next Platform on MRC (May 2026) · NVIDIA on RDMA and RoCE history · SHARP in-network computing
Deployments · Meta’s GenAI infrastructure · Meta SIGCOMM ‘24, RDMA over Ethernet at scale · Llama 3’s 466 training interruptions · Spectrum-X at xAI Colossus · Colossus at 2 GW and 555K GPUs · Microsoft Fairwater
Endpoints, optics and scale-up · CNBC on the AMD Helios launch (Jul 20, 2026) · Helios UALink-over-Ethernet details · Broadcom Thor Ultra · AMD Pollara · NVIDIA’s ~$900M Enfabrica deal · LightCounting on $100B AI optics · NVIDIA co-packaged optics · Spectrum-XGS · June 2026 Top500














