The Rise of Inference Engineering
Squeezing more tokens out of the same silicon has become its own engineering discipline — and it now decides which AI products have margins
Thesis. The price of a fixed level of AI capability falls about 10x a year. That is the most important number in AI economics, and hardware does not explain it: GPU price-performance improves roughly 2x per generation. The other ~5x per year comes out of the serving stack — batching, memory management, quantization, caching, speculation — applied to inference workloads that now consume about two-thirds of all AI compute (Deloitte) on capex that passed $500B in 2025. The deflation that makes AI products viable is an engineering output, produced by a specific and small population of people. How small? I counted: across every public careers board I could parse in August 2026, the entire visible open market for the role is roughly 58 positions — while posted salary ranges reach $850K and NVIDIA’s largest transaction ever, a reported ~$20B for Groq, was substantially a license-and-hire of one inference team. Data science went through this in 2012, data engineering in 2019, ML engineering in 2020. The same arc is running again, compressed, against a capital base a hundred times larger.
The Fourth Wave
New engineering disciplines get minted the same way every time. A layer of the stack becomes economically decisive, a job title crystallizes around it, salaries run ahead of the talent supply, and a tooling industry forms underneath. Data science ran the arc first: coined around 2008 at Facebook and LinkedIn, anointed by Harvard Business Review in 2012, postings up 256% from 2013 to 2018. Data engineering followed once pipelines mattered more than notebooks — the fastest-growing tech job of 2019 at +50% year over year (Dice), with Snowflake’s record IPO and Databricks, $134B at its late-2025 round, as the tooling wave behind it. ML engineering ran it third: Indeed’s #1 US job in 2019, postings up 344% in three years.
Four waves, one pattern: term emerges, anointing moment, posting boom, tooling industry
The fourth is forming now, and the job is easy to state: take a trained model as a fixed artifact and maximize useful tokens per second, per GPU, per dollar, per joule. Anthropic’s posting for a “Performance Engineer, Inference Systems” describes it as “sizing the gap between actual fleet performance and theoretical rooflines.” The title hasn’t settled — inference engineer, performance engineer, model serving engineer — and LinkedIn doesn’t track it as a category yet, which is roughly where “data scientist” stood in 2010. Supply is thin for a structural reason. The work sits at the intersection of GPU microarchitecture, distributed systems, and numerical methods, a combination no university program or bootcamp produces. Most of the people who can do it trained themselves inside a handful of labs, chip companies, and open-source serving projects.
Tokens Went Quadrillion
Start with demand, which has outrun every forecast. Google processed 9.7 trillion tokens a month in May 2024, 480 trillion a year later, and over 3.2 quadrillion by May 2026 — call it 330x in two years. Microsoft ran 100 trillion tokens through a single quarter of FY25, up 5x. OpenRouter, which routes across 400+ models, went from 5 to 25 trillion tokens a week in six months. What changed is that tokens stopped being a chat phenomenon: reasoning models spend thousands of tokens thinking before they answer, and agents reread their entire transcript on every step. Jensen Huang’s claim in early 2025 that reasoning AI would need “100 times more” compute sounded like salesmanship. The token disclosures since then have made it look like arithmetic.
Each gridline is 10x. When demand compounds like this, the marginal cost of a token becomes the most important line in the P&L
The Price of Intelligence Falls 10x a Year
While demand grew 330x, the price of any fixed level of capability collapsed. GPT-3-class tokens cost $60 per million in late 2021 and $0.06 three years later — 1,000x, or 10x a year (a16z). Epoch AI, measuring the price to hit fixed benchmark scores, found declines between 9x and 900x per year depending on the task, with the caveat that the steepest declines are recent and may not persist. Sam Altman stated it as an operating assumption: the cost to use a given level of AI falls about 10x every 12 months.
Where does the deflation come from? Not primarily hardware — 2x per generation doesn’t compound to 10x a year. The balance comes from the serving stack: batching, memory management, quantization, caching, speculation, the techniques in the back half of this piece. The cleanest tell is on Alphabet’s earnings calls, where the company said it cut Gemini’s serving cost 78% in a single year while capex rose. Someone engineered that number, and this piece is about the people who do.
Constant capability, collapsing price. Hardware accounts for a fraction of the slope; serving software does the rest
Inference Is Now Most of AI Compute
The workload mix flipped underneath the buildout. Deloitte puts inference at about two-thirds of all AI compute in 2026, up from one-third in 2023. Training is a periodic capital project; inference is a permanent operating cost that grows with usage, and it is what the capex increasingly buys — roughly $500B of hyperscaler spend in 2025, with 2026 guidance totaling $660–690B. The inference market itself is sized at $106B in 2025, heading for $255B by 2030 (MarketsandMarkets). This is what gives serving teams their unusual leverage. Against $100B of deployed infrastructure, a 15% fleet-wide throughput improvement frees about $15B of effective capacity. Few engineering roles have ever had a larger denominator under them.
Inference went from a third of AI compute to two-thirds in three years, on a capex base approaching $700B a year
What the Market Pays
Anthropic’s posting for an inference performance engineer carries a salary range of $350,000–$850,000. OpenAI’s equivalent posts at $295,000–$555,000, and levels.fyi puts senior IC total compensation at both labs between $600K and $1.15M+ with equity. The reported tail runs much further: Altman said Meta offered his staff “$100 million signing bonuses”; the WSJ reported Meta’s package for Andrew Tulloch — a performance engineer, not a research celebrity — at up to $1.5B over six years (Meta disputed the details); and in December 2025 NVIDIA paid a reported ~$20B to license Groq’s low-latency inference technology and hire its leadership, three months after Groq had raised at $6.9B. For calibration, the average San Francisco data scientist earned $167K at the peak of that wave.
Posted ranges, reported totals, and the reported tail — reported figures flagged as such throughout
Counting the Field
Prices like those imply scarcity, so it is worth asking directly: how many of these engineers actually exist? No labor statistic tracks the title, so I counted. On August 9, 2026 I pulled every public careers board I could parse — the ATS endpoints behind the major labs, serving companies, and infrastructure players — and tallied open roles whose title or core description is inference-serving performance: inference engineer, model serving, inference runtime, GPU kernel work, performance engineering in a serving context. Research scientists, hardware ops, and go-to-market roles were excluded. Seventeen boards were countable. The total came to roughly 58 open roles, and 24 of them sit at two companies, OpenAI and Anthropic.
That number is a floor. NVIDIA, AMD, Meta, and Mistral run JavaScript-only portals that defeat a hand count, and NVIDIA’s own search surfaces many live TensorRT-LLM and inference-performance reqs. But even doubled or tripled, the shape holds: the visible open market for the discipline that manufactures AI’s cost curve is measured in dozens of seats, set against 2026 capex guidance of $660–690B. Two details in the data say more than the total. Fireworks — a $17.5B serving company — lists zero dedicated inference-performance openings; at the top of this market, roles fill through networks before they ever reach a board. And there is no standard title: seventeen boards produced inference engineer, performance engineer, kernel engineer, model serving, inference runtime, and inference systems, with no two companies using the same one. Postings for “data scientist” looked exactly like this in 2011, the year before HBR named the job.
Data Gravity hand count, August 9, 2026. Floors marked ≥ where boards truncate; NVIDIA, AMD, Meta, and Mistral could not be counted
The Constraint: Memory, Not Compute
To understand what these ~58 postings are paying for, one piece of physics is enough. Generating a token in the decode phase means streaming essentially all of the model’s active weights from HBM to the compute units — for one token. That is about two floating-point operations per byte fetched, against an H100 that does ~990 dense FP16 TFLOP/s on 3.35 TB/s of bandwidth: the chip needs ~295 FLOPs per byte to stay busy. Decoding one request at a time uses a low single-digit percentage of the machine while the compute units wait on memory. And the imbalance is widening by design — peak FLOP/s has grown 3.0x every two years industry-wide against 1.6x for DRAM bandwidth (Gholami et al.). V100 to B200: compute up ~18x, bandwidth up ~8.6x.
Every technique below attacks this constraint in one of two ways: move fewer bytes per token, or spread each byte across more tokens.
Compute keeps outgrowing the memory that feeds it. Decode is bandwidth-bound, so unoptimized inference wastes most of the silicon
1. Continuous Batching
The highest-leverage idea in serving is scheduling. If one request can’t saturate the GPU, run many at once, so each pass through the weights produces tokens for dozens of users. The catch is that naive batching holds every request until the slowest one finishes. Orca (OSDI ‘22) fixed this with iteration-level scheduling — requests join and leave the batch between individual decode steps — and reported 36.9x throughput against NVIDIA’s FasterTransformer at the same latency. Anyscale’s public benchmark on OPT-13B decomposed the gain: continuous batching alone was worth ~8x over naive static batching, and adding vLLM’s memory management took it to 23x. Every serious serving stack is now built around iteration-level scheduling; it is the largest single efficiency lever in production inference.
One scheduling idea, 23x the throughput from the same silicon
2. PagedAttention
Big batches created a memory problem. Every in-flight request carries a KV cache — the attention keys and values of its context — which grows token by token and can reach gigabytes per sequence. Servers used to allocate it in contiguous slabs sized for the maximum possible length, and the vLLM team measured the result: 60–80% of KV memory wasted on fragmentation and over-reservation. Their fix, PagedAttention (SOSP 2023), is the oldest idea in operating systems pointed at attention: allocate the cache in small pages, on demand, shared across requests. Waste fell below 4%, the recovered memory became batch size, and throughput rose 2–4x over the prior state of the art — up to 24x over HuggingFace Transformers. LMSYS’s Chatbot Arena served the same traffic on half the GPUs. A 40-year-old OS idea, applied in the right place, halved a production service’s compute bill.
Under 4% waste means larger batches from the same HBM — the source of the throughput gains below
3. Quantization
If decode speed is set by bytes moved, the bluntest lever is smaller bytes. Weights went from 16 bits to 8 to 4 in about three years, and the surprise is how little accuracy it costs when done carefully. GPTQ showed one-shot 4-bit quantization at 175B scale back in 2022; AWQ refined which weights need protecting. By 2025 it wasn’t an aftermarket modification anymore — OpenAI shipped gpt-oss-120b natively 4-bit, putting a 117B-parameter model on one 80 GB GPU, and Blackwell added a hardware 4-bit format. NVIDIA’s own conversion of DeepSeek-R1 to it measured ≤1% accuracy degradation. Each halving of bits roughly doubles effective memory bandwidth and halves the GPUs required; a16z credits quantization with at least 4x of the historical cost decline. The skill is knowing, and proving with evals, where 4 bits is enough.
FP16 to FP4 takes a 70B model from multi-GPU to one card with headroom, at ~1% measured accuracy cost
4. Prefix Caching
Production traffic repeats itself constantly — system prompts, retrieved documents, conversation history — and agents resend their whole transcript on every step. Prefix caching stores the KV cache of shared prefixes and reuses it, turning repeat prefill compute into a lookup; SGLang’s RadixAttention automates the reuse and gets up to 6.4x throughput on multi-call workloads. What makes this technique interesting is that it escaped the infrastructure layer and became pricing. Cached input tokens sell at 10% of the base price at Anthropic and DeepSeek, 50% at OpenAI. When an optimization gets its own line on a public price sheet it has become market structure, and application developers now engineer their prompts around it. Where that cached state physically lives — and what it is doing to the NAND market — was the subject of the storage piece.
A systems optimization, surfaced as an exchange rate: a cached token costs 10–50% of a fresh one
5. Disaggregated Serving
Prefill and decode are opposite workloads sharing a GPU. Prefill is compute-bound and parallel; decode is memory-bound and latency-sensitive; a long prefill stalls every decode stream on the card. Disaggregation splits them onto separate GPU pools, ships the KV cache between them, and lets each pool scale against its own latency target. DistServe (OSDI ‘24) measured 7.4x more requests served within latency limits, and gave the field its scoreboard metric — goodput, throughput that actually meets its latency promises. Moonshot runs this architecture in production for Kimi at +75% requests on the same hardware. NVIDIA productized the pattern as Dynamo, now its flagship serving layer. From academic paper to NVIDIA’s headline software product took about two years; this field industrializes its research nearly as fast as it publishes it.
Two pools, each tuned to its own physics — and a new metric, goodput, that only counts throughput when latency promises hold
The Rest of the Toolkit
Four more levers compound on top of these five, each a variation on the same two moves. Architectural attention changes shrink the KV cache before it exists: grouped-query attention, the Llama default, cuts it 8x, and DeepSeek’s multi-head latent attention compresses it 93.3% while raising maximum throughput 5.76x — an under-appreciated input to how DeepSeek priced its API. Speculative decoding turns cheap draft tokens into verified ones with provably unchanged outputs; three research generations took it from 2–3x to 6.5x and made it hold up at production batch sizes. Kernel engineering extracts more from silicon that already exists — FlashAttention-3 lifted attention from 35% to 75% of H100 peak with identical math, and PyTorch’s gpt-fast reached 9.6x on an unmodified A100 by stacking compilation, quantization, and speculation. And mixture-of-experts moved frontier models from activating 27.6% of their weights per token to ~4–5% in two years, while distillation put o1-mini-beating reasoning in a 32B model. Every one of these went from paper to production default in under three years. The field is closer to its beginning than its middle.
The compounding levers beyond the core five — all recent, all still improving
The Tooling Wave
In every prior wave the practitioners got expensive first and the tools that leverage them got funded second. This time both are happening at once. Inside roughly eighteen months: Fireworks AI at $17.5B on more than $1B of annualized revenue; Baseten reported at $13B, five months after pricing at $5B; Together AI at $8.3B on $1.15B of bookings; Cerebras at a ~$66B first-day market cap; Modal at $4.65B; OpenRouter at $1.3B. Even the open-source engines became companies — vLLM’s creators raised a $150M seed at $800M as Inferact, and the SGLang team spun out at a reported ~$400M — while the frameworks themselves consolidated into infrastructure: vLLM runs Amazon Rufus and LinkedIn, SGLang is xAI’s default engine, HuggingFace’s TGI went into maintenance mode, and NVIDIA elevated Dynamo to a first-class product alongside CUDA.
Why the category attracts this capital comes down to unit economics. DeepSeek published a theoretical 545% cost-profit ratio on its serving fleet — actual margins are lower, since the figure assumes all traffic billed at full R1 pricing, but the disclosure showed what engineered serving efficiency does to a P&L. Alphabet’s 78% Gemini cost reduction is the same lever at hyperscale. In this cycle, serving efficiency is not a cost center. It is where AI gross margin comes from.
The tooling wave, funded in real time — before the discipline powering it has a standardized job title
What Breaks the Thesis
Three real risks, in descending order. First, commoditization from above. Dynamo, vLLM, and SGLang are absorbing the median case the way dbt absorbed the median data-engineering task, and MLOps is the cautionary precedent — the #1 job of 2019 whose middle was largely absorbed by SageMaker and Vertex within five years. If the defaults get good enough, the middle of the skill distribution gets cheap even while the top stays scarce. The difference worth watching: optimization value here scales with fleet size, and the fleets are growing 70% a year, which MLOps never had underneath it. Second, hardware relief. If HBM bandwidth scales faster than expected, or memory-centric architectures gain real ground, the memory wall softens and some of the software leverage compresses with it. The trend line points the other way so far, but it is a bet on physics roadmaps, not a law. Third, the field’s own numbers run hot. NVIDIA’s “30x” for Blackwell inference compares a rack-scale FP4 system against an air-cooled FP8 Hopper baseline; independent analysis puts chip-level gains nearer 4x, with the rest unlocked by software and system design. The gains in this piece are real and compounding, but vendor marketing routinely overstates them — one more reason the people who can measure honestly are worth what they cost.
Summary
The center of gravity in AI economics has moved from training models to serving them. Inference is two-thirds of AI compute on a capex base heading toward $700B a year, and the price of constant-capability tokens falls 10x annually — a decline manufactured mostly in software, one technique at a time. The techniques are young and still steepening: batching went to 23x, speculative decoding from 2x to 6.5x in three years, FlashAttention from 35% to 75% of peak, active parameters from 28% to 4%, and each moved from paper to production default in under three years. Against that stands the supply side this piece measured directly: roughly 58 open roles across seventeen countable careers boards, two dozen of them at OpenAI and Anthropic, with no standard job title among them. That ratio — a few dozen visible seats leveraged against hundreds of billions in deployed capital — is the whole investment case in one line. The value sits in three places: the practitioners, who are being priced like the scarce input they are; the platforms that leverage them across thousands of customers, which is where the prior waves’ Snowflakes and Databrickses came from; and the serving margin itself, which now determines which AI applications survive their own growth. The watch list from here: whether the postings census grows into a boom the way data science’s did after 2012, where the vLLM and SGLang contributors land, how far cached-token discounts deepen, and which providers can show — as Google did with 78% — that they manufacture their own margin. Squeezing more intelligence out of the same silicon is now the central economic activity of the AI buildout, and it is being done by hand, by a group of people you could fit in one lecture hall.
Data Gravity covers AI infrastructure, compute economics, and durable software systems. Related coverage: Why KV Cache and Memory Drive AI Economics and Why AI Is Becoming a Storage Workload.
















