Will inference become bigger than AI training?

In our data center market deck, you will find everything you need to understand the market
SUMMARY
Yes. AI inference is becoming bigger than AI training, and on several of the measures that matter most to the infrastructure market it already appears to have crossed over.
The crossover depends on what we measure. Inference already looks larger in aggregate AI compute and cloud infrastructure spending, while training was still slightly ahead in McKinsey’s latest physical data-center power estimate.
The important structural difference is repetition. Training happens in enormous bursts, but inference keeps running after a model launches and can be triggered millions or billions of times across users, applications and machines.
Reasoning models are making each inference request heavier. A short visible answer can hide long chains of internal computation, tool calls, retries and intermediate work.
Agents raise the ceiling further because one human instruction can turn into dozens or hundreds of model actions. Software can keep invoking models long after a person would have stopped typing prompts.
Falling inference costs are unlikely to shrink total demand. Better chips, caching and software efficiency are making more workloads economically viable, which encourages companies to run AI more often rather than simply spend less.
Long context, video, robotics and continuous machine intelligence all push inference toward heavier, longer-running workloads. Text chat may eventually look like one of the cheaper forms of inference.
Training is not becoming small. Frontier runs are still scaling fast, and the largest future jobs could consume several gigawatts each, but losing share to inference is very different from shrinking in absolute terms.
Hyperscaler behavior supports the same direction. Meta is deploying inference-focused custom silicon, Amazon is pushing Bedrock and Trainium, Microsoft is squeezing more throughput out of Copilot workloads, and Nvidia is selling its roadmap partly on lower inference-token economics.
Our base case is that inference reaches roughly 1.5 to 2 times the infrastructure footprint of training by 2030. If autonomous agents and continuous multimodal systems become genuinely useful at scale, that range could prove conservative.
What does “bigger than AI training” actually mean?
AI inference is becoming bigger than AI training in the sense that matters most for the infrastructure market: total compute used, money spent and electricity consumed over time.
There is still an important measurement problem. One frontier training run can use an absurd amount of computing power, while one inference request can be tiny. But training happens in large bursts. Inference keeps running after the model launches, potentially millions or billions of times a day.
We therefore looked at four measures that can give different answers: compute, cloud spending, data-center power and the size of individual workloads.
The latest numbers are already leaning toward inference. Gartner currently expects 55% of the $42 billion AI-optimized infrastructure-as-a-service market in 2026 to support inference, rising to 59% in 2027. Deloitte expects inference to account for roughly two-thirds of AI compute this year. McKinsey's physical data-center estimate is more conservative: it put global inference demand at 20.9 GW in 2025 versus 23.1 GW for training.
Those estimates are close enough to tell us where we are. The crossover is happening around now, although it has arrived earlier in compute and cloud spending than in physical power demand.
| Measure | Where inference stands |
|---|---|
| AI compute | Probably already larger |
| AI cloud infrastructure spending | Already larger |
| Data-center power | Close to training |
| Largest individual workload | Training remains much larger |
Has AI inference already overtaken AI training today?
AI inference has probably already overtaken AI training on aggregate compute and cloud spending, although training still appears slightly larger when we look at global data-center power.
Deloitte's latest technology outlook estimates that inference moved from about one-third of AI compute in 2023 to half in 2025 and roughly two-thirds today. Gartner's latest forecast reaches a similar conclusion from actual spending: 55% of AI-optimized IaaS spending is expected to go toward inference this year.
McKinsey gives us the useful counterpoint. Its data-center model estimated 23.1 GW of training demand and 20.9 GW of inference demand in 2025. Training was therefore about 10% larger on that measure.
The interesting part comes next. McKinsey expects inference demand to grow at 35% a year through 2030, compared with 22% for training. Its model reaches 93.3 GW for inference and 62.2 GW for training by 2030.
We would therefore call the market roughly at the crossover today. Anyone saying training still overwhelmingly dominates AI infrastructure is working from an increasingly outdated picture.
| Estimate | Inference | Training | What it tells us |
|---|---|---|---|
| Deloitte AI compute | ~67% | ~33% | Inference already ahead |
| Gartner AI IaaS spending | 55% | 45% | Inference already ahead |
| McKinsey 2025 power demand | 20.9 GW | 23.1 GW | Training still slightly ahead |
| McKinsey 2030 power demand | 93.3 GW | 62.2 GW | Inference clearly ahead |
If you want more recent data on this point, please see our latest data center market report.

This market map, featured in our data center market deck, highlights top companies and startups in the data center market
Why is AI inference growing faster than training now?
AI inference is growing faster because people are asking models to do much more work after training ends.
The first wave of generative AI was relatively simple. A user typed a prompt, a model generated an answer and the job ended. That could involve hundreds or a few thousand tokens.
Today's products increasingly search the web, inspect files, generate code, call tools, read the result, correct mistakes and continue working. The amount of hidden computation behind one visible answer can be many times larger.
Microsoft's latest numbers show how seriously this is affecting infrastructure. The company recently said it improved inference throughput for its most-used Copilot models by 40% through hardware and software optimization. In the same quarter, Microsoft added another gigawatt of data-center capacity while continuing to describe AI demand as capacity-constrained.
Amazon is seeing the same change at much larger commercial scale. Amazon Bedrock now has hundreds of thousands of customers, and the company says most Bedrock inference runs on Trainium. Bedrock added more customers in six months than during its first two years after launch, while customer spending in the latest quarter exceeded all previous quarters combined.
So inference demand is no longer mainly about more people opening chatbots. Existing users and applications are simply doing more computing every time they interact with AI.
Are reasoning models making AI inference much more expensive?
Reasoning models are pushing AI inference higher because they can spend far more compute solving a problem before giving the user an answer.
With older chatbots, inference was mostly about producing the next token until the response finished. Reasoning systems can keep working internally, reconsider approaches, call tools and generate much larger amounts of intermediate computation.
That changes the economics quite dramatically. A short final answer can hide a large inference job.
Coding makes this easy to see. A basic coding assistant might generate one function. A coding agent can read a repository, search through files, modify code, run tests, inspect errors, try another solution and repeat that cycle. The user may end up receiving a 30-line patch after the model processed hundreds of thousands of tokens.
Nvidia has increasingly designed its systems around this exact problem. Its Rubin platform is being marketed around a reduction of up to 10 times in inference token cost compared with Blackwell on selected workloads. Nvidia would not be making token economics one of its central hardware metrics if inference still behaved like the lightweight chatbot workload of a few years ago.
Reasoning gives the industry another scaling lever. Better answers can now come from spending more compute during deployment, not only during training.

As this chart shows, and as featured in our data center market deck, search interest in data centers has increased significantly
Will AI agents make inference even bigger than reasoning models?
AI agents could push inference much higher because one human request can trigger dozens or hundreds of model actions.
Imagine asking a chatbot for three hotels in Tokyo. That is a small inference workload. Now imagine asking an agent to plan the entire trip. The agent might search flights, compare hotels, check neighborhoods, inspect train times, change the itinerary after discovering a conflict and return later to watch prices.
One user instruction has turned into a chain of model calls.
Software engineering already gives us a working version of this pattern. Coding agents can operate for long periods, repeatedly reading files, calling terminals and testing their own work. Parallel agents can increase consumption again by exploring several approaches at once.
Amazon has even started talking about the extra CPU demand created by agentic AI. The company recently disclosed that Meta committed to tens of millions of Graviton cores partly for CPU-intensive work behind real-time reasoning, code generation and multi-step agent orchestration.
This creates a much bigger ceiling than human prompting alone. People can only type so quickly. Software can keep calling models all day.
Could cheaper AI inference actually reduce total compute demand?
Cheaper AI inference is unlikely to reduce total compute demand anytime soon because usage is growing faster than the cost per task is falling.
The efficiency gains are already huge. Microsoft recently improved throughput on major Copilot models by 40%. Amazon says Trainium2 offers roughly 30% better price-performance than comparable GPUs for its target workloads, while Trainium3 improves another 30% to 40% over Trainium2. Nvidia says Rubin can reduce inference token costs by as much as 10 times versus Blackwell on some workloads.
Yet none of these companies is talking about needing fewer data centers.
Lower costs are making previously uneconomic workloads viable. If an AI coding task becomes five times cheaper, developers can run several agents in parallel. If analyzing a company's entire document archive becomes affordable, the company can stop sampling documents and process everything. If video inference costs fall enough, applications can examine every frame.
We are already seeing that effect in Amazon's chip business. Its overall chips business has moved above a $25 billion annual revenue run rate and is growing at triple-digit percentages year over year even while successive chip generations reduce the cost of individual workloads.
Inference therefore looks more like a market where efficiency unlocks consumption than one where efficiency eliminates demand.
If you want more recent data on this point, please see our latest data center market report.

This chart, featured in our data center market deck, illustrates yearly venture capital funding for data center startups
Does longer AI context really add much inference compute?
Longer context is quietly making AI inference heavier because models increasingly read whole working environments rather than short prompts.
Coding agents can ingest large repositories. Enterprise assistants retrieve long internal documents. Research agents accumulate histories across many searches. Multimodal systems add images, audio and eventually video to that context.
The amount of information processed can therefore be much larger than the response the user sees.
Caching prevents the cost from rising in a straight line. Providers increasingly store and reuse previously processed context instead of recomputing everything each time. Anthropic, OpenAI and other model providers now price cached input differently because repeated context has become important enough to affect inference economics.
Still, the direction is pretty clear. AI applications are moving from “answer this question” toward “understand this workspace and keep working inside it.”
That gives inference another source of growth even if the number of users eventually slows.
Could AI video and robots become bigger inference workloads than chatbots?
AI video, robotics and other continuous systems could eventually consume far more inference than text chat because they process much more information for much longer periods.
Text is remarkably compact. A chatbot can generate a useful answer with a few thousand tokens. Video generation has to create large sequences of image information while keeping movement and objects consistent across time. Video understanding works in the other direction, repeatedly processing frames rather than one short prompt.
Robotics pushes this further. A useful robot has to interpret cameras and other sensors, understand its surroundings and choose actions repeatedly while it moves. Autonomous vehicles already operate this way: inference happens continuously while the vehicle is running.
The same pattern could appear in personal AI devices. A model that occasionally answers a question creates limited demand. A model that listens, watches a screen or monitors an environment for hours creates a very different workload.
We should therefore be careful about extrapolating the future inference market from today's chatbot tokens. Text may end up being one of the cheaper forms of AI inference.

This chart, featured in our data center market deck, shows how Equinix is capturing share in data centers
Is AI training actually slowing down?
AI training is still growing extremely fast, even as inference takes a larger share of total AI compute.
Epoch AI's research puts frontier training compute growth at roughly four to five times per year historically. Its latest power work says the largest training runs now exceed 100 MW and could require roughly 4 to 16 GW by 2030 if current scaling continues.
Those numbers are enormous. A single frontier training operation could eventually demand several gigawatts for months.
There are signs that the composition of model development is changing, though. More compute is going into post-training, reinforcement learning, synthetic data and reasoning rather than one giant pre-training run. Deloitte also argues that growth in traditional training compute has probably slowed from the exceptional rates seen in 2023 and 2024.
We should not confuse falling share with falling demand. Training can lose ground to inference while continuing to grow dramatically in absolute terms.
Could giant frontier training runs eventually retake the lead?
Frontier training could occasionally jump ahead after a huge new cluster comes online, but a lasting return to training dominance looks unlikely once inference spreads across enough users and applications.
The difference comes from repetition. A frontier model might be trained once over several months and then receive further post-training. After release, that same model family can serve millions of users and software agents every day.
Training also has a harder physical constraint. Thousands or hundreds of thousands of accelerators need to communicate quickly enough to behave like one machine. Epoch AI expects the largest training jobs could eventually require several gigawatts concentrated around a single workload, although distributed training across multiple sites is becoming more practical.
Inference can spread much more easily. The same model can serve requests from data centers across North America, Europe and Asia.
So a future 5 GW training run would be a spectacular infrastructure project. It would still compete against inference running continuously across an entire global fleet.
If you want more recent data on this point, please see our latest data center market report.

This chart, featured in our data center market deck, illustrates yearly funding for data center startups
Are Amazon, Microsoft and Meta actually building for an inference-heavy AI market?
Amazon, Microsoft and Meta are already spending real money around inference, which gives us much stronger evidence than another forecast about where AI infrastructure might go.
Meta is particularly clear. The company says it has deployed hundreds of thousands of MTIA chips for inference across advertising and organic content. Its next MTIA generations will cover broader workloads, but Meta says MTIA 400, 450 and 500 will primarily support generative-AI inference in production in the near term and into 2027.
Amazon gives us a second independent example. Bedrock now has hundreds of thousands of customers and runs most of its inference on Trainium. Amazon's Trainium and broader custom-chip business has expanded so quickly that its overall chips business recently passed a $25 billion annual revenue run rate. OpenAI and Anthropic have also made multi-gigawatt commitments to Trainium capacity, although those deployments span training and other AI workloads as well as inference.
Microsoft is approaching the same problem through fleet optimization. Its recent 40% throughput gain on major Copilot models came while the company continued adding gigawatts of capacity.
Three hyperscalers with very different AI strategies are optimizing around the same thing: serving far more AI work at lower cost.
| Company | What it is doing around inference now |
|---|---|
| Meta | Hundreds of thousands of MTIA inference chips already deployed |
| Amazon | Most Bedrock inference runs on Trainium; Bedrock has hundreds of thousands of customers |
| Microsoft | Improved throughput on major Copilot inference workloads by 40% |
| Nvidia | Rubin targets up to 10× lower inference token cost on selected workloads |
Will inference create a much bigger market for custom AI chips?
AI inference is making custom chips much more attractive because shaving a small amount off the cost of a workload becomes valuable when that workload runs billions of times.
Training changes quickly. Researchers experiment with architectures, numerical formats and communication patterns, so flexible GPUs remain extremely useful.
Production inference can be more predictable. Meta knows roughly what kinds of ranking, recommendation and generative-AI workloads its own services will run at huge scale. That makes it worthwhile to design MTIA around those workloads rather than buying only general-purpose accelerators.
Amazon is showing how quickly that can become a serious business. Trainium2 is largely sold out, Trainium3 has been close to fully subscribed, and Trainium commitments now stretch across major AI labs and large enterprises. Meta has separately expanded its Broadcom partnership to co-develop future MTIA generations.
This will probably be one of the biggest consequences of an inference-heavy market. A custom accelerator does not have to replace Nvidia everywhere. It only has to beat a GPU economically on a workload that its owner runs often enough.

This chart, featured in our data center market deck, compares the main business model options for hyperscale data center operators
Does more inference threaten Nvidia's AI-chip dominance?
More inference creates a real opening for Nvidia's competitors, but it is currently making Nvidia much richer at the same time.
The competitive risk comes from specialization. Google can run large workloads on TPUs, Amazon on Trainium and Meta on MTIA. Those companies control both the software and the infrastructure, so even modest savings per request can justify custom silicon.
Nvidia is responding directly rather than relying on its training lead. Fiscal 2026 Data Center revenue reached $193.7 billion, up 68%, and the company's Rubin roadmap now emphasizes inference economics heavily. Rubin promises up to a 10-fold reduction in inference token cost compared with Blackwell for selected workloads.
Nvidia has also agreed to supply Meta with millions of Blackwell and Rubin GPUs over multiple generations even as Meta scales MTIA. That combination is revealing. Custom chips can take meaningful inference workloads without removing the need for Nvidia's more flexible accelerators.
Inference probably makes the AI-chip market less uniform. Nvidia can remain the largest supplier while hyperscalers increasingly move predictable, high-volume workloads onto their own chips.
If you want more recent data on this point, please see our latest data center market report.
Why is inference becoming the commercially bigger side of AI?
AI inference is becoming commercially bigger because this is where companies repeatedly charge customers for using models rather than spending money to create them.
Training behaves much like extremely expensive R&D. A company can spend billions building a better model, but the financial return depends on whether customers later want to use it.
Inference is production. Every API request, Copilot task, search response, generated video or agent action can become a billable event or support a paid product.
The customer base is also much larger. Only a small number of companies will ever train frontier foundation models. Potentially millions of organizations can use them. A bank does not need to train a frontier model to run AI across fraud analysis, customer support, software development and document processing. An insurer can process millions of claims with AI without owning a giant training cluster.
Amazon's Bedrock trajectory gives us a concrete example. Bedrock now has hundreds of thousands of customers, added more customers in six months than during its first two years, and recently generated more customer spending in one quarter than in all previous quarters combined.
As AI moves from experimentation into day-to-day software, more infrastructure naturally ends up on the serving side.

This chart, featured in our data center market deck, shows the revenue mix across customer segments in the data center market
Will most AI inference eventually move from data centers onto phones and PCs?
Some AI inference will move onto phones, PCs, vehicles and robots, but heavy reasoning and generative workloads still look firmly tied to data centers.
Running models locally makes sense when latency, privacy or connectivity matter. Apple, Qualcomm, AMD and Intel are all putting increasingly capable neural processors into consumer devices, and small language or vision models can already handle useful tasks without calling the cloud.
At the same time, frontier workloads keep getting heavier. Reasoning models need more runtime compute. Agents need long contexts and repeated tool calls. High-quality image and video generation asks for much more memory and processing than a typical phone can provide.
Deloitte currently expects most AI computation to remain on advanced, power-hungry chips located in data centers or enterprise systems even as inference becomes roughly two-thirds of AI compute.
The likely outcome is a split. Devices handle cheap and frequent tasks locally, while difficult requests move to much larger cloud models. If billions of devices begin running local models as well, edge AI could add another layer of inference demand rather than simply taking workloads away from data centers.
What could stop AI inference from becoming much bigger than training?
AI inference would fall short of today's expectations if companies fail to find enough AI applications worth running constantly.
Agents are the biggest uncertainty. They look impressive in software engineering and research, but we still do not know how reliable they will become across high-stakes business processes. If autonomous AI remains brittle, companies may use models as occasional assistants rather than continuous workers.
Economics could also bite. Inference costs can become uncomfortable when an application runs millions of long-context reasoning tasks. Companies will keep asking whether the productivity gain is worth the compute bill.
Efficiency adds another variable. Distillation, quantization, caching, sparse models, better accelerators and smarter software can cut the compute required per task dramatically. Inference demand only keeps racing ahead if new usage grows faster than those savings.
Training meanwhile has its own upside. If another round of frontier scaling produces major capability gains, AI labs could justify even larger clusters and more expensive model-development programs.
We still think inference wins, but the gap depends much more on actual AI usage than on how many GPUs the industry can manufacture.

This chart, featured in our data center market deck, shows how hyperscale AI-ready campus technology has evolved over time
How much bigger than AI training could inference become by 2030?
AI inference becoming roughly 1.5 to 2 times larger than training by 2030 looks like a reasonable base case today, with agents and continuous multimodal AI creating meaningful upside beyond that.
McKinsey provides the cleanest physical comparison. Its model goes from 20.9 GW of inference and 23.1 GW of training in 2025 to 93.3 GW of inference and 62.2 GW of training in 2030. That would make inference roughly 1.5 times larger.
Gartner's newest forecast points in the same direction through spending. Inference represents an expected 55% of AI-optimized IaaS spending today and 59% in 2027. Its previous longer-range work had inference moving above 65% by 2029, which would imply close to twice the rest of the market.
Deloitte is already more aggressive, putting inference at about two-thirds of current AI compute.
As seen above, those forecasts come from different measurement methods, yet they all put inference on the same side of the crossover. We would need a major disappointment in AI usage or an extraordinary acceleration in frontier training for that direction to reverse.
| Measure | Current / recent position | 2030 direction |
|---|---|---|
| McKinsey data-center power | Inference ~0.9× training | ~1.5× training |
| Gartner AI IaaS spending | 55% inference | Rising share |
| Deloitte AI compute | ~67% inference | Already around 2× training |
| Our base case | Around crossover now | ~1.5–2× training |
If you want more recent data on this point, please see our latest data center market report.
Will inference become bigger than AI training?
Yes. AI inference is becoming the bigger AI infrastructure workload, and on several important measures it has already crossed training.
The evidence is much stronger today than it was even recently. Gartner now puts inference above half of AI-optimized cloud infrastructure spending. Deloitte puts it around two-thirds of AI compute. McKinsey still had training slightly ahead on data-center power in 2025, but its model has inference growing much faster and reaching roughly 1.5 times training demand by 2030.
The behavior of the companies actually buying and designing the chips points the same way. Meta has hundreds of thousands of inference-focused MTIA accelerators deployed. Amazon says most Bedrock inference runs on Trainium and its chips business has passed a $25 billion annual revenue run rate. Microsoft is squeezing 40% more inference throughput from major Copilot models while continuing to add gigawatts of capacity. Nvidia's next hardware generation is being sold partly on how dramatically it can cut the cost of inference tokens.
Training will remain enormous. Frontier runs could eventually consume several gigawatts each, and labs are still scaling model-development compute aggressively.
Inference has the larger long-term multiplier, though. Every successful model can be used by millions of people, thousands of companies, fleets of agents and eventually machines that run AI continuously. Reasoning makes each request heavier. Agents turn one request into many. Video and robotics can turn occasional inference into continuous inference.
Our base case is therefore fairly clear: by the end of the decade, inference should consume roughly 1.5 to 2 times as much AI infrastructure as training. If autonomous agents become genuinely useful at scale, even that could look conservative.
The first AI compute race was about how much hardware companies could throw at building smarter models. The much larger recurring business is now forming around how often everyone can afford to run them.

In our data center market deck, we identify pain points entrepreneurs should prioritize
OUR METHODOLOGY
This analysis tests whether AI inference is becoming bigger than AI training across the infrastructure measures that matter most: aggregate compute, cloud infrastructure spending, data-center power and the size of individual workloads.
We did not force those measures into one score. A dollar of cloud spending is not the same thing as a gigawatt of power demand, and neither maps perfectly onto a share of total compute. Instead, we looked for convergence across independent measures and paid attention to the places where they still gave different answers.
We gave the most weight to evidence tied to real infrastructure behavior: hyperscaler spending and capacity decisions, deployed hardware, workload growth, chip economics, power requirements and disclosures from the companies operating these systems. Independent forecasts were used alongside those signals rather than treated as the answer by themselves.
For the present-day crossover, Gartner, Deloitte and McKinsey provide the main quantitative anchors. Gartner measures AI-optimized infrastructure spending, Deloitte estimates the share of AI compute used for inference, and McKinsey gives the cleanest training-versus-inference comparison in physical data-center power demand.
For the forward view, we separated evidence already visible in the market from forces that could accelerate inference further. Reasoning, agents, longer context, multimodal systems and continuous machine inference were treated as growth drivers rather than assumptions required for the current conclusion.
The 2030 range is our synthesis rather than a number lifted from one forecast. We anchored it in McKinsey's physical power projection, checked it against Gartner's spending mix and Deloitte's compute estimate, then asked whether the infrastructure decisions being made by Microsoft, Amazon, Meta and Nvidia were consistent with that direction.
We also used Epoch AI to frame the continuing scale of frontier training. That matters because inference can become the larger share of AI infrastructure while training still grows dramatically in absolute terms.
For the economics of longer context and caching, we used pricing documentation from OpenAI and Anthropic. For edge-versus-cloud inference, Apple's foundation-model work provides a useful first-hand reference point showing how lighter workloads can move onto devices while heavier reasoning and agentic workloads still depend on server infrastructure.
Key sources used for this analysis include: Gartner on AI-optimized IaaS spending, Deloitte on AI compute and inference share, McKinsey on training and inference power demand, Microsoft's FY2026 Q3 earnings, Amazon's Q2 2026 results, Amazon on Trainium economics and demand, Amazon on its custom-silicon business, AWS on Meta's Graviton commitment, Meta on MTIA deployment, Broadcom on its expanded Meta partnership, Nvidia on Rubin inference economics, Nvidia's fiscal 2026 results, Nvidia on Meta's multi-generation GPU deployment, Epoch AI on frontier-training power demand, AWS on Trainium2, OpenAI API pricing and cached input, Anthropic pricing and prompt caching, and Apple's foundation-model architecture.

This chart, featured in our data center market deck, shows the revenue mix by region across Europe, Asia, North America, Africa, and South America in the data center market
Related blog posts
Who is the author of this content?
NEW MARKET PITCH TEAM
We track new markets so founders and investors can move fasterWe build living "market pitch" documents for emerging markets: AI, synthetic biology, new proteins, and more. Instead of outdated PDFs or hallucinated LLM answers, our clients get a clean, visual, always-updated view of what's really happening: key players, deals, regulations, and signals that matter. Learn more about us.